[Paper review] DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Jumyung Song·2026년 9월 24일

Paper review

목록 보기
1/6
post-thumbnail

Summary

  • MLA (Multi-head Latent Attention): Keys and values are jointly compressed into a single low-dimensional latent vector, and only this latent is cached. The KV cache shrinks to the level of GQA, yet performance is actually better than MHA.
  • DeepSeekMoE: Instead of every token sharing all experts in the FFN layer, the model uses a few shared experts plus routed experts selected by affinity scores. On top of this, device-limited routing and three balance losses reduce the communication cost of expert parallelism.
  • Effect: Out of 236B total parameters, only 21B are activated per token. Compared with DeepSeek 67B, training cost drops by 42.5%, the KV cache shrinks by 93.3%, and maximum generation throughput increases 5.76×.

1. Problem & Motivation

As LLMs grow in parameter count, emergent capabilities appear and performance improves. The price is a larger training compute budget, and at inference time, generating each token requires caching the keys and values of all previous tokens. This KV cache consumes memory capacity and memory bandwidth, which in turn lowers throughput.

GQA and MQA are the standard existing remedies. They reduce the cache by having multiple attention heads share the same keys and values, either in groups (GQA) or all together (MQA). Because heads share K and V, however, both perform worse than MHA. To address this, DeepSeek-V2 proposes two techniques.

  1. Multi-head Latent Attention: An attention mechanism equipped with low-rank key-value joint compression
  1. DeepSeekMoE: Fine-grained experts with shared expert isolation

2. Core Idea ①: Multi-head Latent Attention (MLA)

2.1 Low-Rank Key-Value Joint Compression

The core idea of low-rank key-value joint compression is to avoid caching K and V directly. Instead, the model caches only a latent vector ctKV\mathbf{c}_t^{KV} from which K and V can be reconstructed.

The token embedding ht\mathbf{h}_t is compressed into ctKV\mathbf{c}_t^{KV}. At inference time only ctKV\mathbf{c}_t^{KV} is cached, and ktC\mathbf{k}_t^{C} and vtC\mathbf{v}_t^{C} are recovered through WUKW^{UK} and WUVW^{UV}.

ctKV=WDKVht,ktC=WUKctKV,vtC=WUVctKV\mathbf{c}_t^{KV} = W^{DKV}\mathbf{h}_t,\quad \mathbf{k}_t^{C} = W^{UK}\mathbf{c}_t^{KV},\quad \mathbf{v}_t^{C} = W^{UV}\mathbf{c}_t^{KV}
  • ht\mathbf{h}_t: input hidden state of the tt-th token
  • WDKVW^{DKV}: down-projection that compresses the input into ctKV∈Rdc\mathbf{c}_t^{KV} \in \mathbb{R}^{d_c}, where dc≪nhdhd_c \ll n_h d_h
  • WUK,WUVW^{UK}, W^{UV}: up-projections that reconstruct each head's K and V from the latent

Since only ctKV\mathbf{c}_t^{KV} is cached at inference, each token needs just dcd_c elements per layer, far fewer than the 2nhdh2 n_h d_h elements required by standard MHA. In practice, K and V are not even recomputed at every step. Because ktC\mathbf{k}_t^C is multiplied with the query and vtC\mathbf{v}_t^C is multiplied with WOW^O, WUKW^{UK} can be absorbed into the query projection and WUVW^{UV} into WOW^O. As a result, attention can be computed directly on the latent without explicitly reconstructing K and V.

In this way, low-rank key-value joint compression reduces memory access at the cost of more GPU computation. Is trading memory access for extra computation actually worthwhile?

The latency for a GPU to process one step can be expressed as follows.

T≈max⁡(BBWmem, FPpeak)T \approx \max\left(\frac{B}{BW_{\text{mem}}},\ \frac{F}{P_{\text{peak}}}\right)
  • BB: amount of data read from and written to memory (bytes)
  • BWmemBW_{\text{mem}}: memory bandwidth (bytes/s)
  • FF: number of floating-point operations (FLOPs)
  • PpeakP_{\text{peak}}: peak compute throughput of the GPU (FLOP/s)

During decoding, attention consists of simple multiply-add operations within matrix multiplications, so the workload is memory-bound. Increasing the FLOP count therefore does not increase TT. In other words, it is far more effective to reduce BB at the cost of a larger FF.

2.2 Query Compression: Saving Activation Memory

Queries are compressed into a low-dimensional latent in the same way.

ctQ=WDQht,qtC=WUQctQ\mathbf{c}_t^{Q} = W^{DQ}\mathbf{h}_t,\quad \mathbf{q}_t^{C} = W^{UQ}\mathbf{c}_t^{Q}

Queries are not cached, so this compression is not for inference. Its purpose is to reduce activation memory during training, that is, the intermediate values that must be stored for backpropagation.

2.3 Decoupled RoPE

RoPE multiplies keys by a position-dependent rotation matrix, which inserts a position-dependent matrix between WUKW^{UK} and the query. As a result, WUKW^{UK} can no longer be pre-absorbed into WUQW^{UQ}, and the keys of all past tokens would have to be reconstructed at every inference step, eliminating the benefit of MLA.

To solve this, MLA separates the part of the query and key responsible for positional information.

[qt,1R;… ;qt,nhR]=RoPE(WQRctQ),ktR=RoPE(WKRht)qt,i=[qt,iC; qt,iR],kt,i=[kt,iC; ktR]ot,i=∑j=1tSoftmaxj ⁣(qt,i⊤kj,idh+dhR)vj,iC\begin{aligned} [\mathbf{q}_{t,1}^R;\dots;\mathbf{q}_{t,n_h}^R] &= \mathrm{RoPE}(W^{QR}\mathbf{c}_t^Q), \qquad \mathbf{k}_t^R = \mathrm{RoPE}(W^{KR}\mathbf{h}_t) \\ \mathbf{q}_{t,i} &= [\mathbf{q}_{t,i}^C;\ \mathbf{q}_{t,i}^R], \qquad \mathbf{k}_{t,i} = [\mathbf{k}_{t,i}^C;\ \mathbf{k}_t^R] \\ \mathbf{o}_{t,i} &= \sum_{j=1}^{t}\mathrm{Softmax}_j\!\left(\frac{\mathbf{q}_{t,i}^\top \mathbf{k}_{j,i}}{\sqrt{d_h + d_h^R}}\right)\mathbf{v}_{j,i}^C \end{aligned}
  • The RoPE-applied query qR\mathbf{q}^R is kept per head, while the RoPE-applied key kR\mathbf{k}^R is shared across all heads.
  • The cached quantities are ctKV\mathbf{c}_t^{KV} and ktR\mathbf{k}_t^R, totaling (dc+dhR) l(d_c + d_h^R)\,l elements per token.
  • With the RoPE and non-RoPE parts separated, the dot product qt,i⊤kj,i\mathbf{q}_{t,i}^\top \mathbf{k}_{j,i} splits into qR\mathbf{q}^R multiplied with kR\mathbf{k}^R and qC\mathbf{q}^C multiplied with kC\mathbf{k}^C. The latter still allows WUKW^{UK} to be absorbed. Meanwhile, kjR=RoPE(WKRhj)\mathbf{k}_j^R = \mathrm{RoPE}(W^{KR}\mathbf{h}_j) is not reconstructed from the latent; it is computed directly from hj\mathbf{h}_j and cached with RoPE already applied. Consequently, no key reconstruction is needed at all.

2.4 KV Cache Comparison

AttentionKV cache per token (# elements)Capability
MHA2nhdhl2 n_h d_h lStrong
GQA2ngdhl2 n_g d_h lModerate
MQA2dhl2 d_h lWeak
MLA(dc+dhR) l≈92dhl(d_c + d_h^R)\,l \approx \tfrac{9}{2} d_h lStronger

DeepSeek-V2 sets dc=4dhd_c = 4d_h and dhR=dh/2d_h^R = d_h/2. As a result, the MLA cache is the same size as GQA with only 2.25 groups, while performance exceeds MHA. This is the paper's most important result.

3. Core Idea ②: DeepSeekMoE

3.1 Basic Structure

DeepSeekMoE consists of two main ideas.

Fine-grained expert segmentation: Each expert, an FFN obtained by splitting the original FFN into several pieces, is made smaller, and many of them are activated at once. The rationale is that the larger number of possible expert combinations allows the model to learn more diverse features.

Shared expert isolation: In addition to the experts activated differently for each token, shared experts that every token always passes through are introduced to capture common knowledge. This aims to reduce redundant knowledge learned across routed experts.

ht′=ut+∑i=1NsFFNi(s)(ut)+∑i=1Nrgi,t FFNi(r)(ut),gi,t={si,t,si,t∈Topk⁡({sj,t∣1≤j≤Nr}, Kr),0,otherwise,si,t=Softmax⁡i(ut⊤ei)\begin{aligned} \mathbf{h}_t' &= \mathbf{u}_t + \sum_{i=1}^{N_s} \mathrm{FFN}_i^{(s)}(\mathbf{u}_t) + \sum_{i=1}^{N_r} g_{i,t}\,\mathrm{FFN}_i^{(r)}(\mathbf{u}_t), \\ g_{i,t} &= \begin{cases} s_{i,t}, & s_{i,t} \in \operatorname{Topk}\left(\{s_{j,t} \mid 1 \le j \le N_r\},\ K_r\right), \\ 0, & \text{otherwise}, \end{cases} \\ s_{i,t} &= \operatorname{Softmax}_i\left(\mathbf{u}_t^\top \mathbf{e}_i\right) \end{aligned}

Each layer of DeepSeek-V2 has 2 shared experts and 160 routed experts, and 6 routed experts are activated per token.

3.2 Device-Limited Routing

Under expert parallelism, experts are distributed across multiple GPUs. The more devices a token's selected experts are spread over, the higher the communication cost, since communication frequency is proportional to the number of devices covered by the target experts.

Routing is therefore split into two stages. The model first selects the MM devices holding the experts with the highest affinity scores, and then performs top-K selection only among the experts on those devices. The paper reports that M=3M = 3 achieves performance comparable to unrestricted routing.

3.3 Auxiliary Losses for Load Balance

Three balance losses are used to prevent tokens from concentrating on particular experts or GPUs.

1) Expert-level balance loss: prevents imbalance among experts.

LExpBal=α1∑i=1NrfiPi,fi=NrKrT∑t=1T1(Token t selects Expert i),Pi=1T∑t=1Tsi,t\mathcal{L}_{\text{ExpBal}} = \alpha_1 \sum_{i=1}^{N_r} f_i P_i,\qquad f_i = \frac{N_r}{K_r T}\sum_{t=1}^{T}\mathbb{1}(\text{Token } t \text{ selects Expert } i),\qquad P_i = \frac{1}{T}\sum_{t=1}^{T} s_{i,t}
  • fif_i is how often expert ii is actually selected, and PiP_i is the average affinity score of expert ii.

2) Device-level balance loss: prevents imbalance in computation across devices.

LDevBal=α2∑i=1Dfi′Pi′,fi′=1∣Ei∣∑j∈Eifj,Pi′=∑j∈EiPj\mathcal{L}_{\text{DevBal}} = \alpha_2 \sum_{i=1}^{D} f_i' P_i',\qquad f_i' = \frac{1}{|\mathcal{E}_i|}\sum_{j\in\mathcal{E}_i} f_j,\qquad P_i' = \sum_{j\in\mathcal{E}_i} P_j

3) Communication balance loss: balances the number of tokens each device receives, since device-limited routing only bounds communication on the sending side.

LCommBal=α3∑i=1Dfi′′Pi′′,fi′′=DMT∑t=1T1(Token t is sent to Device i),Pi′′=∑j∈EiPj\mathcal{L}_{\text{CommBal}} = \alpha_3 \sum_{i=1}^{D} f_i'' P_i'',\qquad f_i'' = \frac{D}{MT}\sum_{t=1}^{T}\mathbb{1}(\text{Token } t \text{ is sent to Device } i),\qquad P_i'' = \sum_{j\in\mathcal{E}_i} P_j

Although the three losses target different things, they all share the same form. Each ff represents the actual selection ratio relative to the case where devices or experts are chosen uniformly, and the resulting losses encourage all devices and experts to be selected evenly.

3.4 Token-Dropping Strategy

Balance losses alone cannot guarantee strict load balance. During training, a device-level token-dropping strategy is therefore applied. A computational budget (capacity) is set for each device, and when it is exceeded, tokens with the lowest affinity scores are dropped first. To reduce the mismatch between training and inference, about 10% of the sequences are never dropped.

4. Pre-Training

4.1 Setup

  • Data: High-quality data from diverse sources was added, with quality-based filtering and bias removal. The model was trained on 8.1T tokens in total.
  • Model: 60 layers, nh=128n_h = 128, dh=128d_h = 128, dc=512d_c = 512. Out of 236B total parameters, 21B are activated per token.
  • Training: AdamW with a warmup-and-step-decay schedule. Training used DeepSeek's in-house HAI-LLM framework, and thanks to the small number of activated parameters, the model could be trained without tensor parallelism.
  • Long context: YaRN was applied to the decoupled key kR\mathbf{k}^R to extend the context window from 4K to 128K. Long-context capability was verified with the Needle In A Haystack test.

4.2 Evaluation

  • DeepSeek-V2 outperforms DeepSeek 67B on almost all benchmarks and ranked among the top open-source models at the time of release.
  • Training costs: 172.8K GPU hours per trillion tokens, 42.5% less than DeepSeek 67B (300.6K).
  • Inference efficiency: The KV cache is quantized to an average of 6 bits per element, and the smaller cache allows serving with larger batch sizes. As a result, maximum generation throughput increased 5.76×.

5. Alignment

5.1 SFT (Supervised Fine-Tuning)

1.5M instruction instances were used: 1.2M for helpfulness and 0.3M for safety. Data quality was improved to reduce hallucinated responses and strengthen writing ability.

5.2 Reinforcement Learning: GRPO

In conventional RL, computing the advantage requires training an additional critic model to estimate the baseline. Since the critic is typically about the same size as the policy, this hurts training efficiency. GRPO removes the critic and instead estimates the baseline from the reward statistics of GG outputs sampled for the same question.

A^i=ri−mean({r1,…,rG})std({r1,…,rG})\hat{A}_i = \frac{r_i - \mathrm{mean}(\{r_1,\dots,r_G\})}{\mathrm{std}(\{r_1,\dots,r_G\})}

Training proceeds in two stages.

  1. Reasoning alignment: Trained on code and math reasoning tasks with ri=RMreasoning(oi)r_i = RM_{\text{reasoning}}(o_i).
  2. Human preference alignment: Combines multiple rewards: ri=c1⋅RMhelpful(oi)+c2⋅RMsafety(oi)+c3⋅RMrule(oi)r_i = c_1 \cdot RM_{\text{helpful}}(o_i) + c_2 \cdot RM_{\text{safety}}(o_i) + c_3 \cdot RM_{\text{rule}}(o_i)

For training efficiency, a hybrid engine that adopts different parallelism strategies for training and inference was introduced, and large batch sizes were used to speed up inference sampling.

5.3 Evaluation Results

  • Base → SFT: Math and code performance improved substantially because of their large share in the SFT data.
  • SFT → RL: Further gains were observed on reasoning-related tasks, especially math and code generation.
  • Amount of SFT data: Prior work suggested that around 10K instances are sufficient. In DeepSeek-V2, however, reducing the data degraded instruction-following ability. The interpretation is that acquiring a given capability requires a minimum amount of data, and that amount does not approach zero.

6. My Take

6.1 Why does MLA outperform MHA?

MLA does not produce K and V directly; it compresses them once and then reconstructs them. The structure inevitably loses information, yet MLA outperformed MHA on many tasks in the experiments. What I was most curious about was why. (At inference time, weight absorption removes the explicit reconstruction step, but the representational bottleneck remains the same.)

I came up with two possible explanations.

  1. Contribution of the query? In KV-sharing methods such as GQA, the diversity of queries may matter more for performance than that of K/V. However, MLA compresses the query as well, so the query is unlikely to be the cause.
  2. The shared latent as a regularizer? Compressing K and V into a single common vector may lead the model to learn more generalized representations. Projecting into a low-rank space keeps only the most important features.

Of the two, I find the second more plausible, but the paper includes no ablation that directly tests this, so the exact cause remains an open question. Varying dcd_c and observing the change in performance could help verify this hypothesis to some extent. I also suspect that for tasks with longer contexts and much more data and samples, MHA might still achieve better raw performance when GPU efficiency is not a concern.

6.2 Why is splitting experts into finer pieces beneficial?

DeepSeekMoE uses far more, smaller experts than conventional MoE and activates many of them at once. But if the total number of parameters and the activation ratio are the same, I wondered why there should be any theoretical difference in expressiveness between splitting experts finely and keeping them large.

  • The number of combinations certainly increases, but does each expert really learn a separate sub-skill? Common information could also be learned across routed experts, and since humans cannot directly inspect this, I wondered whether sub-skills are actually learned in a separated way.
  • If some experts have overlapping roles, is there a way to distinguish or measure them?
  • Also, if the purpose of separating experts is for shared experts to learn common features and routed experts to learn finer-grained ones, would giving the experts different sizes make the features each one learns more clearly distinguishable?

6.3 Are shared experts really necessary?

I wondered whether common knowledge could be learned through routing alone, without shared experts. From the outside, it is impossible to tell which skill each expert is responsible for. An expert that has learned common knowledge might then naturally receive high affinity scores for every token and effectively behave like a shared expert. In that case, however, the balance losses would work to suppress such concentration. So it may be more accurate to view shared experts as a mechanism that resolves the conflict between the balance losses and the learning of common knowledge.

References

  • DeepSeek-AI. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434, 2024.
  • Dai et al. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models. arXiv:2401.06066, 2024.
  • Shao et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (GRPO). arXiv:2402.03300, 2024.
  • Ainslie et al. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. EMNLP 2023.

This post was translated from the original Korean version with the help of AI, so some errors may remain.

0개의 댓글