
DeepSeek-V2는 K와 V를 latent vector로 압축하는 MLA와 Routed expert, Shared expert를 구분하는 MoE 방식으로 high inference efficiency와 low training cost를 보였다

MQA(Multi-Query Attention)는 Query는 head마다 다르게 두고 Key와 Value는 모든 head가 동일한 값을 공유하여 incremental decoding step을 가속화한다.

GQA(Grouped Query Attention)은 Query head를 G개의 그룹으로 나누고 각 그룹마다 동일한 K, V head 하나를 공유하는 방식이다.

Linear attention은 kernel feature map을 통해 메모리와 시간복잡도를 O(N)으로 만든다.

Roformer는 position 정보를 더하던 기존 방식들과 달리 회전 행렬을 곱하는 RoPE 방식을 도입하였다.

LLM의 성능을 compute, parameter, dataset size에 대한 power law로 나타낼 수 있다