Existing position embedding methods add the relative position information of tokens to the context representation, which does not fit well with methods that decompose the computation, such as the Linear Transformer.
RoPE rotates the query/key by an angle mθ proportional to the position m. Since the inner product of two rotated vectors depends only on the relative position m−n, the relative position ends up being expressed even though the encoding is done with absolute positions.
In particular, the d dimensions are split into d/2 two-dimensional subspaces, each rotated at a different speed θi=10000−2(i−1)/d. Thanks to this design, a long-term decay property emerges in which the upper bound of the inner product decreases as the distance grows.
1. Problem & Motivation
Self-attention in the Transformer is order-agnostic, so the order information of tokens must be injected separately. How order is handled differs by model family.
RNN-based: encodes order by recursively computing the hidden state along the time dimension.
CNN-based: appears to have no position information, but there is research showing that position information can be learned implicitly through padding.
PLM-based: captures contextual representations with self-attention, but position information must be injected through a separate position embedding.
There are two main types of position embedding.
Absolute PE: encodes the position with a pre-defined function or a learnable vector and adds it to the input
Relative PE: injects the relative position information between two tokens into the attention computation
The problem is that both methods add position information to the context representation. In particular, relative PE adds a separate term corresponding to m−n for every query-key pair, so it is not compatible with linear self-attention, which computes attention in a decomposed form such as ϕ(Q)(ϕ(K)TV).
To overcome these limitations of existing position embeddings, the paper proposes RoPE (Rotary Position Embedding), which has the following characteristics.
RoPE uses a rotation matrix to encode the absolute position while at the same time creating a relative position dependency within the self-attention formulation, and the dependency between tokens decreases as the relative distance increases.
2. Background & Related Work
2.1. Absolute Position Embedding
A d-dimensional vector pi∈Rd, determined by the position i, is added directly to the token representation xi and then projected.
ft:t∈{q,k,v}(xi,i):=Wt:t∈{q,k,v}(xi+pi)
The following sinusoidal function is used for pi.
Replaces the key-side absolute PE pn with the relative vector p~m−n
Replaces the query-side pm with position-independent learnable vectors u, v
Uses different Wk, Wk for content-based keys and position-based keys
There are many other relative PE variants (T5, DeBERTa, etc.), but they all share the approach of changing how position information is added to each term of the expanded attention score.
3. Proposed Approach: RoPE
3.1. Formulation
The goal is to make the inner product of the query and key depend only on the relative position m−n.
⟨fq(xm,m),fk(xn,n)⟩=g(xm,xn,m−n)
In other words, the problem is to find an f such that absolute positions m and n are fed into fq and fk respectively, yet their inner product is determined only by the contents of the two tokens and their relative position.
3.2. RoPE - 2D case
First, consider the 2D case. If we treat every vector and matrix as complex numbers, the following solution satisfies the condition given in 3.1.
Multiplying by the complex number eimθ is equivalent to rotating by mθ in the 2D plane. Taking the conjugate and multiplying gives eimθ⋅e−inθ=ei(m−n)θ, so the absolute position disappears and only the relative position remains. Written in matrix form:
Here, the order of operations in f matters: the query/key projection is applied first, and the result is then multiplied by the rotation matrix.
3.3. RoPE - General Form
Since a solution was found for the 2D case in 3.2, the idea is generalized to d dimensions. In general d dimensions (d even), the d dimensions are split into d/2 two-dimensional subspaces, and a 2D rotation with a different angle θi is applied to each subspace. Since the inner product is linear, the per-subspace results can simply be summed. The rotation matrix generalized to d dimensions can be written as follows.
The formula above can be illustrated intuitively as follows.
As shown, in RoPE each subspace rotates at a different speed even within a single token vector. The leading dimensions, which rotate fastest with θ1=1, are sensitive to position differences between nearby tokens, while the trailing dimensions, which rotate very slowly with θd/2≈1/10000, represent relationships between tokens that are far apart.
The attention score with RoPE applied is as follows.
Since the rotation matrix satisfies (RΘ,md)TRΘ,nd=RΘ,n−md, the relative position appears explicitly.
The differences between RoPE and existing position embeddings can be summarized as follows.
Additive vs Multiplicative: existing PEs add position information, whereas RoPE multiplies by a rotation matrix.
It represents relative position without adding any new terms to the expanded attention score.
RΘd is an orthogonal matrix and does not change the norm of the vector, so the position encoding process does not destabilize training.
3.4. Properties
The main properties of RoPE can be summarized as the following two.
① Long-term decay
Setting θi=10000−2(i−1)/d makes the inner product decrease as the relative distance grows. This matches the intuition that tokens farther apart are less related.
② RoPE with linear attention
Linear attention computes similarity with non-negative feature maps ϕ(⋅), φ(⋅) instead of softmax. Since RoPE injects position information through rotation and preserves the norm of the hidden representation, it can be combined with linear attention simply by multiplying by the rotation after applying the feature map.
Efficient Computation RΘ,md is a sparse matrix that is mostly zeros, so a standard matrix multiplication would waste a lot of computation.. Instead, it is computed with two element-wise multiplications and an addition.
Since the cos and sin vectors can be precomputed for each position, the additional cost is O(d).
Long-term decay
When the upper bound of the attention score is plotted against relative distance, the relationship was found to be as follows.
As the graph shows, there is some oscillation, but overall the upper bound decreases as the relative distance increases. Note, however, that this means the upper bound of the inner product decreases, not that the actual attention score always decreases.
4. Experiments and Evaluation
4.1. Machine translation
This task evaluated sequence-to-sequence translation on WMT 2014 English-German. Only the position encoding of the self-attention layer in Transformer-base was replaced with RoPE, and the remaining hyperparameters were kept identical to the original model to isolate the effect of the position encoding.
4.2. Pre-training language modeling
To assess the quality of the contextual representation, the pre-training loss was compared. Similar to the model setup in 4.1, the original sinusoidal position encoding was replaced with RoPE and the losses were compared.
In both cases, the model with RoPE converges faster and reaches a lower loss. In particular, the result on the right experimentally demonstrates compatibility with linear attention.
4.3. Fine-tuning on GLUE
The pre-trained RoFormer was fine-tuned on several GLUE tasks to measure its performance.
RoFormer leads on MRPC, STS-B, and QQP, but BERT leads on SST-2, QNLI, and MNLI. In other words, it did not improve consistently across all tasks.
4.4. Evaluation on Chinese Data
To see whether the advantages of RoPE show up on long text, experiments were conducted on Chinese data. For the evaluation, RoFormer was built by replacing the absolute position embedding of WoBERT with RoPE. It was also compared with models that differ in tokenization unit and position embedding.
At the same length the differences are small, but RoFormer with the maximum length extended to 1024 achieved the best performance. From this result, RoPE can be seen as advantageous for handling long sequences.
5. Conclusion
RoFormer proposed RoPE, which multiplies the query and key by a rotation matrix proportional to their position. Although it encodes absolute positions, the inner product of the two vectors depends only on the relative position, and because it is a multiplicative approach that adds no new terms to the attention formula, it can be applied directly to linear attention as well. In addition, by setting the per-subspace rotation speed to θi=10000−2(i−1)/d, it obtains long-term decay with distance. In experiments, it showed advantages in pre-training convergence speed and on long-document tasks, but results on GLUE were mixed across tasks.
6. My Take
1. Despite its mathematical advantages, GLUE results are mixed
RoFormer appears mathematically more effective than existing positional embeddings, but it did not show a clear overall advantage in GLUE fine-tuning. In fact, it scored lower than BERT on MNLI (84.6 → 80.2), SST-2, and QNLI. Most GLUE tasks consist of short sentences, which seems to make it hard for the ability to handle long-range positional relationships to show. The tasks where it did well may be ones that are particularly sensitive to token encoding or positional embedding rather than to the model's general performance. In other words, it is possible that differences in positional embedding alone did not substantially change the model's performance itself. Conversely, given that the difference did show up on a long-input task like CAIL2019-SCM, it seems right to view RoPE's advantage as coming from long-context processing rather than short-sentence understanding.
Also, BERT's QQP score of 71.2 is the GLUE test set F1 from the original BERT paper, and the paper does not make clear whether it was compared with RoFormer's 86.4 under the same conditions (split, metric). I think it is hard to read the 15-point gap directly as the effect of RoPE.
References
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., & Liu, Y. (2024). RoFormer: Enhanced Transformer with Rotary Position Embedding. Neurocomputing. arXiv:2104.09864
Vaswani, A. et al. (2017). Attention Is All You Need.arXiv:1706.03762
Shaw, P., Uszkoreit, J., & Vaswani, A. (2018). Self-Attention with Relative Position Representations.arXiv:1803.02155
Dai, Z. et al. (2019). Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.arXiv:1901.02860
Devlin, J. et al. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.arXiv:1810.04805
This post was translated from the original Korean version with the help of AI, so some errors may remain.