Back to posts

Rereading the Transformer Upgrade Path

Part 2

Rereading the Transformer Upgrade Path (2): How RoPE Writes Relative Position into Attention

Starting from a single identity for two-dimensional rotations, this article explains how RoPE uses the same sin/cos signals to rotate Queries and Keys so that position enters the attention score only through relative displacement.

Sources. This article primarily draws on Jianlin Su’s “Transformer Upgrade Path 2: Rotary Position Embedding That Combines the Best of Both Worlds”, together with the RoFormer, Linear Transformer, and Performer papers.

Thirty-second summary. In the previous article, “Rereading the Transformer Upgrade Path (1): Tracing Sinusoidal Positional Encoding to Its Source”, every sin/cos pair acts like a “position clock” that represents a position as coordinates on the unit circle. RoPE uses the same position angles but no longer adds a position vector to the token. Instead, it places the sin/cos values inside a rotation operator and directly rotates the content-derived Query and Key. Position mm rotates by mθm\theta, and position nn by nθn\theta; when the rotated vectors are compared by an inner product, the shared absolute rotation cancels, leaving only n−mn-m. RoPE therefore does not multiply an already constructed attention-logit matrix by distance weights. It changes how Queries and Keys are compared before those logits exist.

First establish the global picture: RoPE performs only four steps inside one attention computation.

  1. Produce position-free vectors qm=WQxmq_m=W_Qx_m and kn=WKxnk_n=W_Kx_n from the token representations;
  2. For the ii-th feature pair, θi\theta_i is the rotation increment in radians for each one-token step, so the accumulated angle at position mm is mθim\theta_i; for positions m,nm,n, obtain the corresponding cos⁡(mθi),sin⁡(mθi)\cos(m\theta_i),\sin(m\theta_i) and cos⁡(nθi),sin⁡(nθi)\cos(n\theta_i),\sin(n\theta_i);
  3. Rotate the Query and Key within the feature dimensions of each attention head: q~m=Rmqm, k~n=Rnkn\tilde q_m=R_mq_m,\ \tilde k_n=R_nk_n;
  4. Only then compute sm,n=q~m⊤k~ns_{m,n}=\tilde q_m^\top\tilde k_n and assemble all sm,ns_{m,n} into the N×NN\times N logit matrix.

RmR_m is a dh×dhd_h\times d_h block rotation acting on the feature dimensions of one token. It is not the N×NN\times N attention matrix. At no point does RoPE take an elementwise product between a logit matrix and a rotation matrix.

We will repeatedly use the following notation:

SymbolMeaning
m,nm,nAbsolute positions of the Query and Key
Δ=m−n\Delta=m-nDisplacement of the Query relative to the Key
xmx_mToken representation at position mm
qm,kn,vnq_m,k_n,v_nQuery, Key, and Value before positional information is introduced
dhd_hQuery/Key dimensionality of one attention head, assumed even in this article
R(α)R(\alpha)Matrix that rotates a vector counterclockwise by α\alpha in the two-dimensional plane
RmR_mBlock matrix that rotates every dimension pair in a complete attention head by mθim\theta_i
θi\theta_iRadians rotated by the ii-th dimension pair for each token step

The core RoPE path ends with the quick reference in Section 8. Section 9 discusses ways to combine RoPE with linear attention; it is not required for understanding RoPE itself and can be skipped entirely.

1. What Gap in Sinusoidal PE Is RoPE Trying to Fill?

First, put positional encoding aside. Standard self-attention produces three groups of vectors from token representations:

qm=WQxm,kn=WKxn,vn=WVxn.q_m=W_Qx_m, \qquad k_n=W_Kx_n, \qquad v_n=W_Vx_n.

A Query can be understood as “what position mm is looking for,” a Key as “which features position nn can be found by,” and a Value as the information actually retrieved after that position has been selected. Their inner product

sm,n=qm⊤kns_{m,n}=q_m^\top k_n

determines how well the Query and Key match.

The problem is that if neither qmq_m nor knk_n carries position, then this match compares only content and does not know how many steps separate the two tokens.

The previous article introduced Sinusoidal positional encoding, which first constructs a position vector pmp_m and then adds it to the token representation. To describe additive PE and RoPE in the same language, write “adding a position vector” as a position-dependent transformation:

Fadd(xm,m)=xm+pm.F_{\mathrm{add}}(x_m,m)=x_m+p_m.

The second argument mm tells the function which pmp_m to use. Projecting the transformed tokens into Queries and Keys gives

q^m=WQFadd(xm,m)=qm+WQpm,k^n=WKFadd(xn,n)=kn+WKpn.\begin{aligned} \hat q_m &=W_QF_{\mathrm{add}}(x_m,m) =q_m+W_Qp_m,\\ \hat k_n &=W_KF_{\mathrm{add}}(x_n,n) =k_n+W_Kp_n. \end{aligned}

Now expand their inner product without skipping any terms:

q^m⊤k^n=qm⊤kn+qm⊤WKpn+(WQpm)⊤kn+(WQpm)⊤(WKpn).\begin{aligned} \hat q_m^\top\hat k_n ={}&q_m^\top k_n +q_m^\top W_Kp_n\\ &+(W_Qp_m)^\top k_n +(W_Qp_m)^\top(W_Kp_n). \end{aligned}

The four terms are content—content, content—position, position—content, and position—position. The previous article showed that the Sinusoidal vectors themselves have a clean relative-phase structure—for example, pm⊤pnp_m^\top p_n depends only on m−nm-n. But the complete score above also passes through WQ,WKW_Q,W_K and mixes in the other three terms, so it is not forced into a form in which position appears only through m−nm-n.

This naturally suggests the next step. If position ultimately exists to affect the Query—Key comparison, can we define a more suitable position transformation directly on the Query and Key? Denote these two transformations by

qmpos=FQ(qm,m),knpos=FK(kn,n).q_m^{\mathrm{pos}}=F_Q(q_m,m), \qquad k_n^{\mathrm{pos}}=F_K(k_n,n).

RoPE’s core motivation

We want to find a pair of position transformations, FQF_Q and FKF_K. They first write the absolute positions m,nm,n into the Query and Key; after the transformed vectors are compared by an inner product, the position variables should appear only through the relative displacement m−nm-n:

FQ(qm,m)⊤FK(kn,n)=g(qm,kn,m−n).\boxed{ F_Q(q_m,m)^\top F_K(k_n,n) =g(q_m,k_n,m-n) }.

The left-hand side contains the transformations we still need to design; the right-hand side states the form we want their inner product to take. In other words, we work backward from the relative-position target instead of choosing FQ,FKF_Q,F_K first and merely hoping that the desired structure emerges.

This does not require the score to depend on distance alone: the right-hand side still retains the Query and Key content qm,knq_m,k_n. Only the position variables are restricted, and they may appear only through m−nm-n. RoPE’s concrete choice is to let both FQF_Q and FKF_K use rotations determined by position.

This objective has an intuitive test. If both tokens are shifted backward by cc positions, their relative distance does not change:

(m+c)−(n+c)=m−n.(m+c)-(n+c)=m-n.

We therefore also want the content match to remain unchanged after this shared shift. All that remains is to construct a simple transformation with this property.

2. Properties of Two-Dimensional Rotation Matrices

2.1 What Exactly Does a Rotation Matrix Do?

In the two-dimensional plane, a vector

u=[u1u2]u= \begin{bmatrix} u_1\\ u_2 \end{bmatrix}

becomes

R(α)u=[cos⁡α−sin⁡αsin⁡αcos⁡α][u1u2]=[u1cos⁡α−u2sin⁡αu1sin⁡α+u2cos⁡α]R(\alpha)u = \begin{bmatrix} \cos\alpha&-\sin\alpha\\ \sin\alpha&\cos\alpha \end{bmatrix} \begin{bmatrix} u_1\\ u_2 \end{bmatrix} = \begin{bmatrix} u_1\cos\alpha-u_2\sin\alpha\\ u_1\sin\alpha+u_2\cos\alpha \end{bmatrix}

after a counterclockwise rotation by α\alpha.

For example, take u=[1,0]⊤u=[1,0]^\top, which initially points along the positive horizontal axis. After a rotation by π/2\pi/2,

R(π/2)[10]=[0−110][10]=[01],R(\pi/2) \begin{bmatrix} 1\\0 \end{bmatrix} = \begin{bmatrix} 0&-1\\1&0 \end{bmatrix} \begin{bmatrix} 1\\0 \end{bmatrix} = \begin{bmatrix} 0\\1 \end{bmatrix},

so it now points along the positive vertical axis. Rotation changes the direction of a vector, not its length.

2.2 Why Does the Transpose Equal the Reverse Rotation?

Transposing the rotation matrix swaps its rows and columns:

R(α)⊤=[cos⁡αsin⁡α−sin⁡αcos⁡α].R(\alpha)^\top = \begin{bmatrix} \cos\alpha&\sin\alpha\\ -\sin\alpha&\cos\alpha \end{bmatrix}.

On the other hand,

cos⁡(−α)=cos⁡α,sin⁡(−α)=−sin⁡α,\cos(-\alpha)=\cos\alpha, \qquad \sin(-\alpha)=-\sin\alpha,

so

R(−α)=[cos⁡αsin⁡α−sin⁡αcos⁡α].R(-\alpha) = \begin{bmatrix} \cos\alpha&\sin\alpha\\ -\sin\alpha&\cos\alpha \end{bmatrix}.

The two matrices are identical entry by entry. Therefore,

R(α)⊤=R(−α).\boxed{R(\alpha)^\top=R(-\alpha)}.

The geometry is just as simple: after rotating a hand by α\alpha, restoring it to its original direction requires a rotation by −α-\alpha.

2.3 Why Can Two Consecutive Rotations Add Their Angles?

First rotating by β\beta and then by α\alpha corresponds to the matrix product

R(α)R(β).R(\alpha)R(\beta).

To expose every term, temporarily abbreviate cos⁡α,sin⁡α\cos\alpha,\sin\alpha as cα,sαc_\alpha,s_\alpha, and cos⁡β,sin⁡β\cos\beta,\sin\beta as cβ,sβc_\beta,s_\beta:

R(α)R(β)=[cα−sαsαcα][cβ−sβsβcβ]=[cαcβ−sαsβ−(cαsβ+sαcβ)sαcβ+cαsβcαcβ−sαsβ].\begin{aligned} R(\alpha)R(\beta) &= \begin{bmatrix} c_\alpha&-s_\alpha\\ s_\alpha&c_\alpha \end{bmatrix} \begin{bmatrix} c_\beta&-s_\beta\\ s_\beta&c_\beta \end{bmatrix}\\ &= \begin{bmatrix} c_\alpha c_\beta-s_\alpha s_\beta &-(c_\alpha s_\beta+s_\alpha c_\beta)\\ s_\alpha c_\beta+c_\alpha s_\beta &c_\alpha c_\beta-s_\alpha s_\beta \end{bmatrix}. \end{aligned}

Now apply the angle-sum identities for sine and cosine:

cos⁡(α+β)=cαcβ−sαsβ,sin⁡(α+β)=sαcβ+cαsβ.\begin{aligned} \cos(\alpha+\beta) &=c_\alpha c_\beta-s_\alpha s_\beta,\\ \sin(\alpha+\beta) &=s_\alpha c_\beta+c_\alpha s_\beta. \end{aligned}

The product above becomes

R(α)R(β)=[cos⁡(α+β)−sin⁡(α+β)sin⁡(α+β)cos⁡(α+β)]=R(α+β).R(\alpha)R(\beta) = \begin{bmatrix} \cos(\alpha+\beta)&-\sin(\alpha+\beta)\\ \sin(\alpha+\beta)&\cos(\alpha+\beta) \end{bmatrix} =R(\alpha+\beta).

Hence,

R(α)R(β)=R(α+β).\boxed{R(\alpha)R(\beta)=R(\alpha+\beta)}.

If there is only one sentence to remember, it is this: rotations multiply; their angles add.

3. How Does RoPE Make the Inner Product Depend Only on Relative Position?

Now take one two-dimensional Query and Key:

qm=[q1q2],kn=[k1k2].q_m= \begin{bmatrix} q_1\\q_2 \end{bmatrix}, \qquad k_n= \begin{bmatrix} k_1\\k_2 \end{bmatrix}.

Choose a fixed per-step rotation angle θ\theta. This θ\theta is not a new signal that appears from nowhere. It is the same per-step phase increment used by Sinusoidal PE. For the ii-th dimension pair, define

sm,i=sin⁡(mθi),cm,i=cos⁡(mθi).s_{m,i}=\sin(m\theta_i), \qquad c_{m,i}=\cos(m\theta_i).

Sinusoidal PE treats these two numbers as a pair of coordinates in a position vector. RoPE places the same two numbers inside a rotation operator:

pm(i)=[sm,icm,i]⏟Sinusoidal PE: position coordinates,R(mθi)=[cm,i−sm,ism,icm,i]⏟RoPE: position operator.\underbrace{ p_m^{(i)}= \begin{bmatrix} s_{m,i}\\c_{m,i} \end{bmatrix} }_{\text{Sinusoidal PE: position coordinates}}, \qquad \underbrace{ R(m\theta_i)= \begin{bmatrix} c_{m,i}&-s_{m,i}\\ s_{m,i}&c_{m,i} \end{bmatrix} }_{\text{RoPE: position operator}}.

The precise relationship is therefore not “first construct a Sinusoidal position vector and then convert it into RoPE.” Instead, both methods generate sin/cos values from the same position angles; Sinusoidal PE uses them as coordinates to add, whereas RoPE uses them as rotation coefficients to multiply. Implementations normally cache these sin/cos values directly, without explicitly constructing pmp_m or the full matrix RmR_m.

Continue with one two-dimensional pair. At position mm, RoPE rotates the Query by mθm\theta; at position nn, it rotates the Key by nθn\theta:

q~m=R(mθ)qm,k~n=R(nθ)kn.\tilde q_m=R(m\theta)q_m, \qquad \tilde k_n=R(n\theta)k_n.

Take the inner product of the rotated vectors:

q~m⊤k~n=(R(mθ)qm)⊤(R(nθ)kn)=qm⊤R(mθ)⊤R(nθ)kn=qm⊤R(−mθ)R(nθ)kn=qm⊤R((n−m)θ)kn.\begin{aligned} \tilde q_m^\top\tilde k_n &=(R(m\theta)q_m)^\top(R(n\theta)k_n)\\ &=q_m^\top R(m\theta)^\top R(n\theta)k_n\\ &=q_m^\top R(-m\theta)R(n\theta)k_n\\ &=q_m^\top R((n-m)\theta)k_n. \end{aligned}

The three steps use, respectively:

  1. (AB)⊤=B⊤A⊤(AB)^\top=B^\top A^\top;
  2. R(α)⊤=R(−α)R(\alpha)^\top=R(-\alpha);
  3. R(α)R(β)=R(α+β)R(\alpha)R(\beta)=R(\alpha+\beta).

Only n−mn-m remains as a position variable in the final expression:

q~m⊤k~n=qm⊤R((n−m)θ)kn.\boxed{ \tilde q_m^\top\tilde k_n =q_m^\top R((n-m)\theta)k_n }.

Although the Query and Key are each rotated according to their absolute positions m,nm,n, when they meet in the inner product they see only their relative position.

If both positions are shifted by cc, then similarly

R((m+c)θ)⊤R((n+c)θ)=R(−(m+c)θ)R((n+c)θ)=R((n−m)θ).\begin{aligned} R((m+c)\theta)^\top R((n+c)\theta) &=R(-(m+c)\theta)R((n+c)\theta)\\ &=R((n-m)\theta). \end{aligned}

The shared cc disappears. This is the matrix form of “a global shift does not change a relative relationship.”

Query and Key hands rotate together before and after a shared position shift, while their relative angle stays unchanged
Figure 1 | A global shift changes absolute angles but not the relative rotation. To isolate the rotation caused by position, both the Query and Key start from the same reference direction, with θ = 30° used for illustration. The two position pairs are (1, 3) and (4, 6); the latter adds 3 to both positions, but the position difference remains 2, so the relative angle remains 2θ. Author-created mechanism diagram.

4. RoPE Does Not Reweight an Existing Attention-Logit Matrix

The previous section proved that position appears only through m−nm-n, but we must first rule out a natural misconception. RoPE does not compute the original logit matrix and then multiply every sm,ns_{m,n} by a distance-dependent weight.

4.1 The Two Operations Use a Different Order

Elementwise distance weighting would have the form

S=QK⊤,S′=S⊙W,S=QK^\top, \qquad S'=S\odot W,

where S,W∈RN×NS,W\in\mathbb{R}^{N\times N}. Each original logit can only be enlarged, reduced, or sign-flipped by the scalar Wm,nW_{m,n}.

RoPE uses a different order. It first rotates the Query and Key inside the dhd_h-dimensional feature space of each token and only then constructs the logit matrix:

q~m=Rmqm,k~n=Rnkn,Sm,n′=q~m⊤k~n.\tilde q_m=R_mq_m, \qquad \tilde k_n=R_nk_n, \qquad S'_{m,n}=\tilde q_m^\top\tilde k_n.

Rm,Rn∈Rdh×dhR_m,R_n\in\mathbb{R}^{d_h\times d_h} mix each pair of feature coordinates. They are not the same object as the N×NN\times N matrix SS, and no elementwise product is taken between them. The simplest distinction is that scalar reweighting can never turn an original zero logit into a nonzero value, whereas relative rotation can change the coordinate alignment and do exactly that.

4.2 How Does Relative Rotation Change Content Matching?

We can now answer the more intuitive question: for the same Query and Key, why does changing the relative position change their matching score?

Let

Δ=m−n.\Delta=m-n.

Because n−m=−Δn-m=-\Delta, the two-dimensional RoPE score can be written as

qm⊤R(−Δθ)kn.q_m^\top R(-\Delta\theta)k_n.

To keep every step visible, first write out the two factors on the right:

kn=[k1k2],k_n= \begin{bmatrix} k_1\\k_2 \end{bmatrix},

and

R(−Δθ)=[cos⁡(−Δθ)−sin⁡(−Δθ)sin⁡(−Δθ)cos⁡(−Δθ)]=[cos⁡(Δθ)sin⁡(Δθ)−sin⁡(Δθ)cos⁡(Δθ)].\begin{aligned} R(-\Delta\theta) &= \begin{bmatrix} \cos(-\Delta\theta)&-\sin(-\Delta\theta)\\ \sin(-\Delta\theta)&\cos(-\Delta\theta) \end{bmatrix}\\ &= \begin{bmatrix} \cos(\Delta\theta)&\sin(\Delta\theta)\\ -\sin(\Delta\theta)&\cos(\Delta\theta) \end{bmatrix}. \end{aligned}

The second equality uses the fact that cosine is even and sine is odd: cos⁡(−x)=cos⁡x\cos(-x)=\cos x and sin⁡(−x)=−sin⁡x\sin(-x)=-\sin x.

First multiply the matrix and the vector:

R(−Δθ)kn=[cos⁡(Δθ)sin⁡(Δθ)−sin⁡(Δθ)cos⁡(Δθ)][k1k2]=[k1cos⁡(Δθ)+k2sin⁡(Δθ)−k1sin⁡(Δθ)+k2cos⁡(Δθ)].\begin{aligned} R(-\Delta\theta)k_n &= \begin{bmatrix} \cos(\Delta\theta)&\sin(\Delta\theta)\\ -\sin(\Delta\theta)&\cos(\Delta\theta) \end{bmatrix} \begin{bmatrix} k_1\\k_2 \end{bmatrix}\\ &= \begin{bmatrix} k_1\cos(\Delta\theta)+k_2\sin(\Delta\theta)\\ -k_1\sin(\Delta\theta)+k_2\cos(\Delta\theta) \end{bmatrix}. \end{aligned}

Now write the transposed Query explicitly as a row vector:

qm⊤=[q1q2].q_m^\top= \begin{bmatrix} q_1&q_2 \end{bmatrix}.

The inner product can then be evaluated one line at a time:

qm⊤R(−Δθ)kn=[q1q2][k1cos⁡(Δθ)+k2sin⁡(Δθ)−k1sin⁡(Δθ)+k2cos⁡(Δθ)]=q1[k1cos⁡(Δθ)+k2sin⁡(Δθ)]+q2[−k1sin⁡(Δθ)+k2cos⁡(Δθ)]=q1k1cos⁡(Δθ)+q1k2sin⁡(Δθ)−q2k1sin⁡(Δθ)+q2k2cos⁡(Δθ)=(q1k1+q2k2)cos⁡(Δθ)+(q1k2−q2k1)sin⁡(Δθ).\begin{aligned} q_m^\top R(-\Delta\theta)k_n &= \begin{bmatrix} q_1&q_2 \end{bmatrix} \begin{bmatrix} k_1\cos(\Delta\theta)+k_2\sin(\Delta\theta)\\ -k_1\sin(\Delta\theta)+k_2\cos(\Delta\theta) \end{bmatrix}\\ &=q_1[k_1\cos(\Delta\theta)+k_2\sin(\Delta\theta)]\\ &\quad+q_2[-k_1\sin(\Delta\theta)+k_2\cos(\Delta\theta)]\\ &=q_1k_1\cos(\Delta\theta) +q_1k_2\sin(\Delta\theta)\\ &\quad-q_2k_1\sin(\Delta\theta) +q_2k_2\cos(\Delta\theta)\\ &=(q_1k_1+q_2k_2)\cos(\Delta\theta)\\ &\quad+(q_1k_2-q_2k_1)\sin(\Delta\theta). \end{aligned}

The RoPE score therefore contains two parts:

(q1k1+q2k2)⏟original same-direction matchcos⁡(Δθ)+(q1k2−q2k1)⏟cross-coordinate matchsin⁡(Δθ).\boxed{ \underbrace{(q_1k_1+q_2k_2)}_{\text{original same-direction match}} \cos(\Delta\theta) + \underbrace{(q_1k_2-q_2k_1)}_{\text{cross-coordinate match}} \sin(\Delta\theta) }.

The easiest misconception here is to imagine RoPE as “multiplying the original dot product by a distance cosine.” Only the first term has that form. The second term also cross-compares the first coordinate of the Query with the second coordinate of the Key, and the second coordinate of the Query with the first coordinate of the Key.

What does this expression actually say? Two controlled experiments make it concrete. To keep the rotations easy to calculate mentally, temporarily take θ=π/2\theta=\pi/2, so a one-position difference means a 90∘90^\circ rotation. This is a toy setting chosen to magnify the mechanism, not the only frequency used by a real model.

ExperimentQuery qqOriginal Key kkΔ\DeltaRotated Key R(−Δθ)kR(-\Delta\theta)kRoPE score
A1[1,0][1,0][1,0][1,0]00[1,0][1,0]11
A2[1,0][1,0][1,0][1,0]11[0,−1][0,-1]00
B1[1,0][1,0][0,1][0,1]00[0,1][0,1]00
B2[1,0][1,0][0,1][0,1]11[1,0][1,0]11

Start with experiment A. The content vectors are identical in both rows; only the relative position changes:

Δ=0:[1,0] [1,0]⊤=1,Δ=1:[1,0] [0,−1]⊤=0.\begin{aligned} \Delta=0:&\quad [1,0]\,[1,0]^\top=1,\\ \Delta=1:&\quad [1,0]\,[0,-1]^\top=0. \end{aligned}

Thus even when the content is unchanged, relative position rotates the Key first and then changes how well it matches the Query.

Now consider experiment B. The Query and Key are initially orthogonal, so their score at the same position is 00. When Δ=1\Delta=1, R(−π/2)R(-\pi/2) turns the Key from [0,1][0,1] into [1,0][1,0], exactly aligning it with the Query, and the score becomes 11.

These examples make only one point: RoPE lets relative position determine the coordinate directions in which two pieces of content are compared. It does not add a fixed reward or penalty for distance, nor does it guarantee that the score decreases with distance.

A real attention head uses many different θi\theta_i values at once, while q,kq,k are content vectors learned by the model. The position difference determines how each coordinate pair rotates; the content determines whether those rotated coordinates match. The results from all dimensions are then added together.

5. From One Dimension Pair to a Complete Attention Head

In a real model, one attention head usually has tens or hundreds of dimensions. RoPE groups these dimensions into pairs:

(q0,q1), (q2,q3), …, (qdh−2,qdh−1).(q_0,q_1), \ (q_2,q_3), \ \ldots, \ (q_{d_h-2},q_{d_h-1}).

The ii-th pair has its own per-step rotation angle θi\theta_i. The complete rotation can be written as a block-diagonal matrix:

Rm=diag⁡(R(mθ0),R(mθ1),…,R(mθdh/2−1)).R_m = \operatorname{diag} \left( R(m\theta_0), R(m\theta_1), \ldots, R(m\theta_{d_h/2-1}) \right).

A common frequency schedule is inherited from Sinusoidal PE:

θi=b−2i/dh,\theta_i=b^{-2i/d_h},

with the classic base b=10000b=10000. Larger θi\theta_i values rotate quickly and are more sensitive to local position differences; smaller values rotate slowly and vary more gradually over longer distances.

The complete score is the sum of the scores from all two-dimensional blocks:

q~m⊤k~n=∑i=0dh/2−1[Aicos⁡(Δθi)+Bisin⁡(Δθi)],\tilde q_m^\top\tilde k_n = \sum_{i=0}^{d_h/2-1} \left[ A_i\cos(\Delta\theta_i) +B_i\sin(\Delta\theta_i) \right],

where Ai,BiA_i,B_i are determined by the content in the ii-th Query/Key pair.

Rotation also preserves vector length. Because

Rm⊤Rm=I,R_m^\top R_m=I,

we have

∥Rmqm∥2=qm⊤Rm⊤Rmqm=qm⊤qm=∥qm∥2.\lVert R_mq_m\rVert^2 =q_m^\top R_m^\top R_mq_m =q_m^\top q_m =\lVert q_m\rVert^2.

Position changes direction, not the magnitude of the Query or Key.

The multi-frequency score is still a content-dependent finite trigonometric sum. It may oscillate with distance; it is not guaranteed to decrease strictly, nor are distant tokens guaranteed to receive less attention. A more precise statement is that geometric frequencies supply multiscale phase structure and, when many frequencies are mixed, tend to produce oscillatory decorrelation. This is not a hard-coded distance penalty.

6. Where Does RoPE Enter Standard Attention?

For one attention head, the computation can be written in five steps.

Step 1: Produce Q, K, and V from the Input

qm=WQxm,kn=WKxn,vn=WVxn.q_m=W_Qx_m, \qquad k_n=W_Kx_n, \qquad v_n=W_Vx_n.

Step 2: Rotate Only Q and K

Here RmR_m is not a new, undefined matrix. It is the complete-head rotation from the previous section. To save the reader from searching backward, write it once more:

Rm=diag⁡ ⁣(R(mθ0),…,R(mθdh/2−1)),Rn=diag⁡ ⁣(R(nθ0),…,R(nθdh/2−1)).\begin{aligned} R_m &=\operatorname{diag}\!\left( R(m\theta_0),\ldots,R(m\theta_{d_h/2-1}) \right),\\ R_n &=\operatorname{diag}\!\left( R(n\theta_0),\ldots,R(n\theta_{d_h/2-1}) \right). \end{aligned}

RmR_m rotates every Query dimension pair according to position mm, and RnR_n rotates every Key dimension pair according to position nn. Therefore,

qmrope=Rmqm,knrope=Rnkn.q_m^{\mathrm{rope}}=R_mq_m, \qquad k_n^{\mathrm{rope}}=R_nk_n.

Step 3: Compute Every Query—Key Score

sm,n=(qmrope)⊤knropedh.s_{m,n} =\frac{(q_m^{\mathrm{rope}})^\top k_n^{\mathrm{rope}}}{\sqrt{d_h}}.

Step 4: Add the Mask, Then Apply Softmax

Here we use an additive mask. Define Mm,nM_{m,n} by

Mm,n={0,position m may read position n,−∞,position m may not read position n.M_{m,n}= \begin{cases} 0, & \text{position }m\text{ may read position }n,\\ -\infty, & \text{position }m\text{ may not read position }n. \end{cases}

It is added to the score before the exponential:

am,n=exp⁡(sm,n+Mm,n)∑jexp⁡(sm,j+Mm,j).a_{m,n} = \frac{\exp(s_{m,n}+M_{m,n})} {\sum_j\exp(s_{m,j}+M_{m,j})}.

The plus sign is therefore intentional. When reading is allowed, exp⁡(sm,n+0)=exp⁡(sm,n)\exp(s_{m,n}+0)=\exp(s_{m,n}); when it is forbidden, exp⁡(sm,n−∞)=0\exp(s_{m,n}-\infty)=0. Implementations usually replace −∞-\infty with a very large negative number representable by the chosen data type.

If an implementation uses a Boolean 0/10/1 mask, that is a different representation. It is usually converted into the 0/−∞0/-\infty additive mask used here, or applied through an equivalent multiplication after exponentiation. In this article, MM is defined from the start in the pre-softmax logit space.

Step 5: Use the Weights to Aggregate the Unrotated Values

om=∑nam,nvn.o_m=\sum_n a_{m,n}v_n.

Values usually do not need to be rotated. RoPE’s purpose is to make the pairing score—“how much information should position mm take from position nn?”—depend on relative position. Once am,na_{m,n} contains that information, it can directly select and combine the Values.

In code-like form:

q, k, v = project(x)
q_rope = q * cos(position) + rotate_half(q) * sin(position)
k_rope = k * cos(position) + rotate_half(k) * sin(position)

score   = q_rope @ k_rope.T / sqrt(head_dim)
weight  = softmax(score + mask)
output  = weight @ v

For every two-dimensional block [a,b][a,b],

rotate_half⁡([a,b])=[−b,a],\operatorname{rotate\_half}([a,b])=[-b,a],

which is exactly the part of the rotation matrix multiplied by sin⁡\sin.

One boundary still needs to be explicit: standard attention must compute every m,nm,n combination and form an N×NN\times N score matrix. RoPE adds positional structure; it does not automatically turn the O(N2)O(N^2) complexity of standard attention into linear complexity.

7. Core Conclusions and Boundaries

At this point, the core path needed to understand RoPE itself is complete. The preceding derivations establish or directly demonstrate that:

  1. RoPE and Sinusoidal PE use the same kind of position angles and sin/cos signals, but assign them different roles: RoPE places them inside a rotation operator, whereas Sinusoidal PE places them inside a position vector;
  2. Two-dimensional RoPE is a length-preserving rotation;
  3. Rm⊤Rn=Rn−mR_m^\top R_n=R_{n-m}, so the position variables enter the Query—Key score only through relative displacement;
  4. Shifting the Query and Key together does not change their relative rotation;
  5. RoPE rotates Q/K before the logit matrix is formed; it does not apply elementwise distance weights to existing logits;
  6. The RoPE score is not merely the original dot product multiplied by a distance cosine, because it also contains cross-coordinate content matching.

These derivations do not prove that:

  1. RoPE is the unique solution to the relative-position objective;
  2. the RoPE score decreases strictly and monotonically with distance;
  3. a model using RoPE must prefer nearby tokens;
  4. being able to compute rotation angles beyond the training length guarantees reliable length extrapolation;
  5. RoPE reduces the O(N2)O(N^2) complexity of standard attention.

8. Quick Reference

QuestionShortest answer
Where do RoPE’s sin/cos values come from?It generates sin/cos from the same kind of position angles and frequencies as Sinusoidal PE, but uses them directly as rotation coefficients.
What does RoPE rotate?The projected Query and Key in each attention head.
Where does the rotation matrix act?In the dhd_h-dimensional feature space of one token, not on the N×NN\times N attention matrix.
Why pair every two dimensions?Two dimensions are exactly enough to represent a length-preserving planar rotation.
Why does relative position appear?R(mθ)⊤R(nθ)=R((n−m)θ)R(m\theta)^\top R(n\theta)=R((n-m)\theta).
Does RoPE add a position vector to content?No. It directly rotates the content-derived Q/K vectors.
Does RoPE multiply the logit matrix elementwise?No. It rotates Q/K first and constructs the logits from the rotated vectors.
Does RoPE merely multiply the dot product by a cosine?No. It also introduces cross-coordinate matching controlled by the sine term.
Why is the Value usually not rotated?Relative position already enters the attention weights that select the Values, so the Values can be aggregated directly.
Does standard attention become linear as a result?No. It must still compute N2N^2 Query—Key scores.

The single most important sentence from the core RoPE path is:

RoPE uses sin/cos signals from the same source as Sinusoidal PE to express position as rotations acting on Queries and Keys; when the two are compared, their absolute rotations cancel and only relative displacement remains.

9. Optional Reading: RoPE and Linear Attention

The remainder addresses a separate question: once RoPE is understood, can it be combined with the multiplication reordering used by linear attention? This is not part of RoPE’s definition and does not affect the preceding conclusions about standard attention. Readers interested only in RoPE itself can skip directly to the references.

This section introduces two additional symbols: ϕ,φ\phi,\varphi denote the feature maps applied to Queries and Keys in linear attention.

9.1 Why Is Linear Attention Called “Linear”?

Here “linear” means that the amount of computation grows linearly with sequence length NN, not that the entire module contains no nonlinear functions.

To make the contrast with standard attention explicit, temporarily omit RoPE and the mask. Standard attention is

omstd=∑n=1Nexp⁡ ⁣(qm⊤kn/dh)vn∑n=1Nexp⁡ ⁣(qm⊤kn/dh).o_m^{\mathrm{std}} = \frac{ \sum_{n=1}^N \exp\!\left(q_m^\top k_n/\sqrt{d_h}\right)v_n }{ \sum_{n=1}^N \exp\!\left(q_m^\top k_n/\sqrt{d_h}\right) }.

Every coefficient exp⁡(qm⊤kn/dh)\exp(q_m^\top k_n/\sqrt{d_h}) is jointly produced by one specific Query—Key pair. A length-NN sequence has N2N^2 such pairs, so all Keys and Values cannot first be compressed into a finite Query-independent summary.

Linear attention changes this crucial pairwise coefficient. It chooses or approximates a similarity that separates into two feature vectors:

sim⁡(qm,kn)=ϕ(qm)⊤φ(kn).\operatorname{sim}(q_m,k_n) =\phi(q_m)^\top\varphi(k_n).

Let

um=ϕ(qm),zn=φ(kn).u_m=\phi(q_m), \qquad z_n=\varphi(k_n).

Its normalized output is then

omlin=∑n=1N(um⊤zn)vn∑n=1Num⊤zn.o_m^{\mathrm{lin}} = \frac{ \sum_{n=1}^N(u_m^\top z_n)v_n }{ \sum_{n=1}^N u_m^\top z_n }.

Place the two forms side by side and the difference is concentrated in the middle column:

Standard attentionLinear attention
Query—Key coefficientexp⁡(qm⊤kn/dh)\exp(q_m^\top k_n/\sqrt{d_h})um⊤znu_m^\top z_n
Form every pair first?Yes, N2N^2 pairs in totalNot necessary; Keys and Values can be aggregated first
NormalizationSoftmax for each QueryDivide by ∑num⊤zn\sum_nu_m^\top z_n
Computation in sequence lengthO(N2)O(N^2)O(N)O(N) when the feature dimension is fixed

The key to linear attention is therefore not merely that “softmax is absent from the formula.” It is that umu_m and znz_n can be computed separately, after which the associativity of matrix multiplication changes the order of summation.

Because umu_m does not depend on the summation index nn, it can be moved outside the sum:

∑n=1N(um⊤zn)vn=um⊤(∑n=1Nznvn⊤),∑n=1Num⊤zn=um⊤(∑n=1Nzn).\begin{aligned} \sum_{n=1}^N(u_m^\top z_n)v_n &=u_m^\top\left(\sum_{n=1}^Nz_nv_n^\top\right),\\ \sum_{n=1}^Nu_m^\top z_n &=u_m^\top\left(\sum_{n=1}^Nz_n\right). \end{aligned}

We can therefore compute two summary quantities for the entire sequence first:

SV=∑n=1Nznvn⊤,S1=∑n=1Nzn,S_V=\sum_{n=1}^Nz_nv_n^\top, \qquad S_1=\sum_{n=1}^Nz_n,

and then let every Query read them:

om=um⊤SVum⊤S1.\boxed{ o_m=\frac{u_m^\top S_V}{u_m^\top S_1} }.

No explicit N×NN\times N attention matrix is constructed. When the feature dimension is fixed, the computation is linear in sequence length NN.

In the causal setting, replace the full-sequence sums with prefix states:

SV,m=SV,m−1+zmvm⊤,S1,m=S1,m−1+zm.S_{V,m}=S_{V,m-1}+z_mv_m^\top, \qquad S_{1,m}=S_{1,m-1}+z_m.

Position mm can read only the state accumulated up to that position, so it cannot see future tokens.

9.2 How Does RoPE Enter Linear Attention?

9.2.1 Why Does RoPE Preserve the Multiplication Reordering?

Assume that the feature dimensions can also be paired. Rotate the Query and Key features separately:

um′=Rmum,zn′=Rnzn.u_m'=R_mu_m, \qquad z_n'=R_nz_n.

Their similarity still depends only on the relative rotation:

(um′)⊤zn′=um⊤Rn−mzn.(u_m')^\top z_n' =u_m^\top R_{n-m}z_n.

At the same time, the numerator can still be reordered:

∑n[(um′)⊤zn′]vn=(um′)⊤(∑nzn′vn⊤).\sum_n[(u_m')^\top z_n']v_n =(u_m')^\top\left(\sum_nz_n'v_n^\top\right).

We may therefore aggregate

SV′=∑nzn′vn⊤S_V'=\sum_nz_n'v_n^\top

before allowing the rotated Query to read it. RoPE’s position transformation acts on each Query and Key separately; it does not require a complete N×NN\times N score matrix to exist first. The multiplication reordering that matters most to linear attention therefore remains valid.

9.2.2 Why Does the Denominator Become a Problem?

Many linear-attention methods choose ϕ,φ\phi,\varphi with nonnegative outputs. Then

um⊤zn≥0,u_m^\top z_n\geq 0,

so the denominator is a sum of nonnegative similarities and the output can be interpreted as a weighted average of Values.

Rotation does not preserve the property that “every coordinate is nonnegative.” For example,

R(π)[10]=[−10].R(\pi) \begin{bmatrix} 1\\0 \end{bmatrix} = \begin{bmatrix} -1\\0 \end{bmatrix}.

Thus, even if um,znu_m,z_n originally contain only nonnegative coordinates, the rotated similarity

(um′)⊤zn′(u_m')^\top z_n'

may be negative. If these values are summed directly in the denominator, positive and negative terms can cancel, and the denominator may even approach 00.

9.2.3 What Treatment Does the RoFormer Paper Use?

The linear-attention form presented in RoFormer uses the rotated features only in the numerator, while the denominator retains the original nonnegative features:

om=(um′)⊤(∑nzn′vn⊤)um⊤(∑nzn).\boxed{ o_m = \frac{ (u_m')^\top\left(\sum_nz_n'v_n^\top\right) }{ u_m^\top\left(\sum_nz_n\right) } }.

This preserves two properties:

  1. Both numerator and denominator can still be computed in linear complexity through pre-aggregation.
  2. The denominator continues to use unrotated nonnegative features, reducing the risk of division by zero caused by positive—negative cancellation.

The cost must also be stated clearly. The effective coefficient of Value vnv_n is

wm,n=(um′)⊤zn′∑jum⊤zj.w_{m,n} = \frac{(u_m')^\top z_n'}{\sum_j u_m^\top z_j}.

wm,nw_{m,n} may be negative, and in general it does not satisfy

∑nwm,n=1.\sum_nw_{m,n}=1.

The output is therefore still a content aggregation with a normalized scale, but it is no longer a strict probability-weighted average.

9.2.4 Rotate Before or After the Feature Map?

When reading different implementations, we also encounter two orders that look similar but are genuinely different:

OrderFormMain property
Map, then rotateRmϕ(qm)R_m\phi(q_m)The relative-rotation identity holds exactly, but nonnegative features become signed
Rotate, then mapϕ(Rmqm)\phi(R_mq_m)With positive random features, it can approximate the rotated softmax kernel while preserving nonnegativity, but this is a kernel approximation rather than the exact identity in the row above

For example, Performer’s FAVOR+ uses positive random features to approximate the softmax kernel. Applying RoPE to the original Queries and Keys before mapping the rotated vectors through these features can approximate standard RoPE attention while retaining linear computation. This is a different combination from the RoFormer linear formula in the previous subsection; both should not be collapsed into the single phrase “add RoPE.”

9.3 What Does “Suitable for Linear Attention” Actually Require?

What linear attention truly requires is not one particular positional-encoding name, but a position-aware similarity that can be factored as

sim⁡(qm,kn,m,n)=am(qm)⊤bn(kn).\operatorname{sim}(q_m,k_n,m,n) =a_m(q_m)^\top b_n(k_n).

As long as the Query-side ama_m and Key-side bnb_n can be computed separately, the Keys and Values still have a chance to be aggregated first.

RoPE satisfies this condition naturally:

am(qm)=Rmϕ(qm),bn(kn)=Rnφ(kn).a_m(q_m)=R_m\phi(q_m), \qquad b_n(k_n)=R_n\varphi(k_n).

Many relative-position biases that are added directly to a complete attention matrix, by contrast, cannot be used until every pairwise score (m,n)(m,n) already exists, so the same reordering does not apply directly.

This is not a uniqueness theorem for RoPE. Any other position function that can be separated may also work with linear attention. For example, cosFormer uses separable cosine position reweighting to retain nonnegativity while enabling linear computation. A more accurate conclusion is therefore:

RoPE’s advantage is that it uses absolute-position rotations applied separately to the Query and Key to construct an explicit relative-position interaction. This separable structure is naturally suited to linear attention, but it is not the only possible structure.

The optional material can be compressed into three points:

  1. linear attention depends on a similarity that can be computed separately on the Query and Key sides;
  2. RoPE rotates Queries and Keys separately, so it can preserve this multiplication reordering;
  3. “the computation remains linear” does not imply that the rotated weights remain nonnegative or sum to one—the exact behavior depends on the feature map, rotation order, and normalization scheme.

The single most important sentence from the optional section: RoPE’s relationship to linear attention comes from its separable structure—its transformations can act independently on Queries and Keys. This is a compatibility result, not a prerequisite for defining or understanding RoPE.

References and Image Sources

  1. Jianlin Su (2021). “Transformer Upgrade Path 2: Rotary Position Embedding That Combines the Best of Both Worlds”. This article begins from the questions and RoPE construction in that post, then reorganizes the teaching sequence around two-dimensional rotation, standard attention, and linear attention.
  2. Su, J. et al. (2021). RoFormer: Enhanced Transformer with Rotary Position Embedding. Standard RoPE formulation, properties, and its linear-attention scheme.
  3. Katharopoulos, A. et al. (2020). Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. Linear attention through kernel features and associativity.
  4. Choromanski, K. et al. (2021). Rethinking Attention with Performers. Softmax-attention approximation with positive orthogonal random features.
  5. Qin, Z. et al. (2022). cosFormer: Rethinking Softmax in Attention. Its separable cosine position reweighting shows that RoPE is not the only possible positional structure for linear attention.
  6. Figure 1 is an author-created mechanism diagram.

Series

Rereading the Transformer Upgrade Path

This is Part 2. Later notes will continue through relative position encodings, attention variants, and long-context methods.

Back to post index