Technical note

Positional Embedding Advance - RoPE

Advance to PE technologies: RoPE.

  • Deep Learning
  • Transformer

Fundamentals

the full RoPE rotation matrix is high-dimensional. Mnay studies use the 2D rotation matrix as example, which only one small block inside it.

The key idea is:

RoPE rotates a high-dimensional vector by splitting it into many 2D pairs and rotating each pair by a position-dependent angle.

So for a vector with dimension d=8d=8, RoPE does not use one 2×22\times2 matrix. It uses four 2×22\times2 rotation blocks inside one big 8×88\times8 matrix.


1. Start with ordinary 2D rotation

For a 2D vector

[x0x1]\begin{bmatrix} x_0 \\ x_1 \end{bmatrix}

a rotation by angle ϕ\phi is:

[x0′x1′]=[cos⁡ϕ−sin⁡ϕsin⁡ϕcos⁡ϕ][x0x1]\begin{bmatrix} x_0' \\ x_1' \end{bmatrix} = \begin{bmatrix} \cos\phi & -\sin\phi \\ \sin\phi & \cos\phi \end{bmatrix} \begin{bmatrix} x_0 \\ x_1 \end{bmatrix}

So:

x0′=x0cos⁡ϕ−x1sin⁡ϕx_0'=x_0\cos\phi-x_1\sin\phi x1′=x0sin⁡ϕ+x1cos⁡ϕx_1'=x_0\sin\phi+x_1\cos\phi

That is the small 2D matrix I showed earlier.

But RoPE applies this idea repeatedly across many pairs of dimensions.


2. High-dimensional RoPE rotation

Suppose my attention head dimension is d=8d=8. A query vector may look like:

q=[q0q1q2q3q4q5q6q7]q= \begin{bmatrix} q_0\\ q_1\\ q_2\\ q_3\\ q_4\\ q_5\\ q_6\\ q_7 \end{bmatrix}

RoPE groups it into pairs:

(q0,q1),(q2,q3),(q4,q5),(q6,q7)(q_0,q_1),\quad(q_2,q_3),\quad(q_4,q_5),\quad(q_6,q_7)

Each pair is rotated separately.

For position pp, RoPE computes several angles:

ϕp,0=pθ0\phi_{p,0}=p\theta_0 ϕp,1=pθ1\phi_{p,1}=p\theta_1 ϕp,2=pθ2\phi_{p,2}=p\theta_2 ϕp,3=pθ3\phi_{p,3}=p\theta_3

Each pair gets its own angle. So the full rotation matrix is:

Rp=[cos⁡ϕp,0−sin⁡ϕp,0000000sin⁡ϕp,0cos⁡ϕp,000000000cos⁡ϕp,1−sin⁡ϕp,1000000sin⁡ϕp,1cos⁡ϕp,100000000cos⁡ϕp,2−sin⁡ϕp,2000000sin⁡ϕp,2cos⁡ϕp,200000000cos⁡ϕp,3−sin⁡ϕp,3000000sin⁡ϕp,3cos⁡ϕp,3]R_p= \begin{bmatrix} \cos\phi_{p,0} & -\sin\phi_{p,0} & 0 & 0 & 0 & 0 & 0 & 0 \\ \sin\phi_{p,0} & \cos\phi_{p,0} & 0 & 0 & 0 & 0 & 0 & 0 \\ 0 & 0 & \cos\phi_{p,1} & -\sin\phi_{p,1} & 0 & 0 & 0 & 0 \\ 0 & 0 & \sin\phi_{p,1} & \cos\phi_{p,1} & 0 & 0 & 0 & 0 \\ 0 & 0 & 0 & 0 & \cos\phi_{p,2} & -\sin\phi_{p,2} & 0 & 0 \\ 0 & 0 & 0 & 0 & \sin\phi_{p,2} & \cos\phi_{p,2} & 0 & 0 \\ 0 & 0 & 0 & 0 & 0 & 0 & \cos\phi_{p,3} & -\sin\phi_{p,3} \\ 0 & 0 & 0 & 0 & 0 & 0 & \sin\phi_{p,3} & \cos\phi_{p,3} \end{bmatrix}

This is an 8×88\times8 matrix.

So,

RoPE is high-dimensional, but it is built from many independent 2D rotations.


3. Concrete numerical example

Take a 6-dimensional vector:

x=[102030]x= \begin{bmatrix} 1\\ 0\\ 2\\ 0\\ 3\\ 0 \end{bmatrix}

Group it into three 2D pairs:

(1,0),(2,0),(3,0)(1,0),\quad(2,0),\quad(3,0)

Now suppose this token is at position:

p=4p=4

For illustration, use three RoPE frequencies:

θ0=1,θ1=0.1,θ2=0.01\theta_0=1,\quad \theta_1=0.1,\quad \theta_2=0.01

Then the angles are:

ϕ4,0=4×1=4\phi_{4,0}=4\times1=4 ϕ4,1=4×0.1=0.4\phi_{4,1}=4\times0.1=0.4 ϕ4,2=4×0.01=0.04\phi_{4,2}=4\times0.01=0.04

So the first pair rotates by 44 radians, the second pair by 0.40.4 radians, and the third pair by 0.040.04 radians.

Now rotate each pair.


First pair

(1,0)(1,0)

rotated by 44 radians:

x0′=1cos⁡4−0sin⁡4x_0' = 1\cos 4 - 0\sin 4 x1′=1sin⁡4+0cos⁡4x_1' = 1\sin 4 + 0\cos 4

Since

cos⁡4≈−0.654,sin⁡4≈−0.757\cos 4 \approx -0.654,\quad \sin 4 \approx -0.757

we get:

(1,0)→(−0.654,−0.757)(1,0)\rightarrow(-0.654,-0.757)

Second pair

(2,0)(2,0)

rotated by 0.40.4 radians:

x2′=2cos⁡0.4x_2'=2\cos0.4 x3′=2sin⁡0.4x_3'=2\sin0.4

Since

cos⁡0.4≈0.921,sin⁡0.4≈0.389\cos0.4\approx0.921,\quad \sin0.4\approx0.389

we get:

(2,0)→(1.842,0.779)(2,0)\rightarrow(1.842,0.779)

Third pair

(3,0)(3,0)

rotated by 0.040.04 radians:

x4′=3cos⁡0.04x_4'=3\cos0.04 x5′=3sin⁡0.04x_5'=3\sin0.04

Since

cos⁡0.04≈0.999,sin⁡0.04≈0.040\cos0.04\approx0.999,\quad \sin0.04\approx0.040

we get:

(3,0)→(2.998,0.120)(3,0)\rightarrow(2.998,0.120)

So the whole vector becomes:

x=[102030]x= \begin{bmatrix} 1\\ 0\\ 2\\ 0\\ 3\\ 0 \end{bmatrix} ⇓\Downarrow R4x≈[−0.654−0.7571.8420.7792.9980.120]R_4x \approx \begin{bmatrix} -0.654\\ -0.757\\ 1.842\\ 0.779\\ 2.998\\ 0.120 \end{bmatrix}

That is how RoPE rotates a high-dimensional vector.

It does not rotate the whole vector with one simple 2D angle. Instead, it rotates different 2D slices of the vector by different angles.


4. Why different dimension pairs use different angles

In real RoPE, the frequencies are usually chosen like this:

θm=10000−2m/d\theta_m = 10000^{-2m/d}

where mm is the pair index.

For example, if d=8d=8, then there are four pairs:

m=0,1,2,3m=0,1,2,3

The angles are:

ϕp,m=p⋅10000−2m/d\phi_{p,m}=p\cdot 10000^{-2m/d}

So early dimension pairs rotate quickly, while later dimension pairs rotate slowly.

This is similar to sinusoidal positional encoding. Some dimensions represent fine-grained local position, while others represent broad long-range position.

I can think of the dimension pairs as different clocks:

(q0,q1):fast clock(q_0,q_1): \text{fast clock} (q2,q3):medium clock(q_2,q_3): \text{medium clock} (q4,q5):slow clock(q_4,q_5): \text{slow clock} (q6,q7):very slow clock(q_6,q_7): \text{very slow clock}

So when the position changes, each pair rotates by a different amount.


5. Why this produces relative position information

Now suppose we have query vector qiq_i at position ii, and key vector kjk_j at position jj.

RoPE rotates them:

qi′=Riqiq_i' = R_i q_i kj′=Rjkjk_j' = R_j k_j

Then attention computes:

(qi′)⊤kj′(q_i')^\top k_j'

Substitute the rotated forms:

(qi′)⊤kj′=(Riqi)⊤(Rjkj)(q_i')^\top k_j' = (R_iq_i)^\top(R_jk_j) =qi⊤Ri⊤Rjkj= q_i^\top R_i^\top R_j k_j

For rotation matrices:

Ri⊤Rj=Rj−iR_i^\top R_j = R_{j-i}

So:

(qi′)⊤kj′=qi⊤Rj−ikj(q_i')^\top k_j' = q_i^\top R_{j-i} k_j

This is the beautiful part.

Even though RoPE rotates each vector using its absolute position, the final query-key dot product depends on the relative position:

j−ij-i

So RoPE turns absolute position rotations into relative-position-aware attention.


6. A useful analogy

Imagine every token vector lives in many tiny 2D planes:

Plane 1: (q0,q1)\text{Plane 1: }(q_0,q_1) Plane 2: (q2,q3)\text{Plane 2: }(q_2,q_3) Plane 3: (q4,q5)\text{Plane 3: }(q_4,q_5) Plane 4: (q6,q7)\text{Plane 4: }(q_6,q_7)

At position pp, RoPE rotates the vector in each plane.

But each plane rotates at a different speed.

So position is encoded not as an added vector, but as a set of phases across many 2D planes.

That is why RoPE is called rotary positional embedding.


7. Compact interpretation

For a high-dimensional vector, RoPE does this:

[x0,x1,x2,x3,x4,x5,…][x_0,x_1,x_2,x_3,x_4,x_5,\ldots]

becomes:

[rotate(x0,x1,ϕ0),rotate(x2,x3,ϕ1),rotate(x4,x5,ϕ2),…][ \text{rotate}(x_0,x_1,\phi_0), \text{rotate}(x_2,x_3,\phi_1), \text{rotate}(x_4,x_5,\phi_2), \ldots ]

So the 2D rotation matrix is just one repeated building block.

The full RoPE rotation is a block-diagonal high-dimensional rotation matrix:

Rp=diag(R(ϕp,0),R(ϕp,1),R(ϕp,2),…)R_p= \text{diag} \left( R(\phi_{p,0}), R(\phi_{p,1}), R(\phi_{p,2}), \ldots \right)

where each block is:

R(ϕ)=[cos⁡ϕ−sin⁡ϕsin⁡ϕcos⁡ϕ]R(\phi)= \begin{bmatrix} \cos\phi & -\sin\phi\\ \sin\phi & \cos\phi \end{bmatrix}

That is the cleanest mental model:

RoPE rotates a high-dimensional vector by performing many small 2D rotations on pairs of dimensions. Each pair rotates at a different frequency, and the rotation angle depends on the token position.

Math behind RoPE

RoPE is not just a random intuition validated by experiments. It has a clean mathematical motivation. But the claim that “RoPE is better for language models” is largely empirical and mechanistic.

So RoPE has theory for what it encodes, but not a complete theory for why it always improves model performance.


1. The theoretical goal of RoPE

The original RoPE paper starts from a desired property:

⟨qmPE,knPE⟩\langle q_m^{PE}, k_n^{PE} \rangle

should depend on the relative position

n−mn-m

rather than depending independently on mm and nn.

In other words, we want:

⟨fq(q,m),fk(k,n)⟩=g(q,k,n−m)\langle f_q(q,m), f_k(k,n) \rangle = g(q,k,n-m)

This is a formal design target, not merely intuition. The RoFormer paper explicitly formulates RoPE as a way to encode absolute position through rotation while making the self-attention score depend on relative position. (arXiv)

Using the rotation matrices introduced in the Fundamentals chapter, RoPE encodes the query and key as qmPE=Rmqq_m^{PE}=R_mq and knPE=Rnkk_n^{PE}=R_nk. The derivation above then gives:

(qmPE)⊤knPE=q⊤Rn−mk(q_m^{PE})^\top k_n^{PE} = q^\top R_{n-m}k

This is the central theoretical justification: RoPE converts absolute rotations at positions mm and nn into a relative rotation depending on n−mn-m.


2. Why rotation is mathematically natural here

Positions in a sequence behave like translations:

m→m+1m \rightarrow m+1 m→m+2m \rightarrow m+2 m→m+cm \rightarrow m+c

Rotations also compose naturally:

RaRb=Ra+bR_aR_b=R_{a+b} Ra−1=R−aR_a^{-1}=R_{-a}

So if representing position mm by a rotation RmR_m, then comparing two positions automatically gives:

Rm−1Rn=Rn−mR_m^{-1}R_n=R_{n-m}

This is why RoPE is elegant: relative position appears automatically from the algebra of rotations.

In more abstract language, RoPE uses a representation of the additive position group through rotations. In simpler language:

moving forward in sequence position corresponds to rotating in hidden-vector space.

That is a real theoretical structure.


3. The connection to complex numbers and Fourier features

There is another theoretical background: RoPE is closely related to complex exponentials.

A 2D vector

(x1,x2)(x_1,x_2)

can be written as a complex number:

z=x1+ix2z=x_1+ix_2

Rotating by angle θ\theta is the same as multiplying by:

eiθe^{i\theta}

So RoPE can be viewed as:

zmPE=zeimθz_m^{PE}=z e^{im\theta}

For two positions mm and nn:

e−imθeinθ=ei(n−m)θe^{-im\theta}e^{in\theta} = e^{i(n-m)\theta}

Again, the relative distance n−mn-m appears naturally.

This is why RoPE is related to sinusoidal positional encoding. Both use sine/cosine waves or complex phases. The difference is that sinusoidal PE is usually added to token embeddings, while RoPE multiplies/rotates query and key vectors. The RoFormer paper specifically notes that RoPE incorporates relative position information by multiplying with sinusoidal functions rather than directly adding position to contextual representations. (arXiv)


4. The high-dimensional theory

For a high-dimensional vector, RoPE uses many 2D rotation planes:

Rm=diag(R(mθ0),R(mθ1),R(mθ2),…)R_m= \text{diag} \left( R(m\theta_0), R(m\theta_1), R(m\theta_2), \ldots \right)

Each frequency θi\theta_i gives a different positional scale.

So the full attention score is roughly:

(Rmq)⊤(Rnk)=∑iqi⊤R((n−m)θi)ki(R_m q)^\top(R_n k) = \sum_i q_i^\top R((n-m)\theta_i)k_i

This gives the model many “relative-distance channels”:

θ0:fast local position signal\theta_0: \text{fast local position signal} θ1:medium-scale position signal\theta_1: \text{medium-scale position signal} θ2:longer-range position signal\theta_2: \text{longer-range position signal}

and so on.

This is another reason RoPE is not arbitrary. It is a multiscale Fourier-like positional system.


5. What the original paper theoretically claims

The original RoPE paper claims several mathematical/structural advantages:

First, RoPE encodes absolute position using a rotation matrix while making attention explicitly sensitive to relative position. Second, because it does not require a learned lookup table for each position, it is flexible with respect to sequence length. Third, the authors analyze a long-term decay property: under their frequency setting, certain RoPE inner-product terms decay as relative distance increases. Fourth, they argue that RoPE can be used with linear self-attention more naturally than some previous additive relative-position methods. (arXiv)

So yes, the original paper did not simply say:

“Let’s rotate vectors; maybe it works.”

It said something closer to:

“We want relative-position-dependent attention scores. Rotation gives us that property exactly. Then we validate that this useful inductive bias improves performance.”


6. But the performance advantage is not fully theoretically explained

The theory explains what RoPE does geometrically. It does not fully explain why RoPE makes a large Transformer better at language modeling, reasoning, or long-context understanding.

Even the RoFormer authors acknowledge limitations. They state that although they provide theoretical grounding and experiments, they do not have a thorough explanation for why RoPE converges faster than baseline positional encoding methods, nor a faithful explanation for why it performs better on long texts than peer models. (arXiv)

RoPE is partly a theoretically motivated design, and partly an empirical success.


7. A modern nuance: the “distance decay” explanation is incomplete

One common explanation is:

RoPE works because attention naturally decays with distance.

That is only partly true.

Later analysis argues that this decay explanation is too simple. A 2024 paper, Round and Round We Go! What makes Rotary Positional Encodings useful?, argues that RoPE does not necessarily make activations decay with distance in general. For example, they show that decay may appear under special query/key assumptions, but not necessarily for random Gaussian query/key vectors. They also show that RoPE can help construct attention heads that attend to specific relative positions. (arXiv)

This is important.

RoPE should not be understood as merely:

farther token⇒lower attention\text{farther token} \Rightarrow \text{lower attention}

A better understanding (more precise) is:

RoPE gives attention heads a geometric mechanism for comparing content at specific relative offsets.\text{RoPE gives attention heads a geometric mechanism for comparing content at specific relative offsets.}

8. Length extrapolation is also not fully solved by RoPE

RoPE is often associated with better length extrapolation than learned absolute position embeddings, but it does not magically solve length extrapolation.

A NeurIPS 2023 study comparing positional encodings found that the exact impact of different PE schemes on extrapolation remains unclear, and in their downstream length-generalization tasks, common explicit PE methods including Rotary, ALiBi, and absolute position embeddings were not always well suited; their NoPE baseline performed surprisingly well. (NeurIPS Proceedings)

So we should avoid saying:

RoPE theoretically guarantees long-context generalization.

It does not.

A more accurate statement is:

RoPE has better structural properties than learned absolute position tables for variable-length use, but long-context generalization still depends on training distribution, frequency scale, architecture, data, and fine-tuning.


Other absolute PE + relative PE ideas/concepts

For non-RoPE methods, the usual way to combine absolute PE and relative PE is to let them play different roles:

absolute PE: “where am I?”\text{absolute PE: “where am I?”} relative PE: “how am I positioned with respect to you?”\text{relative PE: “how am I positioned with respect to you?”}

The central design choice is whether to mix them inside token representations or keep them disentangled as separate attention-score terms.

One nuance first: RoPE is often described as encoding both absolute and relative position, but standard RoPE is better described as absolute-indexed in construction and relative in the attention logit. The query/key vector at position ii is rotated by an absolute phase RiR_i, but when attention computes (Riqi)⊤(Rjkj)(R_i q_i)^\top(R_j k_j), the absolute rotations combine into Rj−iR_{j-i}. So the pairwise attention score mainly sees relative offset j−ij-i, not an independent absolute-position label.


1. Naive combination: add absolute PE to tokens, add relative PE to attention

This is the simplest combination.

Let

xi=content embedding at position ix_i = \text{content embedding at position } i ai=absolute position embeddinga_i = \text{absolute position embedding} bi−j=relative position biasb_{i-j} = \text{relative position bias}

Then use:

hi=xi+aih_i = x_i + a_i

and attention score:

sij=(hiWQ)(hjWK)⊤d+bi−js_{ij} = \frac{(h_iW_Q)(h_jW_K)^\top}{\sqrt d} + b_{i-j}

So the model receives absolute position through hih_i, and relative position through bi−jb_{i-j}.

This is conceptually simple:

attention score=content/absolute-position mixed similarity+relative-distance bias\text{attention score} = \text{content/absolute-position mixed similarity} + \text{relative-distance bias}

The disadvantage is that the absolute position vector is mixed into the token vector before attention. Expanding the dot product shows the entanglement:

(hiWQ)(hjWK)⊤=(xiWQ+aiWQ)(xjWK+ajWK)⊤(h_iW_Q)(h_jW_K)^\top = (x_iW_Q + a_iW_Q)(x_jW_K + a_jW_K)^\top =xiWQ(xjWK)⊤⏟content-content+xiWQ(ajWK)⊤⏟content-absolute position+aiWQ(xjWK)⊤⏟absolute position-content+aiWQ(ajWK)⊤⏟absolute position-position= \underbrace{x_iW_Q(x_jW_K)^\top}_{\text{content-content}} + \underbrace{x_iW_Q(a_jW_K)^\top}_{\text{content-absolute position}} + \underbrace{a_iW_Q(x_jW_K)^\top}_{\text{absolute position-content}} + \underbrace{a_iW_Q(a_jW_K)^\top}_{\text{absolute position-position}}

So this approach combines absolute and relative position, but it also creates mixed content-position interactions that may or may not be useful.

This method is common because it is easy: for example, I can take a BERT/ViT-style model with learned absolute embeddings and add a relative position bias term to the attention logits.


2. Relative key/value combination: absolute PE at input, relative PE inside attention

A more expressive version uses relative vectors, not just scalar biases.

A Shaw-style relative attention score can look like:

sij=qi⊤(kj+ri−jK)ds_{ij} = \frac{q_i^\top(k_j + r^K_{i-j})}{\sqrt d}

where ri−jKr^K_{i-j} is a learned relative-position key vector.

The output can also use a relative value vector:

oi=∑jαij(vj+ri−jV)o_i = \sum_j \alpha_{ij}(v_j + r^V_{i-j})

Then we can still add absolute PE to the input:

hi=xi+aih_i = x_i + a_i

So the full system has:

absolute position: hi=xi+ai\text{absolute position: } h_i=x_i+a_i relative position in attention score: qi⊤ri−jK\text{relative position in attention score: } q_i^\top r^K_{i-j} relative position in aggregated value: ri−jV\text{relative position in aggregated value: } r^V_{i-j}

This is more expressive than scalar relative bias because the relative position can interact with the query vector. Shaw et al. introduced relative position representations for self-attention and found that relative position improved machine translation over absolute position in their experiments; interestingly, they also reported that combining relative and absolute position gave no further translation improvement in that setting. (arXiv)

So the lesson is important:

Combining absolute and relative PE is possible, but it is not automatically better.

Sometimes relative position already contains the useful ordering signal.


3. Disentangled combination: keep content, absolute position, and relative position separate

This is the most conceptually clean approach.

Instead of first doing:

hi=xi+aih_i=x_i+a_i

we keep content and position as different objects:

xi=contentx_i = \text{content} ai=absolute positiona_i = \text{absolute position} ri−j=relative positionr_{i-j} = \text{relative position}

Then define attention score as a sum of separate components:

sij=sijcontent+sijabsolute+sijrelatives_{ij} = s^{content}_{ij} + s^{absolute}_{ij} + s^{relative}_{ij}

For example:

sijcontent=(xiWQc)(xjWKc)⊤ds^{content}_{ij} = \frac{(x_iW_Q^c)(x_jW_K^c)^\top}{\sqrt d} sijabsolute=(aiWQa)(ajWKa)⊤ds^{absolute}_{ij} = \frac{(a_iW_Q^a)(a_jW_K^a)^\top}{\sqrt d} sijrelative=bi−js^{relative}_{ij} = b_{i-j}

So:

sij=(xiWQc)(xjWKc)⊤d+(aiWQa)(ajWKa)⊤d+bi−js_{ij} = \frac{(x_iW_Q^c)(x_jW_K^c)^\top}{\sqrt d} + \frac{(a_iW_Q^a)(a_jW_K^a)^\top}{\sqrt d} + b_{i-j}

This gives a clear separation:

content similarity+absolute-position compatibility+relative-distance compatibility\text{content similarity} + \text{absolute-position compatibility} + \text{relative-distance compatibility}

This idea is close to TUPE, “Transformer with Untied Positional Encoding.” TUPE argues that simply adding word embeddings and position embeddings creates mixed correlations between heterogeneous sources of information; it instead computes word contextual correlation and positional correlation separately, then adds them together in attention. (OpenReview)

This is a strong design pattern when concerning: additive absolute PE can act like “position noise” inside content attention. Disentangling lets the model use position without forcing content and position to live in the same representation channel.


4. DeBERTa-style combination: relative position in attention, absolute position near decoding/output

DeBERTa gives a very instructive example.

BERT-style models usually do:

hi=xi+aih_i=x_i+a_i

DeBERTa instead represents each word using separate content and position vectors. Its disentangled attention computes attention using content and relative-position information, rather than simply mixing word and position embeddings at the input. The DeBERTa paper also notes that relative information alone may miss absolute-position cues, so it incorporates absolute position information in the enhanced mask decoder used for masked-token prediction. (ar5iv)

Conceptually, DeBERTa separates the jobs:

attention layers: use content + relative position\text{attention layers: use content + relative position} prediction/decoding layer: inject absolute position\text{prediction/decoding layer: inject absolute position}

This is a very elegant pattern.

Why? Because during attention, relative position is often what matters:

“Is this word next to that word?” “Is this patch above or below that patch?” “Is this token far away or nearby?”

But during output prediction, absolute position can matter:

“Is this word near the beginning of the sentence?” “Is this token in the title region?” “Is this patch in the top-left of the image?” “Is this timestep early or late in the sequence?”

So DeBERTa-style design says:

Use relative position for pairwise interaction, but keep absolute position available when the model needs to make position-specific predictions.


5. Head-wise or layer-wise combination

Another practical way is to let different heads or layers specialize.

For attention head hh, use:

sij(h)=qi(h)(kj(h))⊤d+bi−j(h)s_{ij}^{(h)} = \frac{q_i^{(h)}(k_j^{(h)})^\top}{\sqrt d} + b^{(h)}_{i-j}

Some heads can receive relative bias; other heads can use absolute PE; some layers may use both.

For example:

early layers: relative/local position\text{early layers: relative/local position} middle layers: mixed content-position structure\text{middle layers: mixed content-position structure} late layers: task-specific absolute position\text{late layers: task-specific absolute position}

This can help because positional needs are not uniform across the network. Early layers may need local ordering, while later layers may need global layout or beginning/end information.

This idea is especially natural in vision and document models, where some heads should learn local geometry while others learn global layout.


6. 2D combination for vision: absolute coordinates plus relative offsets

For images, absolute and relative position mean slightly different things.

Absolute position:

(rowi,coli)(row_i, col_i)

Relative position:

(rowi−rowj, coli−colj)(row_i-row_j,\ col_i-col_j)

A combined 2D PE can be:

hi=xi+arowi,colih_i = x_i + a_{row_i,col_i} sij=qi⊤kjd+bΔrow,Δcols_{ij} = \frac{q_i^\top k_j}{\sqrt d} + b_{\Delta row,\Delta col}

where

Δrow=rowi−rowj\Delta row=row_i-row_j Δcol=coli−colj\Delta col=col_i-col_j

This gives the model both:

global location: top, bottom, center, corner\text{global location: top, bottom, center, corner}

and

local geometry: above, below, left, right, nearby, far\text{local geometry: above, below, left, right, nearby, far}

But again, combining is not always better. Swin Transformer’s ablation is a useful example. In Swin-T, relative position bias outperformed no position encoding and absolute position embedding, and adding absolute position to relative position was close but did not beat the default relative-only setting on the reported classification/detection/segmentation metrics. The authors also noted that absolute position improved ImageNet classification slightly but hurt COCO detection and ADE20K segmentation, suggesting that relative/local translation-friendly bias is especially valuable for dense vision tasks. (ar5iv)

So in vision:

Absolute PE can help classification or layout-sensitive tasks. Relative PE is often better for detection, segmentation, and local geometric reasoning.


7. Graph combination: node-level absolute PE plus pairwise relative PE

Graphs are another case where combining absolute and relative PE can be very useful.

For a graph Transformer, absolute PE might be a node feature:

ai=Laplacian eigenvector position of node ia_i = \text{Laplacian eigenvector position of node } i

or

ai=centrality/degree/random-walk feature of node ia_i = \text{centrality/degree/random-walk feature of node } i

Relative PE might be a pairwise feature:

rij=shortest-path distance between nodes i,jr_{ij} = \text{shortest-path distance between nodes } i,j

or an edge/path encoding.

Then use:

hi=xi+aih_i = x_i + a_i sij=qi⊤kjd+b(rij)s_{ij} = \frac{q_i^\top k_j}{\sqrt d} + b(r_{ij})

This is very natural:

absolute graph PE: “what structural role does this node have?”\text{absolute graph PE: “what structural role does this node have?”} relative graph PE: “how are these two nodes connected?”\text{relative graph PE: “how are these two nodes connected?”}

Graphormer is an important example of this philosophy: its key idea is that Transformers need graph structural information encoded into the model, and it proposes structural encodings to help model graph-structured data. (arXiv) A later graph-Transformer PE comparison frames the distinction clearly: absolute PEs assign features to each node and are given as input to the Transformer, while relative PEs assign features to pairs of nodes, such as graph distance, and augment the attention block. (arXiv)


8. When combined absolute + relative PE helps

Combined PE helps when the task needs both of these questions answered:

Absolute: “Where is this item globally?”\textbf{Absolute: “Where is this item globally?”} Relative: “How is this item related to that item?”\textbf{Relative: “How is this item related to that item?”}

Case A: The task depends on global position

Absolute PE helps when the same local pattern means different things in different global positions.

Examples:

beginning vs. end of sentence\text{beginning vs. end of sentence} title vs. body in a document\text{title vs. body in a document} top vs. bottom of image\text{top vs. bottom of image} early vs. late timestep\text{early vs. late timestep} root-near vs. leaf-near node in a graph\text{root-near vs. leaf-near node in a graph}

Relative PE alone can say:

token A is 3 positions before token B\text{token A is 3 positions before token B}

but it may not directly say:

token A is near the beginning of the document\text{token A is near the beginning of the document}

The DeBERTa example above uses exactly this division of labor: relative position guides attention, while absolute position remains available near prediction.

Case B: The task also needs local relational geometry

Relative PE helps when pairwise spatial structure matters.

Examples:

adjacent words\text{adjacent words} nearby image patches\text{nearby image patches} neighboring graph nodes\text{neighboring graph nodes} nearby time steps\text{nearby time steps} query span and answer span in QA\text{query span and answer span in QA}

Relative PE is usually better than absolute PE for saying:

“these two things are near each other”\text{“these two things are near each other”}

because relative distance is directly represented rather than needing to be inferred from two absolute indices.

Case C: I need both global anchors and translation-like invariance

This is common in vision.

For object detection, I want the model to recognize a dog whether it appears on the left or right side of the image. That favors relative/local position and some translation invariance. But for document understanding, satellite imagery, medical scans, or UI screenshots, absolute location can also matter. The Swin results above illustrate this trade-off between translation-friendly relative bias and global position cues.

Case D: I want more expressive attention

A pure scalar relative bias says:

sij=qi⊤kj+bi−js_{ij}=q_i^\top k_j+b_{i-j}

This says:

“At this distance, attention should be easier or harder.”

But richer combined methods can say:

“This type of token should attend to that type of token differently depending on both absolute location and relative displacement.”

That is why some methods increase interaction between query, key, and relative position embeddings rather than using only a simple bias. Huang et al. argued that position information was not fully used in existing methods and proposed techniques to encourage stronger interaction between query, key, and relative position embeddings, reporting improved SQuAD1.1 results compared with previous absolute and relative PE approaches. (ACL Anthology)


9. When combined PE may not help

Combined PE can hurt or fail to help when absolute position is redundant, noisy, or harmful.

The Shaw translation and Swin vision results discussed above provide two examples: an added absolute signal may contribute nothing when relative order is sufficient, and it may weaken translation-friendly behavior in dense vision tasks.

Combined PE may also hurt length or resolution extrapolation. Learned absolute PE often binds the model to positions seen during training, while relative PE usually generalizes more naturally to shifted patterns. So if the model must extrapolate to much longer sequences or different image sizes, adding a strong learned absolute embedding can become a liability.


10. Practical rule

A good practical design is:

Use relative PE for attention; add absolute PE only where global anchors matter.\boxed{ \text{Use relative PE for attention; add absolute PE only where global anchors matter.} }

More specifically:

SettingGood PE combination
Text encoder, fixed lengthlearned absolute PE + relative bias can work
Text encoder with strong syntax/QA needsdisentangled content/absolute/relative PE
Decoder-only LLMRoPE or ALiBi usually enough; extra absolute PE can hurt extrapolation
Masked language modelrelative attention + absolute position near decoder/output can help
Image classificationabsolute 2D PE + relative bias may help
Detection/segmentationrelative 2D bias often better than strong absolute PE
Document/layout modelabsolute page coordinates + relative local offsets
Graph Transformernode-level structural PE + pairwise distance/edge bias
Time seriesabsolute time/calendar features + relative lag bias

The cleanest conceptual formula is:

sij=C(xi,xj)⏟content similarity+A(ai,aj)⏟absolute-position compatibility+R(rij)⏟relative-position compatibilitys_{ij} = \underbrace{C(x_i,x_j)}_{\text{content similarity}} + \underbrace{A(a_i,a_j)}_{\text{absolute-position compatibility}} + \underbrace{R(r_{ij})}_{\text{relative-position compatibility}}

This is better than blindly doing:

hi=xi+aih_i=x_i+a_i

when I care about separating content from position.

So the final answer is:

To combine absolute and relative PE, either add absolute PE to token embeddings and add relative PE to attention, or more cleanly, compute content, absolute-position, and relative-position terms separately and add them at the attention-logit level. Combined PE helps when both global location and pairwise geometry matter, but it is not universally better; in tasks where translation invariance or length extrapolation is important, relative-only PE can be superior.

RoPE in long context

In modern long-context LLMs, positional encoding is one of the central technologies that makes 128K, 1M, or even multi-million-token context possible, but it is only one part of the solution.

A good way to think about it:

Long context=long-range PE+training/fine-tuning on long sequences+efficient attention/KV-cache infrastructure+evaluation for long-context recall\text{Long context} = \text{long-range PE} + \text{training/fine-tuning on long sequences} + \text{efficient attention/KV-cache infrastructure} + \text{evaluation for long-context recall}

PE gives the model an address system for token positions. But PE alone does not guarantee that the model can actually retrieve, reason over, or prioritize information across 1M tokens.

Google’s Gemini API documentation says many Gemini models support context windows of 1 million or more tokens, and the Gemini 1.5 report describes models capable of recalling and reasoning over fine-grained information from millions of tokens of context. (Google AI for Developers)


1. Long context is hard for PE

For a normal decoder-only LLM, each token has a position:

0,1,2,…,N−10,1,2,\ldots,N-1

If N=1,000,000N=1{,}000{,}000, the PE system must give meaningful positional information for token 0, token 500,000, token 999,999, and every relative distance between them.

That creates several problems.

First, learned absolute PE does not scale well. A model trained with a learned table for 4K or 8K positions does not automatically have trained embeddings for position 900K. User could make a huge learned table, but it would be expensive, poorly extrapolative, and very data-hungry.

Second, plain RoPE extrapolation can become unstable. RoPE can mathematically compute rotations for any position, but if the model was trained only up to 4K or 8K, positions like 500K are far outside the phase range it learned to use. The Position Interpolation paper argues that directly extrapolating beyond the trained context can produce extreme attention behavior, while interpolating positions back into the trained range is much more stable. (arXiv)

Third, long context is not only a PE problem. Even with perfect PE, full attention over 1M tokens is extremely expensive. The practical system also needs efficient attention kernels, distributed context parallelism, KV-cache management, and training examples that actually teach the model to use long evidence. A 2024 context-parallelism paper, for example, studies scalable million-token inference and reports 1M-context prefill experiments on large models using many GPUs. (arXiv)


For standard RoPE, the relative-position result derived earlier applies directly:

(qi′)⊤kj′=qi⊤Rj−ikj(q_i')^\top k_j'=q_i^\top R_{j-i}k_j

So RoPE gives a very useful property:

absolute position in construction⇒relative position in attention\text{absolute position in construction} \quad\Rightarrow\quad \text{relative position in attention}

This is why RoPE is attractive for LLMs. It does not need a learned table for every position, and the query-key comparison naturally depends on relative offset. The original RoPE paper describes this as encoding absolute position with a rotation matrix while incorporating explicit relative-position dependency in self-attention. (arXiv)

But standard RoPE still has a problem at 1M context: the rotation angles can go far outside the training distribution.

For each dimension pair, RoPE uses something like:

ϕi,m=iθm\phi_{i,m}=i\theta_m

where ii is the token position and θm\theta_m is the frequency for dimension pair mm.

At position i=1,000,000i=1{,}000{,}000, the angle

1,000,000θm1{,}000{,}000\theta_m

may be much larger than anything the model saw during training. So long-context LLMs usually do not just “turn on RoPE to 1M.” They modify how RoPE maps positions to angles.


3. Main strategy: RoPE scaling

The core idea of RoPE scaling is simple:

do not feed the raw position i into RoPE\text{do not feed the raw position } i \text{ into RoPE}

Instead, use an effective position:

i~=f(i)\tilde{i}=f(i)

Then RoPE becomes:

qi′=Ri~qiq_i' = R_{\tilde{i}}q_i kj′=Rj~kjk_j' = R_{\tilde{j}}k_j

The attention comparison becomes:

(qi′)⊤kj′=qi⊤Rj~−i~kj(q_i')^\top k_j' = q_i^\top R_{\tilde{j}-\tilde{i}}k_j

So the model still gets relative position information, but now relative position is measured in a rescaled coordinate system.


4. Simple position interpolation

Suppose a model was trained with context length:

Ltrain=8192L_{\text{train}}=8192

and we want to run it at:

Ltarget=1,048,576L_{\text{target}}=1{,}048{,}576

The scale factor is:

s=LtargetLtrain=128s=\frac{L_{\text{target}}}{L_{\text{train}}}=128

Position interpolation uses:

i~=is\tilde{i}=\frac{i}{s}

So token position 1,000,000 becomes:

i~=1,000,000128=7812.5\tilde{i}=\frac{1{,}000{,}000}{128}=7812.5

That means token #1,000,000 is represented using a RoPE phase similar to a token near position 7,812 in the original model’s positional space.

This is the central trick:

very long real context→compressed RoPE coordinate range\text{very long real context} \rightarrow \text{compressed RoPE coordinate range}

The Position Interpolation paper showed that RoPE-based LLaMA models could be extended to 32K context with limited fine-tuning by down-scaling position indices rather than directly extrapolating them. (arXiv)

But there is a cost.

A real distance of 128 tokens becomes:

128128=1\frac{128}{128}=1

in RoPE space.

A real distance of 1024 tokens becomes:

1024128=8\frac{1024}{128}=8

in RoPE space.

So the model’s positional ruler becomes compressed. This helps avoid out-of-distribution high positions, but it can reduce the model’s ability to distinguish nearby positions precisely. That is why simple interpolation is useful but not ideal for very long contexts.


5. NTK-aware scaling: change the RoPE frequency scale

Another strategy is to modify the RoPE frequency base. Instead of only shrinking positions:

i~=i/s\tilde{i}=i/s

we change the frequency schedule:

θm=base−2m/d\theta_m = \text{base}^{-2m/d}

If we increase the base, many RoPE dimensions rotate more slowly. Slower rotations mean longer usable wavelengths, which helps the model represent longer contexts.

Intuitively:

larger RoPE base⇒slower positional phase change⇒longer context coverage\text{larger RoPE base} \Rightarrow \text{slower positional phase change} \Rightarrow \text{longer context coverage}

But this creates a tradeoff. If we slow all frequencies too much, the model loses fine local position resolution. If we do not slow them enough, far positions remain out of distribution.

That is the central tension in long-context PE:

preserve local precisionvs.support very long distance\text{preserve local precision} \quad \text{vs.} \quad \text{support very long distance}

6. YaRN: scale different frequency bands more carefully

YaRN, short for “Yet another RoPE extensioN method,” improves on simple interpolation. Its goal is to extend context length efficiently without requiring as much long-context training as naive fine-tuning. The YaRN paper says RoPE-based models fail to generalize far past the sequence length they were trained on, and proposes a compute-efficient extension method requiring fewer tokens and fewer training steps than previous methods. (arXiv)

Conceptually, YaRN treats different RoPE dimensions differently.

Some dimensions encode high-frequency, local information:

nearby token order\text{nearby token order}

Other dimensions encode low-frequency, long-distance information:

broad document-scale position\text{broad document-scale position}

A good long-context PE should not compress both equally. User want to preserve local structure while stretching the long-distance part.

So YaRN-style methods are closer to:

θ~m=frequency-specific scaling(θm)\tilde{\theta}_m = \text{frequency-specific scaling}(\theta_m)

rather than:

i~=i/sfor every dimension equally\tilde{i}=i/s \quad \text{for every dimension equally}

This is one of the major reasons modern long-context RoPE extensions are more than simple linear interpolation.


7. LongRoPE: non-uniform scaling for very large contexts

LongRoPE pushes this idea further.

Instead of using one global scale factor, LongRoPE searches for non-uniform rescaling factors across RoPE dimensions and positions. Its paper reports extending pretrained LLM context windows beyond 2 million tokens, using non-uniform positional interpolation, progressive extension, and short-context readjustment. (arXiv)

The important idea is:

not all dimensions should be stretched the same way\text{not all dimensions should be stretched the same way}

and

not all position ranges behave the same way\text{not all position ranges behave the same way}

Early positions, local spans, mid-range spans, and ultra-long spans may need different treatment.

So compared with simple interpolation:

i~=i/s\tilde{i}=i/s

LongRoPE is more like:

ϕi,m=g(i,m)θm\phi_{i,m}=g(i,m)\theta_m

where g(i,m)g(i,m) depends on both the token position and the RoPE dimension/frequency.

This gives much more control over the model’s positional geometry.


8. Continued pretraining: the PE must be learned in context

Even if we design a good RoPE scaling function, the model still needs to learn how to use long sequences.

A model trained mostly on 4K or 8K samples has not practiced:

retrieve evidence from 600K tokens ago\text{retrieve evidence from 600K tokens ago} ignore irrelevant text across hundreds of documents\text{ignore irrelevant text across hundreds of documents} combine facts from distant sections\text{combine facts from distant sections} avoid attention dilution across massive context\text{avoid attention dilution across massive context}

So long-context extension usually includes continued pretraining or fine-tuning on long sequences.

A 2025 ultra-long-context training paper describes extending Llama-3.1-8B-Instruct from 128K to 1M, 2M, and 4M tokens using continued pretraining, YaRN-based RoPE scaling, full attention, and context parallelism. It reports using RoPE scaling factors of 128, 256, and 512 for 1M, 2M, and 4M target lengths, respectively. (arXiv)

That is very important: the PE formula provides the coordinate system, but training teaches the model what to do inside that coordinate system.


9. ALiBi as another long-context PE direction

RoPE is not the only long-context-friendly PE method.

ALiBi adds a linear distance penalty directly to attention scores:

sij=qi⊤kjd−mh(i−j)s_{ij} = \frac{q_i^\top k_j}{\sqrt d} - m_h(i-j)

where mhm_h is a head-specific slope.

Unlike learned absolute PE, ALiBi does not need a learned position table. Unlike RoPE, it does not rotate Q/K vectors. It simply says:

farther tokens receive a stronger penalty\text{farther tokens receive a stronger penalty}

The ALiBi paper was explicitly about “train short, test long.” It reports training a 1.3B model at length 1024 and extrapolating to length 2048, matching a sinusoidal baseline trained at length 2048 while training faster and using less memory. (arXiv)

ALiBi’s advantage is simplicity and extrapolation. Its disadvantage is that it imposes a strong monotonic recency bias. That can be good for language modeling, but it may be less expressive for tasks where the model must retrieve exact distant information from arbitrary positions.

So:

RoPE=richer relative geometry\text{RoPE} = \text{richer relative geometry} ALiBi=simple distance-decay prior\text{ALiBi} = \text{simple distance-decay prior}

For 1M-token models, RoPE-scaling methods have become especially important because they give more control over positional phase geometry.


10. What happens during generation with 1M context

Suppose the prompt has 1M tokens.

During the prefill stage, the model processes all prompt tokens:

x0,x1,…,x999999x_0,x_1,\ldots,x_{999999}

For each layer, it computes queries, keys, and values. With RoPE, it applies position-dependent rotations to Q and K:

qi′=Ri~qiq_i' = R_{\tilde{i}}q_i ki′=Ri~kik_i' = R_{\tilde{i}}k_i

The model usually stores the rotated keys in the KV cache.

Then when generating token #1,000,001, the new query gets its own position:

q1000000′=R1000000~q1000000q_{1000000}'=R_{\widetilde{1000000}}q_{1000000}

It attends to cached keys:

k0′,k1′,…,k999999′k_0',k_1',\ldots,k_{999999}'

The attention score to each previous token is:

(q1000000′)⊤kj′(q_{1000000}')^\top k_j'

Because both query and key have RoPE applied, the score carries relative positional information:

j~−1000000~\widetilde{j}-\widetilde{1000000}

So the new token can, in principle, distinguish information from nearby tokens, medium-distance tokens, and very distant tokens.

But “in principle” matters. The model also has to learn which distant signals are important, and the attention implementation must handle the enormous compute and memory cost.


11. Why 1M context is not the same as perfect 1M-token reasoning

A 1M-token context window means the model can accept that many tokens. It does not mean every token is used equally well.

Google’s Gemini long-context documentation explicitly notes that performance can vary for cases involving multiple pieces of information to retrieve, even though single-needle retrieval can be very strong. (Google AI for Developers)

This distinction is critical.

A model may pass:

“Find one passkey hidden in 1M tokens.”\text{“Find one passkey hidden in 1M tokens.”}

but still struggle with:

“Compare 25 pieces of evidence spread across 1M tokens.”\text{“Compare 25 pieces of evidence spread across 1M tokens.”}

or:

“Resolve contradictions across many distant documents.”\text{“Resolve contradictions across many distant documents.”}

or:

“Maintain exact chronological state across a huge transcript.”\text{“Maintain exact chronological state across a huge transcript.”}

Long-context PE helps the model locate positions, but reasoning over long context also requires attention allocation, compression, working memory, and training on long-context tasks.


12. Summary

For LLMs with very large context windows, PE is usually handled like this:

MethodHow it supports long contextMain weakness
Learned absolute PELearn a vector for each positionPoor extrapolation; huge table for 1M
Sinusoidal PECompute position features for arbitrary lengthModel may not generalize beyond trained lengths
Relative biasBias attention by distanceLarge distances often bucketed/compressed
ALiBiAdd linear distance penaltyStrong recency bias; less expressive
Standard RoPERotate Q/K by position; attention sees relative offsetRaw long positions can be out of distribution
Position InterpolationCompress long positions into trained RoPE rangeLoses local positional resolution
NTK/RoPE base scalingSlow down RoPE rotationsNeeds careful frequency tradeoff
YaRNFrequency-aware RoPE scalingMore complex, usually needs adaptation
LongRoPENon-uniform dimension/position scalingRequires search/fine-tuning strategy
Long-context pretrainingTeaches model to use long sequencesExpensive but necessary for strong performance

The most important formula is:

qi′=Rf(i)qi,kj′=Rf(j)kjq_i' = R_{f(i)}q_i,\quad k_j'=R_{f(j)}k_j

where f(i)f(i) is no longer always just ii. For short-context RoPE:

f(i)=if(i)=i

For long-context RoPE scaling:

f(i)=isf(i)=\frac{i}{s}

or a more advanced non-uniform function:

f(i,m)=position- and dimension-dependent scalingf(i,m)=\text{position- and dimension-dependent scaling}

That is the heart of long-context PE.

Long-context LLMs usually rely on RoPE or RoPE-like positional systems because they avoid fixed position tables and naturally give relative-position-aware attention. To reach 1M tokens, raw RoPE is usually modified through interpolation, frequency/base scaling, YaRN-style scaling, LongRoPE-style non-uniform scaling, and then adapted with long-context continued pretraining. PE provides the coordinate system, but real 1M-context performance also depends on training data, attention efficiency, KV-cache handling, and long-context evaluation.

References

  1. RoFormer: Enhanced Transformer with Rotary Position Embedding
  2. Round and Round We Go! What Makes Rotary Positional Encodings Useful?
  3. The Impact of Positional Encoding on Length Generalization in Transformers
  4. Self-Attention with Relative Position Representations
  5. Rethinking Positional Encoding in Language Pre-training (TUPE)
  6. DeBERTa: Decoding-enhanced BERT with Disentangled Attention
  7. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows
  8. Do Transformers Really Perform Bad for Graph Representation? (Graphormer)
  9. Comparing Graph Transformers via Positional Encodings
  10. Improve Transformer Models with Better Relative Position Embeddings
  11. Long context — Gemini API documentation
  12. Extending Context Window of Large Language Models via Positional Interpolation
  13. Context Parallelism for Scalable Million-Token Inference
  14. YaRN: Efficient Context Window Extension of Large Language Models
  15. LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
  16. From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
  17. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation