Technical note
Positional Embedding Advance - RoPE
Advance to PE technologies: RoPE.
Fundamentals
the full RoPE rotation matrix is high-dimensional. Mnay studies use the 2D rotation matrix as example, which only one small block inside it.
The key idea is:
RoPE rotates a high-dimensional vector by splitting it into many 2D pairs and rotating each pair by a position-dependent angle.
So for a vector with dimension , RoPE does not use one matrix. It uses four rotation blocks inside one big matrix.
1. Start with ordinary 2D rotation
For a 2D vector
a rotation by angle is:
So:
That is the small 2D matrix I showed earlier.
But RoPE applies this idea repeatedly across many pairs of dimensions.
2. High-dimensional RoPE rotation
Suppose my attention head dimension is . A query vector may look like:
RoPE groups it into pairs:
Each pair is rotated separately.
For position , RoPE computes several angles:
Each pair gets its own angle. So the full rotation matrix is:
This is an matrix.
So,
RoPE is high-dimensional, but it is built from many independent 2D rotations.
3. Concrete numerical example
Take a 6-dimensional vector:
Group it into three 2D pairs:
Now suppose this token is at position:
For illustration, use three RoPE frequencies:
Then the angles are:
So the first pair rotates by radians, the second pair by radians, and the third pair by radians.
Now rotate each pair.
First pair
rotated by radians:
Since
we get:
Second pair
rotated by radians:
Since
we get:
Third pair
rotated by radians:
Since
we get:
So the whole vector becomes:
That is how RoPE rotates a high-dimensional vector.
It does not rotate the whole vector with one simple 2D angle. Instead, it rotates different 2D slices of the vector by different angles.
4. Why different dimension pairs use different angles
In real RoPE, the frequencies are usually chosen like this:
where is the pair index.
For example, if , then there are four pairs:
The angles are:
So early dimension pairs rotate quickly, while later dimension pairs rotate slowly.
This is similar to sinusoidal positional encoding. Some dimensions represent fine-grained local position, while others represent broad long-range position.
I can think of the dimension pairs as different clocks:
So when the position changes, each pair rotates by a different amount.
5. Why this produces relative position information
Now suppose we have query vector at position , and key vector at position .
RoPE rotates them:
Then attention computes:
Substitute the rotated forms:
For rotation matrices:
So:
This is the beautiful part.
Even though RoPE rotates each vector using its absolute position, the final query-key dot product depends on the relative position:
So RoPE turns absolute position rotations into relative-position-aware attention.
6. A useful analogy
Imagine every token vector lives in many tiny 2D planes:
At position , RoPE rotates the vector in each plane.
But each plane rotates at a different speed.
So position is encoded not as an added vector, but as a set of phases across many 2D planes.
That is why RoPE is called rotary positional embedding.
7. Compact interpretation
For a high-dimensional vector, RoPE does this:
becomes:
So the 2D rotation matrix is just one repeated building block.
The full RoPE rotation is a block-diagonal high-dimensional rotation matrix:
where each block is:
That is the cleanest mental model:
RoPE rotates a high-dimensional vector by performing many small 2D rotations on pairs of dimensions. Each pair rotates at a different frequency, and the rotation angle depends on the token position.
Math behind RoPE
RoPE is not just a random intuition validated by experiments. It has a clean mathematical motivation. But the claim that “RoPE is better for language models” is largely empirical and mechanistic.
So RoPE has theory for what it encodes, but not a complete theory for why it always improves model performance.
1. The theoretical goal of RoPE
The original RoPE paper starts from a desired property:
should depend on the relative position
rather than depending independently on and .
In other words, we want:
This is a formal design target, not merely intuition. The RoFormer paper explicitly formulates RoPE as a way to encode absolute position through rotation while making the self-attention score depend on relative position. (arXiv)
Using the rotation matrices introduced in the Fundamentals chapter, RoPE encodes the query and key as and . The derivation above then gives:
This is the central theoretical justification: RoPE converts absolute rotations at positions and into a relative rotation depending on .
2. Why rotation is mathematically natural here
Positions in a sequence behave like translations:
Rotations also compose naturally:
So if representing position by a rotation , then comparing two positions automatically gives:
This is why RoPE is elegant: relative position appears automatically from the algebra of rotations.
In more abstract language, RoPE uses a representation of the additive position group through rotations. In simpler language:
moving forward in sequence position corresponds to rotating in hidden-vector space.
That is a real theoretical structure.
3. The connection to complex numbers and Fourier features
There is another theoretical background: RoPE is closely related to complex exponentials.
A 2D vector
can be written as a complex number:
Rotating by angle is the same as multiplying by:
So RoPE can be viewed as:
For two positions and :
Again, the relative distance appears naturally.
This is why RoPE is related to sinusoidal positional encoding. Both use sine/cosine waves or complex phases. The difference is that sinusoidal PE is usually added to token embeddings, while RoPE multiplies/rotates query and key vectors. The RoFormer paper specifically notes that RoPE incorporates relative position information by multiplying with sinusoidal functions rather than directly adding position to contextual representations. (arXiv)
4. The high-dimensional theory
For a high-dimensional vector, RoPE uses many 2D rotation planes:
Each frequency gives a different positional scale.
So the full attention score is roughly:
This gives the model many “relative-distance channels”:
and so on.
This is another reason RoPE is not arbitrary. It is a multiscale Fourier-like positional system.
5. What the original paper theoretically claims
The original RoPE paper claims several mathematical/structural advantages:
First, RoPE encodes absolute position using a rotation matrix while making attention explicitly sensitive to relative position. Second, because it does not require a learned lookup table for each position, it is flexible with respect to sequence length. Third, the authors analyze a long-term decay property: under their frequency setting, certain RoPE inner-product terms decay as relative distance increases. Fourth, they argue that RoPE can be used with linear self-attention more naturally than some previous additive relative-position methods. (arXiv)
So yes, the original paper did not simply say:
“Let’s rotate vectors; maybe it works.”
It said something closer to:
“We want relative-position-dependent attention scores. Rotation gives us that property exactly. Then we validate that this useful inductive bias improves performance.”
6. But the performance advantage is not fully theoretically explained
The theory explains what RoPE does geometrically. It does not fully explain why RoPE makes a large Transformer better at language modeling, reasoning, or long-context understanding.
Even the RoFormer authors acknowledge limitations. They state that although they provide theoretical grounding and experiments, they do not have a thorough explanation for why RoPE converges faster than baseline positional encoding methods, nor a faithful explanation for why it performs better on long texts than peer models. (arXiv)
RoPE is partly a theoretically motivated design, and partly an empirical success.
7. A modern nuance: the “distance decay” explanation is incomplete
One common explanation is:
RoPE works because attention naturally decays with distance.
That is only partly true.
Later analysis argues that this decay explanation is too simple. A 2024 paper, Round and Round We Go! What makes Rotary Positional Encodings useful?, argues that RoPE does not necessarily make activations decay with distance in general. For example, they show that decay may appear under special query/key assumptions, but not necessarily for random Gaussian query/key vectors. They also show that RoPE can help construct attention heads that attend to specific relative positions. (arXiv)
This is important.
RoPE should not be understood as merely:
A better understanding (more precise) is:
8. Length extrapolation is also not fully solved by RoPE
RoPE is often associated with better length extrapolation than learned absolute position embeddings, but it does not magically solve length extrapolation.
A NeurIPS 2023 study comparing positional encodings found that the exact impact of different PE schemes on extrapolation remains unclear, and in their downstream length-generalization tasks, common explicit PE methods including Rotary, ALiBi, and absolute position embeddings were not always well suited; their NoPE baseline performed surprisingly well. (NeurIPS Proceedings)
So we should avoid saying:
RoPE theoretically guarantees long-context generalization.
It does not.
A more accurate statement is:
RoPE has better structural properties than learned absolute position tables for variable-length use, but long-context generalization still depends on training distribution, frequency scale, architecture, data, and fine-tuning.
Other absolute PE + relative PE ideas/concepts
For non-RoPE methods, the usual way to combine absolute PE and relative PE is to let them play different roles:
The central design choice is whether to mix them inside token representations or keep them disentangled as separate attention-score terms.
One nuance first: RoPE is often described as encoding both absolute and relative position, but standard RoPE is better described as absolute-indexed in construction and relative in the attention logit. The query/key vector at position is rotated by an absolute phase , but when attention computes , the absolute rotations combine into . So the pairwise attention score mainly sees relative offset , not an independent absolute-position label.
1. Naive combination: add absolute PE to tokens, add relative PE to attention
This is the simplest combination.
Let
Then use:
and attention score:
So the model receives absolute position through , and relative position through .
This is conceptually simple:
The disadvantage is that the absolute position vector is mixed into the token vector before attention. Expanding the dot product shows the entanglement:
So this approach combines absolute and relative position, but it also creates mixed content-position interactions that may or may not be useful.
This method is common because it is easy: for example, I can take a BERT/ViT-style model with learned absolute embeddings and add a relative position bias term to the attention logits.
2. Relative key/value combination: absolute PE at input, relative PE inside attention
A more expressive version uses relative vectors, not just scalar biases.
A Shaw-style relative attention score can look like:
where is a learned relative-position key vector.
The output can also use a relative value vector:
Then we can still add absolute PE to the input:
So the full system has:
This is more expressive than scalar relative bias because the relative position can interact with the query vector. Shaw et al. introduced relative position representations for self-attention and found that relative position improved machine translation over absolute position in their experiments; interestingly, they also reported that combining relative and absolute position gave no further translation improvement in that setting. (arXiv)
So the lesson is important:
Combining absolute and relative PE is possible, but it is not automatically better.
Sometimes relative position already contains the useful ordering signal.
3. Disentangled combination: keep content, absolute position, and relative position separate
This is the most conceptually clean approach.
Instead of first doing:
we keep content and position as different objects:
Then define attention score as a sum of separate components:
For example:
So:
This gives a clear separation:
This idea is close to TUPE, “Transformer with Untied Positional Encoding.” TUPE argues that simply adding word embeddings and position embeddings creates mixed correlations between heterogeneous sources of information; it instead computes word contextual correlation and positional correlation separately, then adds them together in attention. (OpenReview)
This is a strong design pattern when concerning: additive absolute PE can act like “position noise” inside content attention. Disentangling lets the model use position without forcing content and position to live in the same representation channel.
4. DeBERTa-style combination: relative position in attention, absolute position near decoding/output
DeBERTa gives a very instructive example.
BERT-style models usually do:
DeBERTa instead represents each word using separate content and position vectors. Its disentangled attention computes attention using content and relative-position information, rather than simply mixing word and position embeddings at the input. The DeBERTa paper also notes that relative information alone may miss absolute-position cues, so it incorporates absolute position information in the enhanced mask decoder used for masked-token prediction. (ar5iv)
Conceptually, DeBERTa separates the jobs:
This is a very elegant pattern.
Why? Because during attention, relative position is often what matters:
“Is this word next to that word?” “Is this patch above or below that patch?” “Is this token far away or nearby?”
But during output prediction, absolute position can matter:
“Is this word near the beginning of the sentence?” “Is this token in the title region?” “Is this patch in the top-left of the image?” “Is this timestep early or late in the sequence?”
So DeBERTa-style design says:
Use relative position for pairwise interaction, but keep absolute position available when the model needs to make position-specific predictions.
5. Head-wise or layer-wise combination
Another practical way is to let different heads or layers specialize.
For attention head , use:
Some heads can receive relative bias; other heads can use absolute PE; some layers may use both.
For example:
This can help because positional needs are not uniform across the network. Early layers may need local ordering, while later layers may need global layout or beginning/end information.
This idea is especially natural in vision and document models, where some heads should learn local geometry while others learn global layout.
6. 2D combination for vision: absolute coordinates plus relative offsets
For images, absolute and relative position mean slightly different things.
Absolute position:
Relative position:
A combined 2D PE can be:
where
This gives the model both:
and
But again, combining is not always better. Swin Transformer’s ablation is a useful example. In Swin-T, relative position bias outperformed no position encoding and absolute position embedding, and adding absolute position to relative position was close but did not beat the default relative-only setting on the reported classification/detection/segmentation metrics. The authors also noted that absolute position improved ImageNet classification slightly but hurt COCO detection and ADE20K segmentation, suggesting that relative/local translation-friendly bias is especially valuable for dense vision tasks. (ar5iv)
So in vision:
Absolute PE can help classification or layout-sensitive tasks. Relative PE is often better for detection, segmentation, and local geometric reasoning.
7. Graph combination: node-level absolute PE plus pairwise relative PE
Graphs are another case where combining absolute and relative PE can be very useful.
For a graph Transformer, absolute PE might be a node feature:
or
Relative PE might be a pairwise feature:
or an edge/path encoding.
Then use:
This is very natural:
Graphormer is an important example of this philosophy: its key idea is that Transformers need graph structural information encoded into the model, and it proposes structural encodings to help model graph-structured data. (arXiv) A later graph-Transformer PE comparison frames the distinction clearly: absolute PEs assign features to each node and are given as input to the Transformer, while relative PEs assign features to pairs of nodes, such as graph distance, and augment the attention block. (arXiv)
8. When combined absolute + relative PE helps
Combined PE helps when the task needs both of these questions answered:
Case A: The task depends on global position
Absolute PE helps when the same local pattern means different things in different global positions.
Examples:
Relative PE alone can say:
but it may not directly say:
The DeBERTa example above uses exactly this division of labor: relative position guides attention, while absolute position remains available near prediction.
Case B: The task also needs local relational geometry
Relative PE helps when pairwise spatial structure matters.
Examples:
Relative PE is usually better than absolute PE for saying:
because relative distance is directly represented rather than needing to be inferred from two absolute indices.
Case C: I need both global anchors and translation-like invariance
This is common in vision.
For object detection, I want the model to recognize a dog whether it appears on the left or right side of the image. That favors relative/local position and some translation invariance. But for document understanding, satellite imagery, medical scans, or UI screenshots, absolute location can also matter. The Swin results above illustrate this trade-off between translation-friendly relative bias and global position cues.
Case D: I want more expressive attention
A pure scalar relative bias says:
This says:
“At this distance, attention should be easier or harder.”
But richer combined methods can say:
“This type of token should attend to that type of token differently depending on both absolute location and relative displacement.”
That is why some methods increase interaction between query, key, and relative position embeddings rather than using only a simple bias. Huang et al. argued that position information was not fully used in existing methods and proposed techniques to encourage stronger interaction between query, key, and relative position embeddings, reporting improved SQuAD1.1 results compared with previous absolute and relative PE approaches. (ACL Anthology)
9. When combined PE may not help
Combined PE can hurt or fail to help when absolute position is redundant, noisy, or harmful.
The Shaw translation and Swin vision results discussed above provide two examples: an added absolute signal may contribute nothing when relative order is sufficient, and it may weaken translation-friendly behavior in dense vision tasks.
Combined PE may also hurt length or resolution extrapolation. Learned absolute PE often binds the model to positions seen during training, while relative PE usually generalizes more naturally to shifted patterns. So if the model must extrapolate to much longer sequences or different image sizes, adding a strong learned absolute embedding can become a liability.
10. Practical rule
A good practical design is:
More specifically:
| Setting | Good PE combination |
|---|---|
| Text encoder, fixed length | learned absolute PE + relative bias can work |
| Text encoder with strong syntax/QA needs | disentangled content/absolute/relative PE |
| Decoder-only LLM | RoPE or ALiBi usually enough; extra absolute PE can hurt extrapolation |
| Masked language model | relative attention + absolute position near decoder/output can help |
| Image classification | absolute 2D PE + relative bias may help |
| Detection/segmentation | relative 2D bias often better than strong absolute PE |
| Document/layout model | absolute page coordinates + relative local offsets |
| Graph Transformer | node-level structural PE + pairwise distance/edge bias |
| Time series | absolute time/calendar features + relative lag bias |
The cleanest conceptual formula is:
This is better than blindly doing:
when I care about separating content from position.
So the final answer is:
To combine absolute and relative PE, either add absolute PE to token embeddings and add relative PE to attention, or more cleanly, compute content, absolute-position, and relative-position terms separately and add them at the attention-logit level. Combined PE helps when both global location and pairwise geometry matter, but it is not universally better; in tasks where translation invariance or length extrapolation is important, relative-only PE can be superior.
RoPE in long context
In modern long-context LLMs, positional encoding is one of the central technologies that makes 128K, 1M, or even multi-million-token context possible, but it is only one part of the solution.
A good way to think about it:
PE gives the model an address system for token positions. But PE alone does not guarantee that the model can actually retrieve, reason over, or prioritize information across 1M tokens.
Google’s Gemini API documentation says many Gemini models support context windows of 1 million or more tokens, and the Gemini 1.5 report describes models capable of recalling and reasoning over fine-grained information from millions of tokens of context. (Google AI for Developers)
1. Long context is hard for PE
For a normal decoder-only LLM, each token has a position:
If , the PE system must give meaningful positional information for token 0, token 500,000, token 999,999, and every relative distance between them.
That creates several problems.
First, learned absolute PE does not scale well. A model trained with a learned table for 4K or 8K positions does not automatically have trained embeddings for position 900K. User could make a huge learned table, but it would be expensive, poorly extrapolative, and very data-hungry.
Second, plain RoPE extrapolation can become unstable. RoPE can mathematically compute rotations for any position, but if the model was trained only up to 4K or 8K, positions like 500K are far outside the phase range it learned to use. The Position Interpolation paper argues that directly extrapolating beyond the trained context can produce extreme attention behavior, while interpolating positions back into the trained range is much more stable. (arXiv)
Third, long context is not only a PE problem. Even with perfect PE, full attention over 1M tokens is extremely expensive. The practical system also needs efficient attention kernels, distributed context parallelism, KV-cache management, and training examples that actually teach the model to use long evidence. A 2024 context-parallelism paper, for example, studies scalable million-token inference and reports 1M-context prefill experiments on large models using many GPUs. (arXiv)
2. RoPE became popular for long context
For standard RoPE, the relative-position result derived earlier applies directly:
So RoPE gives a very useful property:
This is why RoPE is attractive for LLMs. It does not need a learned table for every position, and the query-key comparison naturally depends on relative offset. The original RoPE paper describes this as encoding absolute position with a rotation matrix while incorporating explicit relative-position dependency in self-attention. (arXiv)
But standard RoPE still has a problem at 1M context: the rotation angles can go far outside the training distribution.
For each dimension pair, RoPE uses something like:
where is the token position and is the frequency for dimension pair .
At position , the angle
may be much larger than anything the model saw during training. So long-context LLMs usually do not just “turn on RoPE to 1M.” They modify how RoPE maps positions to angles.
3. Main strategy: RoPE scaling
The core idea of RoPE scaling is simple:
Instead, use an effective position:
Then RoPE becomes:
The attention comparison becomes:
So the model still gets relative position information, but now relative position is measured in a rescaled coordinate system.
4. Simple position interpolation
Suppose a model was trained with context length:
and we want to run it at:
The scale factor is:
Position interpolation uses:
So token position 1,000,000 becomes:
That means token #1,000,000 is represented using a RoPE phase similar to a token near position 7,812 in the original model’s positional space.
This is the central trick:
The Position Interpolation paper showed that RoPE-based LLaMA models could be extended to 32K context with limited fine-tuning by down-scaling position indices rather than directly extrapolating them. (arXiv)
But there is a cost.
A real distance of 128 tokens becomes:
in RoPE space.
A real distance of 1024 tokens becomes:
in RoPE space.
So the model’s positional ruler becomes compressed. This helps avoid out-of-distribution high positions, but it can reduce the model’s ability to distinguish nearby positions precisely. That is why simple interpolation is useful but not ideal for very long contexts.
5. NTK-aware scaling: change the RoPE frequency scale
Another strategy is to modify the RoPE frequency base. Instead of only shrinking positions:
we change the frequency schedule:
If we increase the base, many RoPE dimensions rotate more slowly. Slower rotations mean longer usable wavelengths, which helps the model represent longer contexts.
Intuitively:
But this creates a tradeoff. If we slow all frequencies too much, the model loses fine local position resolution. If we do not slow them enough, far positions remain out of distribution.
That is the central tension in long-context PE:
6. YaRN: scale different frequency bands more carefully
YaRN, short for “Yet another RoPE extensioN method,” improves on simple interpolation. Its goal is to extend context length efficiently without requiring as much long-context training as naive fine-tuning. The YaRN paper says RoPE-based models fail to generalize far past the sequence length they were trained on, and proposes a compute-efficient extension method requiring fewer tokens and fewer training steps than previous methods. (arXiv)
Conceptually, YaRN treats different RoPE dimensions differently.
Some dimensions encode high-frequency, local information:
Other dimensions encode low-frequency, long-distance information:
A good long-context PE should not compress both equally. User want to preserve local structure while stretching the long-distance part.
So YaRN-style methods are closer to:
rather than:
This is one of the major reasons modern long-context RoPE extensions are more than simple linear interpolation.
7. LongRoPE: non-uniform scaling for very large contexts
LongRoPE pushes this idea further.
Instead of using one global scale factor, LongRoPE searches for non-uniform rescaling factors across RoPE dimensions and positions. Its paper reports extending pretrained LLM context windows beyond 2 million tokens, using non-uniform positional interpolation, progressive extension, and short-context readjustment. (arXiv)
The important idea is:
and
Early positions, local spans, mid-range spans, and ultra-long spans may need different treatment.
So compared with simple interpolation:
LongRoPE is more like:
where depends on both the token position and the RoPE dimension/frequency.
This gives much more control over the model’s positional geometry.
8. Continued pretraining: the PE must be learned in context
Even if we design a good RoPE scaling function, the model still needs to learn how to use long sequences.
A model trained mostly on 4K or 8K samples has not practiced:
So long-context extension usually includes continued pretraining or fine-tuning on long sequences.
A 2025 ultra-long-context training paper describes extending Llama-3.1-8B-Instruct from 128K to 1M, 2M, and 4M tokens using continued pretraining, YaRN-based RoPE scaling, full attention, and context parallelism. It reports using RoPE scaling factors of 128, 256, and 512 for 1M, 2M, and 4M target lengths, respectively. (arXiv)
That is very important: the PE formula provides the coordinate system, but training teaches the model what to do inside that coordinate system.
9. ALiBi as another long-context PE direction
RoPE is not the only long-context-friendly PE method.
ALiBi adds a linear distance penalty directly to attention scores:
where is a head-specific slope.
Unlike learned absolute PE, ALiBi does not need a learned position table. Unlike RoPE, it does not rotate Q/K vectors. It simply says:
The ALiBi paper was explicitly about “train short, test long.” It reports training a 1.3B model at length 1024 and extrapolating to length 2048, matching a sinusoidal baseline trained at length 2048 while training faster and using less memory. (arXiv)
ALiBi’s advantage is simplicity and extrapolation. Its disadvantage is that it imposes a strong monotonic recency bias. That can be good for language modeling, but it may be less expressive for tasks where the model must retrieve exact distant information from arbitrary positions.
So:
For 1M-token models, RoPE-scaling methods have become especially important because they give more control over positional phase geometry.
10. What happens during generation with 1M context
Suppose the prompt has 1M tokens.
During the prefill stage, the model processes all prompt tokens:
For each layer, it computes queries, keys, and values. With RoPE, it applies position-dependent rotations to Q and K:
The model usually stores the rotated keys in the KV cache.
Then when generating token #1,000,001, the new query gets its own position:
It attends to cached keys:
The attention score to each previous token is:
Because both query and key have RoPE applied, the score carries relative positional information:
So the new token can, in principle, distinguish information from nearby tokens, medium-distance tokens, and very distant tokens.
But “in principle” matters. The model also has to learn which distant signals are important, and the attention implementation must handle the enormous compute and memory cost.
11. Why 1M context is not the same as perfect 1M-token reasoning
A 1M-token context window means the model can accept that many tokens. It does not mean every token is used equally well.
Google’s Gemini long-context documentation explicitly notes that performance can vary for cases involving multiple pieces of information to retrieve, even though single-needle retrieval can be very strong. (Google AI for Developers)
This distinction is critical.
A model may pass:
but still struggle with:
or:
or:
Long-context PE helps the model locate positions, but reasoning over long context also requires attention allocation, compression, working memory, and training on long-context tasks.
12. Summary
For LLMs with very large context windows, PE is usually handled like this:
| Method | How it supports long context | Main weakness |
|---|---|---|
| Learned absolute PE | Learn a vector for each position | Poor extrapolation; huge table for 1M |
| Sinusoidal PE | Compute position features for arbitrary length | Model may not generalize beyond trained lengths |
| Relative bias | Bias attention by distance | Large distances often bucketed/compressed |
| ALiBi | Add linear distance penalty | Strong recency bias; less expressive |
| Standard RoPE | Rotate Q/K by position; attention sees relative offset | Raw long positions can be out of distribution |
| Position Interpolation | Compress long positions into trained RoPE range | Loses local positional resolution |
| NTK/RoPE base scaling | Slow down RoPE rotations | Needs careful frequency tradeoff |
| YaRN | Frequency-aware RoPE scaling | More complex, usually needs adaptation |
| LongRoPE | Non-uniform dimension/position scaling | Requires search/fine-tuning strategy |
| Long-context pretraining | Teaches model to use long sequences | Expensive but necessary for strong performance |
The most important formula is:
where is no longer always just . For short-context RoPE:
For long-context RoPE scaling:
or a more advanced non-uniform function:
That is the heart of long-context PE.
Long-context LLMs usually rely on RoPE or RoPE-like positional systems because they avoid fixed position tables and naturally give relative-position-aware attention. To reach 1M tokens, raw RoPE is usually modified through interpolation, frequency/base scaling, YaRN-style scaling, LongRoPE-style non-uniform scaling, and then adapted with long-context continued pretraining. PE provides the coordinate system, but real 1M-context performance also depends on training data, attention efficiency, KV-cache handling, and long-context evaluation.
References
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Round and Round We Go! What Makes Rotary Positional Encodings Useful?
- The Impact of Positional Encoding on Length Generalization in Transformers
- Self-Attention with Relative Position Representations
- Rethinking Positional Encoding in Language Pre-training (TUPE)
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows
- Do Transformers Really Perform Bad for Graph Representation? (Graphormer)
- Comparing Graph Transformers via Positional Encodings
- Improve Transformer Models with Better Relative Position Embeddings
- Long context — Gemini API documentation
- Extending Context Window of Large Language Models via Positional Interpolation
- Context Parallelism for Scalable Million-Token Inference
- YaRN: Efficient Context Window Extension of Large Language Models
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
- From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation