Technical note
Positional Embedding Intermediate
A first advanced to positional embedding/encoding methods.
Let’s start from a pure content-based attention between token, patch, or node representations:
where represents the querying token at position , and represents the candidate token at position .
Then positional encoding adds another source of information: where things are, either absolutely or relatively.
1. Recap the fundamentals
PE basically increases spatial relationship awareness, either with respect to absolute position or relative distance.
For sequence models, “space” usually means token order:
For vision models, it may mean 2D patch coordinates:
For graphs, it may mean graph distance, centrality, edge relation, or spectral coordinates.
So PE is fundamentally about giving the model access to order, distance, geometry, or structure that is not already available from content embeddings alone.
2. Mechanism A: additive PE does entangle position with content
Additive PE injects positional information into the token representation itself, so content similarity and positional information become entangled inside the attention computation.
In additive PE, we do:
Then attention uses projected versions:
So the attention score becomes:
Expanding:
So additive PE does not merely add position as a separate channel. It mixes content and position before attention.
The so called “noise” is true since the additive PE can make content and position less cleanly separated. But it is not always harmful. The model may learn to use that entanglement productively.
3. Mechanism B: position bias separates spatial relation from content similarity
Position bias cleanly separates positional preference from content representation at the attention-logit level, but the final attention distribution still combines both.
With position bias, attention becomes:
or, for relative position:
Here the first term is content-based:
and the second term is position-based:
So mechanism B keeps the content similarity and spatial relationship prior more explicitly separated.
For example, if is large when two tokens are nearby, the model is encouraged to attend locally. If is small or negative for faraway tokens, attention to distant tokens is discouraged unless the content score is strong enough.
So the attention score becomes something like:
This is a clean interpretation.
However, there is one nuance: after addition, both terms go through the same softmax:
So the final attention weights still combine content and position. Mechanism B separates them before scoring, but the final decision still merges them.
Also, simple scalar biases are less expressive than additive embeddings or RoPE because the bias is often independent of token content.
4. Mechanism C: RoPE applies position as a rotation, but not merely by adding an angle to cosine similarity
Mechanism-C basically applies spatial representation to the tensors. Their angular difference is naturally added to the cosine similarity besides their tensor similarity.
In RoPE, the model rotates query and key vectors according to their positions:
Then attention uses:
Because rotations compose nicely,
so:
This is the key idea: although RoPE assigns each position an absolute rotation, the dot product between two rotated vectors depends on their relative positional difference.
It is not exactly “added.” It is more like the content similarity is computed after one vector has been relatively rotated.
A better interpretation is:
In a 2D subspace, suppose RoPE rotates by angle at position and at position . The relative angle is:
The dot product becomes something like:
So RoPE does not simply do:
Instead, it does:
The relative angle changes how the query and key dimensions align with each other.
That makes RoPE more expressive than a simple additive position bias. A bias says:
RoPE says:
The difference is important:
is a content-independent spatial prior.
But:
is a content-dependent spatial interaction.
So RoPE does not merely say:
“Tokens at this distance should receive this extra score.”
It says:
“The way token compares to token depends on their relative displacement.”
That is a deeper coupling between content and position.
5. A clean comparison of the interpretations
| Mechanism | intuition | More precise version |
|---|---|---|
| Additive PE | Adds position but may introduce position noise | Position is injected into token vectors, causing content and position to become entangled before attention |
| Position bias | Separates spatial relationship from tensor attention | Adds an explicit content-independent spatial prior to attention logits |
| RoPE / rotation | Assigns a spatial angle to tensors | Rotates query/key vectors by position so relative position appears naturally inside the dot product |
The most important distinction is this:
6. One more nuance: standard attention is not exactly cosine similarity
Technically Transformer attention usually uses scaled dot product, not normalized cosine similarity:
Cosine similarity would be:
If and were normalized, dot product and cosine similarity would be equivalent. But in standard Transformers, they are not necessarily normalized.
So it is okay to think of attention as “similarity,” but more precisely it is learned dot-product compatibility.
Therefore, for RoPE, I would not say:
angular difference is added to cosine similarity.
I would say:
relative angular phase changes the query-key dot product, making attention sensitive to relative position.
7. Final verdict to different PE mechanisms
Positional encoding gives the model spatial, temporal, or structural awareness. Additive PE injects position directly into token representations, which entangles content and position. Position bias keeps content similarity and spatial preference more separate by adding a positional term to the attention logits. RoPE applies a position-dependent rotation to query and key vectors, so their dot product naturally depends on relative position. This is not simply adding an angular term to cosine similarity; rather, RoPE changes the geometry of the similarity computation itself.