Technical note

Positional Embedding Intermediate

A first advanced to positional embedding/encoding methods.

  • Deep Learning
  • Transformer

Let’s start from a pure content-based attention between token, patch, or node representations:

sij=qi⊤kjds_{ij}=\frac{q_i^\top k_j}{\sqrt d}

where qiq_i represents the querying token at position ii, and kjk_j represents the candidate token at position jj.

Then positional encoding adds another source of information: where things are, either absolutely or relatively.


1. Recap the fundamentals

PE basically increases spatial relationship awareness, either with respect to absolute position or relative distance.

For sequence models, “space” usually means token order:

i, j, i−ji,\ j,\ i-j

For vision models, it may mean 2D patch coordinates:

(rowi,coli), (rowi−rowj, coli−colj)(row_i, col_i),\ (row_i-row_j,\ col_i-col_j)

For graphs, it may mean graph distance, centrality, edge relation, or spectral coordinates.

So PE is fundamentally about giving the model access to order, distance, geometry, or structure that is not already available from content embeddings alone.


2. Mechanism A: additive PE does entangle position with content

Additive PE injects positional information into the token representation itself, so content similarity and positional information become entangled inside the attention computation.

In additive PE, we do:

zi=xi+piz_i = x_i + p_i

Then attention uses projected versions:

qi=ziWQ=(xi+pi)WQq_i = z_i W_Q = (x_i+p_i)W_Q kj=zjWK=(xj+pj)WKk_j = z_j W_K = (x_j+p_j)W_K

So the attention score becomes:

qi⊤kj=(xiWQ+piWQ)⊤(xjWK+pjWK)q_i^\top k_j = (x_iW_Q + p_iW_Q)^\top(x_jW_K + p_jW_K)

Expanding:

qi⊤kj=(xiWQ)⊤(xjWK)⏟content-content+(xiWQ)⊤(pjWK)⏟content-position+(piWQ)⊤(xjWK)⏟position-content+(piWQ)⊤(pjWK)⏟position-positionq_i^\top k_j = \underbrace{(x_iW_Q)^\top(x_jW_K)}_{\text{content-content}} + \underbrace{(x_iW_Q)^\top(p_jW_K)}_{\text{content-position}} + \underbrace{(p_iW_Q)^\top(x_jW_K)}_{\text{position-content}} + \underbrace{(p_iW_Q)^\top(p_jW_K)}_{\text{position-position}}

So additive PE does not merely add position as a separate channel. It mixes content and position before attention.

The so called “noise” is true since the additive PE can make content and position less cleanly separated. But it is not always harmful. The model may learn to use that entanglement productively.


3. Mechanism B: position bias separates spatial relation from content similarity

Position bias cleanly separates positional preference from content representation at the attention-logit level, but the final attention distribution still combines both.

With position bias, attention becomes:

sij=qi⊤kjd+b(i,j)s_{ij} = \frac{q_i^\top k_j}{\sqrt d} + b(i,j)

or, for relative position:

sij=qi⊤kjd+b(i−j)s_{ij} = \frac{q_i^\top k_j}{\sqrt d} + b(i-j)

Here the first term is content-based:

qi⊤kjd\frac{q_i^\top k_j}{\sqrt d}

and the second term is position-based:

b(i,j)b(i,j)

So mechanism B keeps the content similarity and spatial relationship prior more explicitly separated.

For example, if b(i−j)b(i-j) is large when two tokens are nearby, the model is encouraged to attend locally. If b(i−j)b(i-j) is small or negative for faraway tokens, attention to distant tokens is discouraged unless the content score is strong enough.

So the attention score becomes something like:

attention score=content compatibility+spatial compatibility\text{attention score} = \text{content compatibility} + \text{spatial compatibility}

This is a clean interpretation.

However, there is one nuance: after addition, both terms go through the same softmax:

aij=softmaxj(sij)a_{ij} = \text{softmax}_j(s_{ij})

So the final attention weights still combine content and position. Mechanism B separates them before scoring, but the final decision still merges them.

Also, simple scalar biases are less expressive than additive embeddings or RoPE because the bias is often independent of token content.


4. Mechanism C: RoPE applies position as a rotation, but not merely by adding an angle to cosine similarity

Mechanism-C basically applies spatial representation to the tensors. Their angular difference is naturally added to the cosine similarity besides their tensor similarity.

In RoPE, the model rotates query and key vectors according to their positions:

qi′=Riqiq_i' = R_i q_i kj′=Rjkjk_j' = R_j k_j

Then attention uses:

(qi′)⊤kj′=(Riqi)⊤(Rjkj)(q_i')^\top k_j' = (R_iq_i)^\top(R_jk_j) =qi⊤Ri⊤Rjkj= q_i^\top R_i^\top R_j k_j

Because rotations compose nicely,

Ri⊤Rj=Rj−iR_i^\top R_j = R_{j-i}

so:

(qi′)⊤kj′=qi⊤Rj−ikj(q_i')^\top k_j' = q_i^\top R_{j-i} k_j

This is the key idea: although RoPE assigns each position an absolute rotation, the dot product between two rotated vectors depends on their relative positional difference.

It is not exactly “added.” It is more like the content similarity is computed after one vector has been relatively rotated.

A better interpretation is:

RoPE attention=content similarity measured in a relative-position-dependent coordinate system\text{RoPE attention} = \text{content similarity measured in a relative-position-dependent coordinate system}

In a 2D subspace, suppose RoPE rotates by angle θi\theta_i at position ii and θj\theta_j at position jj. The relative angle is:

Δθ=θj−θi\Delta\theta = \theta_j-\theta_i

The dot product becomes something like:

(qi′)⊤kj′=cos⁡(Δθ)(q1k1+q2k2)+sin⁡(Δθ)(q1k2−q2k1)(q_i')^\top k_j' = \cos(\Delta\theta)(q_1k_1+q_2k_2) + \sin(\Delta\theta)(q_1k_2-q_2k_1)

So RoPE does not simply do:

content similarity+angle difference\text{content similarity} + \text{angle difference}

Instead, it does:

rotated content similarity\text{rotated content similarity}

The relative angle changes how the query and key dimensions align with each other.

That makes RoPE more expressive than a simple additive position bias. A bias says:

sij=qi⊤kj+b(i−j)s_{ij}=q_i^\top k_j + b(i-j)

RoPE says:

sij=qi⊤Rj−ikjs_{ij}=q_i^\top R_{j-i} k_j

The difference is important:

b(i−j)b(i-j)

is a content-independent spatial prior.

But:

qi⊤Rj−ikjq_i^\top R_{j-i} k_j

is a content-dependent spatial interaction.

So RoPE does not merely say:

“Tokens at this distance should receive this extra score.”

It says:

“The way token ii compares to token jj depends on their relative displacement.”

That is a deeper coupling between content and position.


5. A clean comparison of the interpretations

MechanismintuitionMore precise version
Additive PEAdds position but may introduce position noisePosition is injected into token vectors, causing content and position to become entangled before attention
Position biasSeparates spatial relationship from tensor attentionAdds an explicit content-independent spatial prior to attention logits
RoPE / rotationAssigns a spatial angle to tensorsRotates query/key vectors by position so relative position appears naturally inside the dot product

The most important distinction is this:

Additive PE:position modifies token representation\textbf{Additive PE:} \quad \text{position modifies token representation} Position bias:position modifies attention score\textbf{Position bias:} \quad \text{position modifies attention score} RoPE:position modifies the geometry of query-key comparison\textbf{RoPE:} \quad \text{position modifies the geometry of query-key comparison}

6. One more nuance: standard attention is not exactly cosine similarity

Technically Transformer attention usually uses scaled dot product, not normalized cosine similarity:

sij=qi⊤kjds_{ij}=\frac{q_i^\top k_j}{\sqrt d}

Cosine similarity would be:

cos⁡(qi,kj)=qi⊤kj∥qi∥∥kj∥\cos(q_i,k_j)= \frac{q_i^\top k_j}{\|q_i\|\|k_j\|}

If qiq_i and kjk_j were normalized, dot product and cosine similarity would be equivalent. But in standard Transformers, they are not necessarily normalized.

So it is okay to think of attention as “similarity,” but more precisely it is learned dot-product compatibility.

Therefore, for RoPE, I would not say:

angular difference is added to cosine similarity.

I would say:

relative angular phase changes the query-key dot product, making attention sensitive to relative position.


7. Final verdict to different PE mechanisms

Positional encoding gives the model spatial, temporal, or structural awareness. Additive PE injects position directly into token representations, which entangles content and position. Position bias keeps content similarity and spatial preference more separate by adding a positional term to the attention logits. RoPE applies a position-dependent rotation to query and key vectors, so their dot product naturally depends on relative position. This is not simply adding an angular term to cosine similarity; rather, RoPE changes the geometry of the similarity computation itself.