Technical note

Positional Embedding Fundamentals

An introduction to positional embedding/encoding methods.

  • Deep Learning
  • Transformer

The terms of positional embedding and positional encoding are often used interchangeably, but there is a useful distinction:

Positional encoding usually means a fixed, hand-designed function of position, such as sinusoidal encoding.

Positional embedding usually means a learned lookup table:

pi=EmbeddingTable[i]p_i = \text{EmbeddingTable}[i]

In practice, “PE” often refers to both.

The important design questions are:

QuestionPossible answers
What position information is represented?Absolute position, relative distance, 2D coordinates, graph distance, time step
Where is it injected?Input embeddings, attention logits, queries/keys, values, convolution channels
Is it learned or fixed?Learned lookup, fixed sinusoidal, fixed linear bias, learned relative bias
Is it 1D, 2D, 3D, or structural?Text sequence, image grid, video volume, graph topology

1. Why positional encoding is needed

A Transformer’s self-attention layer is powerful because every token can compare itself with every other token. But this creates a problem: plain self-attention does not naturally know sequence order.

Suppose the model receives:

“dog bites man” “man bites dog”

The token set is the same: {dog, bites, man}. But the meaning is completely different. A model that only sees tokens without positions has no built-in way to distinguish those two orders.

Mathematically, a basic self-attention layer computes

Q=XWQ,K=XWK,V=XWVQ=XW_Q,\quad K=XW_K,\quad V=XW_V Attention(X)=softmax(QKTd)V\text{Attention}(X)=\text{softmax}\left(\frac{QK^T}{\sqrt{d}}\right)V

If we permute the input tokens, the output is permuted in the same way. In other words, full self-attention is permutation-equivariant: it treats the sequence as a set unless we inject order. The original Transformer removed recurrence and convolution, so the authors explicitly added positional encodings to the input embeddings to give the model information about token order.

Think of positional encoding as an address system. Token embeddings say what a token is. Positional encodings say where it is.

inputi=token embeddingi+position encodingi\text{input}_i = \text{token embedding}_i + \text{position encoding}_i

Without that address system, the model can learn “dog,” “bites,” and “man,” but not reliably “who came before whom.”


2. General mechanisms of PE

Mechanism A: Add position vectors to token vectors

This is the simplest mechanism:

zi=xi+piz_i = x_i + p_i

Here, (x_i) is the token embedding and (p_i) is the position vector for index (i). This is used by original Transformers, BERT-like encoders, and Vision Transformers.

The advantage is simplicity. The disadvantage is that the model must learn from the combined vector which part is content and which part is position.


Mechanism B: Add position bias to attention scores

Instead of modifying token vectors, we modify the attention score between token (i) and token (j):

sij=qiTkjd+b(i,j)s_{ij} = \frac{q_i^T k_j}{\sqrt{d}} + b(i,j)

Here, (b(i,j)) can depend on relative distance (i-j), 2D patch offset, graph shortest-path distance, or another structural relation.

This is very common in relative PE methods, T5-style relative position bias, ALiBi, and Swin Transformer. T5 uses a simplified relative position scheme in which a scalar bias is added to the attention logit; the T5 paper describes using 32 relative-position buckets whose ranges grow logarithmically.


Mechanism C: Rotate or transform query/key vectors

RoPE, or Rotary Position Embedding, injects position by rotating the query and key vectors as a function of their absolute positions:

qi′=Riqi,kj′=Rjkjq_i' = R_i q_i,\quad k_j' = R_j k_j

Then

(qi′)T(kj′)=qiTRiTRjkj(q_i')^T(k_j') = q_i^T R_i^T R_j k_j

Because rotations compose as relative rotations, (R_i^T R_j) depends on (j-i). That means RoPE encodes absolute positions through rotations while making attention naturally sensitive to relative distances. The RoFormer paper explicitly describes RoPE as encoding absolute position with a rotation matrix while incorporating relative position dependency in self-attention.


Mechanism D: Use coordinates or Fourier features

For images, neural fields, NeRF-like models, audio coordinates, and continuous signals, we often encode real-valued coordinates:

x↦[sin⁡(20πx),cos⁡(20πx),…,sin⁡(2Lπx),cos⁡(2Lπx)]x \mapsto [\sin(2^0\pi x), \cos(2^0\pi x), \ldots, \sin(2^L\pi x), \cos(2^L\pi x)]

This gives an MLP access to high-frequency information. Fourier feature work showed that mapping low-dimensional coordinates through Fourier features helps MLPs learn high-frequency functions that standard MLPs struggle to represent.


Mechanism E: Use graph/topological structure

Graphs do not have a natural left-to-right order. For graph Transformers, “position” often means structural identity: degree, shortest-path distance, centrality, Laplacian eigenvectors, random-walk features, or edge relations.

Graphormer argues that structural encoding is necessary for Transformers to work well on graph representation learning, and introduces structural encodings such as centrality and spatial encodings. Laplacian eigenvector positional encodings are another graph method, proposed as a graph generalization of sinusoidal PEs.


3. The most important PE methods

3.1 Learned absolute positional embeddings

This is the simplest learned method:

pi∈Rdp_i \in \mathbb{R}^d

Each position has its own trainable vector. The input becomes

zi=xi+piz_i = x_i + p_i

Strengths: easy, effective, flexible for fixed maximum length, widely used in BERT-style encoders and early ViTs.

Weaknesses: weak length extrapolation. If the model was trained up to position 512, position 2000 has no learned vector unless user extend or interpolate the table.

Best fit: fixed-length text encoders, classification models, ViT image classifiers at fixed or interpolated resolutions.

Vision Transformer uses learnable 1D position embeddings added to patch embeddings; the authors also note that the 2D image structure is only weakly injected, mainly through patch extraction and position embedding interpolation during resolution changes. (ar5iv)


3.2 Sinusoidal absolute positional encoding

The original Transformer used sine and cosine functions:

PEpos,2i=sin⁡(pos100002i/d)PE_{pos,2i}=\sin\left(\frac{pos}{10000^{2i/d}}\right) PEpos,2i+1=cos⁡(pos100002i/d)PE_{pos,2i+1}=\cos\left(\frac{pos}{10000^{2i/d}}\right)

Each dimension oscillates at a different frequency. Low-frequency dimensions encode broad position; high-frequency dimensions encode fine local changes.

Strengths: no learned parameters, can be computed for arbitrary sequence length, gives each position a unique multiscale pattern.

Weaknesses: extrapolation is possible mathematically, but not guaranteed behaviorally. The model still learns on a finite length distribution.

Best fit: classic Transformer baselines, small/medium sequence models, cases where parameter-free PE is preferred.

The original Transformer paper added positional encodings to input embeddings and found learned positional embeddings gave nearly identical results in their translation experiment. (arXiv)

A useful intuition: sinusoidal PE is a multiscale clock. One dimension ticks quickly, another slowly, another very slowly. Combining them lets the model infer both “where am I?” and “how far am I from another token?”


3.3 Relative position representations

Absolute PE says:

“I am token 17.”

Relative PE says:

“I am 3 tokens after you.”

In relative attention, the attention interaction between (i) and (j) depends on (i-j). Shaw et al. introduced relative position representations for self-attention, efficiently modeling distances between sequence elements, and reported improvements over absolute position representations on WMT translation tasks. (ACL Anthology)

A simplified form is:

sij=qiT(kj+ri−j)ds_{ij} = \frac{q_i^T(k_j + r_{i-j})}{\sqrt{d}}

where (r_{i-j}) is a learned embedding for the relative offset.

Strengths: captures local order and distance naturally; more translation-invariant; good for language, speech, music, and local dependencies.

Weaknesses: can be more expensive than scalar biases; implementation can be tricky; relative-only methods may lose useful absolute information such as beginning/end of document.

Best fit: encoder-decoder Transformers, machine translation, long-sequence encoders, tasks where relative distance matters more than exact absolute index.


3.4 T5-style relative position bias

T5 simplifies relative PE by adding a scalar bias to the attention score:

sij=qiTkjd+bbucket(i−j)s_{ij}=\frac{q_i^T k_j}{\sqrt{d}} + b_{\text{bucket}(i-j)}

Instead of learning a vector for every distance, T5 maps distances into buckets. Nearby distances get fine-grained buckets; far distances are grouped logarithmically.

Strengths: cheap, stable, easy to share across layers or heads, good for encoder-decoder NLP.

Weaknesses: less expressive than vector relative embeddings; bucket design matters; far positions may be compressed too aggressively.

Best fit: text-to-text encoder-decoder models, robust NLP systems, moderate-to-long contexts.

T5 explicitly describes self-attention as order-independent and uses relative position scalar biases added to attention logits, with logarithmically increasing relative-position ranges.


3.5 Transformer-XL relative positional encoding

Transformer-XL was designed to handle long context by reusing hidden states from previous segments. Absolute position embeddings create trouble here because position numbers reset or shift between segments. Transformer-XL instead uses a relative positional scheme plus segment-level recurrence, allowing dependency beyond a fixed segment length and reducing context fragmentation. (arXiv)

Strengths: strong for long-context language modeling; position remains meaningful across cached segments.

Weaknesses: more complex than standard attention; less common in modern decoder-only LLM stacks than RoPE-style methods.

Best fit: recurrent/cached Transformers, long document modeling, streaming text.


3.6 RoPE: Rotary Position Embedding

RoPE is one of the most influential modern PE methods for decoder-only language models.

Instead of adding (p_i) to (x_i), RoPE rotates pairs of query/key dimensions:

[q2m′q2m+1′]=[cos⁡(iθm)−sin⁡(iθm)sin⁡(iθm)cos⁡(iθm)][q2mq2m+1]\begin{bmatrix} q_{2m}' \\ q_{2m+1}' \end{bmatrix} = \begin{bmatrix} \cos(i\theta_m) & -\sin(i\theta_m) \\ \sin(i\theta_m) & \cos(i\theta_m) \end{bmatrix} \begin{bmatrix} q_{2m} \\ q_{2m+1} \end{bmatrix}

Same for keys.

Why it works: the dot product between a rotated query at position (i) and a rotated key at position (j) naturally depends on (j-i). This gives relative-position awareness while using absolute position indices internally.

Strengths: elegant; no learned position table; works naturally with attention; strong for autoregressive LLMs; supports flexible sequence lengths better than learned absolute embeddings.

Weaknesses: long-context extrapolation is not automatically solved; many systems need RoPE scaling, interpolation, NTK-aware scaling, YaRN-like methods, or fine-tuning for much longer contexts. It can also introduce frequency/phase issues at very long positions.

Best fit: decoder-only LLMs, long-context language models, multimodal Transformers with text/image/video positions.

The RoPE paper highlights that RoPE encodes absolute position by rotation while incorporating explicit relative dependency, and reports favorable results on long-text classification benchmarks. (arXiv)


3.7 ALiBi: Attention with Linear Biases

ALiBi does not add positional embeddings. It modifies attention scores with a distance penalty:

sij=qiTkjd−mh(i−j)s_{ij}=\frac{q_i^T k_j}{\sqrt{d}} - m_h(i-j)

for causal attention where (j \le i). Each head (h) has a different slope (m_h).

The intuition is simple:

The farther back a token is, the more attention must “pay” to use it.

Strengths: extremely simple; no position embedding table; efficient; designed for train-short-test-long extrapolation; strong recency bias.

Weaknesses: less expressive than RoPE or learned relative schemes; monotonic distance penalty may be too restrictive for tasks requiring precise long-distance retrieval; mostly natural for causal language modeling.

Best fit: causal decoder Transformers where length extrapolation and efficiency matter.

The ALiBi paper says it biases query-key scores with a penalty proportional to distance and reports that, in their 1.3B-parameter setup, training on length 1024 extrapolated to length 2048 while training faster and using less memory than a sinusoidal baseline trained at length 2048. (OpenReview)


3.8 2D positional embeddings and relative bias for vision

Images are not 1D sequences. Flattening an image into patches creates an artificial order: patch 15 and patch 16 may be adjacent horizontally, but patch 15 and patch 29 may be adjacent vertically depending on image width.

Vision models therefore need spatial PE:

prow,colp_{row,col}

Common choices include:

  • learned 1D patch embeddings,
  • learned 2D row/column embeddings,
  • 2D sinusoidal embeddings,
  • relative 2D position bias,
  • axial RoPE,
  • window-relative bias.

ViT used standard learnable 1D position embeddings and interpolated them when fine-tuning at higher resolution. (ar5iv) Swin Transformer uses windowed self-attention with relative position bias; its ablation reports that relative position bias improved ImageNet classification, COCO detection, and ADE20K segmentation compared with no position encoding or absolute position embedding. (ar5iv)

Strengths: 2D-aware PE respects image geometry better than naïve 1D position; relative bias helps dense prediction tasks.

Weaknesses: interpolation across resolution can be fragile; 2D RoPE variants need care; absolute learned embeddings may overfit to training resolution.

Best fit: ViTs, Swin-like architectures, object detection, segmentation, multimodal vision-language models.


3.9 Graph positional and structural encodings

For graphs, there is no universal “position 1, 2, 3.” Node order is arbitrary. If shuffling node IDs, the graph is the same.

So graph PE must encode structure, not sequence index.

Common graph PEs include:

degree(v),shortest path(u,v),Laplacian eigenvectors,random-walk features\text{degree}(v),\quad \text{shortest path}(u,v),\quad \text{Laplacian eigenvectors},\quad \text{random-walk features}

Graphormer uses structural encodings such as centrality and spatial encodings to make Transformer attention aware of graph topology. (arXiv) Laplacian eigenvector PEs are a spectral method: they treat graph structure like a generalized coordinate system, analogous to Fourier bases on sequences. (arXiv)

Strengths: essential for graph Transformers; can encode topology, roles, distances, and global structure.

Weaknesses: Laplacian eigenvectors can have sign ambiguity and stability issues; shortest-path encodings may be expensive for large graphs; graph PE design is task-dependent.

Best fit: molecular graphs, code graphs, knowledge graphs, social networks, graph-level prediction.


3.10 Fourier features and coordinate encodings

For coordinate-based MLPs, such as NeRF-style neural fields, standard MLPs struggle with high-frequency detail. Fourier features solve this by mapping coordinates into many sinusoidal frequencies.

Example for a coordinate (x):

γ(x)=[sin⁡(20πx),cos⁡(20πx),…,sin⁡(2Lπx),cos⁡(2Lπx)]\gamma(x)= [\sin(2^0\pi x),\cos(2^0\pi x),\ldots,\sin(2^L\pi x),\cos(2^L\pi x)]

Strengths: excellent for continuous coordinates, 3D scenes, implicit neural representations, high-frequency signals.

Weaknesses: frequency scale must be chosen carefully; too high can overfit or create artifacts; not a general replacement for sequence PE.

Best fit: NeRF, 3D reconstruction, coordinate MLPs, audio/image implicit functions, spatial regression.

Fourier feature work showed that Fourier mappings help MLPs learn high-frequency functions and can transform the effective neural tangent kernel into a stationary kernel with tunable bandwidth. (arXiv)


5. Comparison of major PE methods

MethodCore ideaAdvantagesDisadvantagesBest architecture fit
Learned absolute PELearn (p_i) for each positionSimple, strong for fixed lengthPoor extrapolation beyond trained positionsBERT-style encoders, ViT classifiers
Sinusoidal PEFixed multiscale sine/cosine vectorsNo parameters, arbitrary length computableExtrapolation not guaranteed in practiceClassic Transformers, baselines
Shaw relative PELearn embeddings for relative distancesGood relational inductive biasMore complex and costlyTranslation, encoder-decoder, local sequence tasks
T5 relative biasAdd learned scalar bias by distance bucketCheap, robust, efficientLess expressive; bucket design mattersText-to-text Transformers
Transformer-XL PERelative PE for segment recurrenceLong context with cached memoryMore complex architectureStreaming/long-document language modeling
RoPERotate Q/K by positionStrong relative behavior; no position tableLong-context scaling can be trickyModern decoder-only LLMs
ALiBiAdd linear distance penalty to attentionVery simple, extrapolation-friendlyStrong recency bias, less expressiveCausal LMs, train-short-test-long
2D relative biasBias attention by 2D patch offsetGood for image geometry and dense tasksWindow/resolution design mattersViT, Swin, detection, segmentation
Graph PEEncode topology, distance, centralityNecessary for graph TransformersTask-specific, can be unstableGraph Transformers
Fourier featuresEncode continuous coordinates with sin/cosGreat for high-frequency coordinate functionsFrequency tuning requiredNeRF, neural fields, coordinate MLPs

6. Effectiveness across different network architectures

Transformers

Transformers need PE the most because attention itself is content-based and, in full self-attention, order-independent. The original Transformer explicitly needed PE after removing recurrence and convolution. (arXiv)

For encoder-only Transformers, learned absolute PE and relative bias both work. Relative bias is usually better when local distance patterns matter.

For decoder-only LLMs, RoPE and ALiBi are especially effective. RoPE is more expressive; ALiBi is simpler and more extrapolation-oriented.

For encoder-decoder models, T5-style relative bias is elegant because both input-side and output-side relative relations matter.


RNNs and LSTMs

RNNs process tokens one step at a time:

ht=f(xt,ht−1)h_t=f(x_t,h_{t-1})

So order is built into the computation. They usually do not need positional encoding to know sequence order.

But PE can still help if the task requires absolute time information, such as “this happened at hour 23” or “this is the 400th step.”

Best PE choice: usually none, or simple time-step/clock features for irregular time series.


CNNs

CNNs have locality and translation equivariance built in. A 3×3 kernel knows “neighbor above,” “neighbor below,” etc., but it does not naturally know absolute coordinate unless padding, boundaries, or extra channels reveal it.

CoordConv adds explicit coordinate channels, letting convolutional layers know where they are in Cartesian space. The CoordConv paper frames this as giving convolution access to its own input coordinates through extra coordinate channels. (arXiv)

Best PE choice: coordinate channels, 2D sinusoidal features, or learned spatial embeddings when absolute position matters.

Examples: medical imaging, segmentation, object localization, document layout, game boards.


Vision Transformers

ViTs convert images into patch sequences. Because flattening destroys some 2D structure, PE is crucial.

For image classification, learned absolute patch embeddings often work well. For dense tasks like detection and segmentation, 2D relative position bias is often better because it preserves local spatial relationships and some translation-friendly inductive bias. Swin’s ablations support the value of relative position bias for classification, detection, and segmentation. (ar5iv)

Best PE choice: 2D learned PE, 2D sinusoidal PE, 2D relative bias, or 2D RoPE depending on architecture.


Graph Transformers

Graph Transformers need structural PE because node order is arbitrary. A graph model should not change its prediction merely because nodes were renumbered.

Best PE choice: shortest-path distance bias, degree/centrality encoding, Laplacian eigenvectors, random-walk structural features, edge encodings.

Graphormer is a key example: it injects graph structure into a standard Transformer to make it competitive on graph representation tasks. (arXiv)


State-space models and Mamba-like architectures

State-space and recurrent sequence models process sequences through a scan or recurrence-like mechanism, so order is part of the computation. Mamba, for example, is a selective state-space model with linear scaling in sequence length and a hardware-aware recurrent-mode algorithm. (arXiv)

They generally do not need PE in the same way a full self-attention Transformer does. But if the input is a flattened image, irregular time series, or multimodal sequence, extra coordinate/time embeddings may still help.

Best PE choice: usually none for plain sequence order; add domain-specific time, coordinate, or modality-position features when needed.


Coordinate MLPs and neural fields

For coordinate MLPs, PE is often not optional. A raw coordinate like ((x,y,z)) is too low-frequency for representing fine detail. Fourier features give the model a rich basis of spatial frequencies.

Best PE choice: Fourier features, hash-grid encodings, multiresolution coordinate encodings.


7. A subtle modern caveat: NoPE

There is recent work on Transformers without explicit positional encoding, often called NoPE. This mostly applies to causal decoder-only Transformers, where the causal mask itself introduces some order asymmetry. Recent studies report that NoPE can sometimes generalize competitively to longer contexts, but also that it has limitations and that length generalization remains nuanced. (arXiv)

The important distinction:

  • In full bidirectional self-attention, no PE means the model is essentially permutation-equivariant.
  • In causal self-attention, the triangular mask gives a weak implicit positional signal.
  • NoPE is interesting, but it is not a general replacement for PE across architectures.

8. Practical rule of thumb

For a fixed-length text encoder, learned absolute PE is fine. For a robust encoder-decoder NLP system, T5-style relative position bias is a strong default. For modern decoder-only LLMs, RoPE is usually the most expressive standard choice, while ALiBi is attractive when simple length extrapolation is a priority. For vision, use 2D-aware absolute or relative PE; for dense prediction, relative bias is often preferable. For graphs, use structural encodings, not arbitrary node indices. For coordinate MLPs, use Fourier or multiresolution encodings. For RNNs, CNNs, and state-space models, PE is less fundamental because those architectures already contain order or locality biases, but explicit coordinates can still help when absolute position matters.

The core idea is simple:

Positional encoding is not just “adding position.” It is choosing the right inductive bias for how model should understand space, time, distance, order, or structure.