Technical note
The anatomy of 2D positional embedding
An in-depth analysis into 2D positional embedding.
In a Vision Transformer, positional embeddings answer a simple question:
A patch embedding tells the transformer what is in a patch. A positional embedding tells it where that patch came from.
For images, the natural coordinate system is 2D, so each patch has a row coordinate and a column coordinate . A 2D sinusoidal positional embedding encodes both coordinates using fixed sine/cosine functions.
1. Start from the ViT patch grid
Suppose the image is
and the patch size is
Then the patch grid has
patches along height and width.
For a standard ViT example:
so
You therefore have a grid:
Each patch gets embedded into a vector of dimension , giving
Without positional information, the transformer effectively sees these as an unordered set of 196 tokens.
2. First understand 1D sinusoidal embeddings
The original Transformer uses:
and
Here:
- = token position
- = frequency index
- = embedding dimension
A more intuitive notation is
so that
The important idea is that every position is represented at many frequencies.
Some dimensions vary rapidly:
while others vary very slowly:
This gives the model both fine-grained and coarse positional information.
3. Why 1D isn’t quite right for images
You could flatten a image grid into:
For example:
But this destroys some of the explicit 2D structure.
For example, patches
and
are adjacent in the flattened sequence:
even though spatially one is at the far right of one row and the other is at the far left of the next.
A 2D embedding instead represents
directly.
4. The basic idea of 2D sin/cos embeddings
Suppose the ViT hidden dimension is .
A common design is:
where means concatenation.
Half of the dimensions encode the -coordinate:
and half encode the -coordinate:
Each half itself contains sine/cosine pairs.
So effectively:
corresponding to:
across many frequencies.
This is why implementations often require:
5. Mathematical definition
Let
Construct frequencies:
There are slightly different conventions in implementations, but the idea is identical.
For horizontal coordinate :
This has dimension
Similarly:
Then:
has total dimension:
6. Small concrete example
Imagine a tiny patch grid:
The coordinates are:
depending on whether you write coordinates as or .
Let’s use:
Suppose:
Then:
dimensions encode , and 4 encode .
For simplicity, suppose we only use two frequencies:
Then the -embedding could be:
And similarly:
For patch:
we get:
Numerically, roughly:
That vector uniquely characterizes the patch’s location across multiple spatial frequencies.
7. What the full position tensor looks like
For an grid, you create:
For a ViT with :
You then flatten the spatial dimensions:
Your patch embeddings also have shape:
Then simply:
before feeding the tokens into the Transformer.
8. Constructing the coordinate grids
Conceptually, you first create:
and
For a grid.
Then encode every scalar in with sin/cos and every scalar in with sin/cos.
This produces something conceptually like:
and
Concatenate:
giving:
9. A PyTorch implementation
Here’s a relatively transparent implementation:
import torch
def get_1d_sincos_pos_embed(pos, embed_dim):
"""
pos: (...,) coordinate tensor
returns: (..., embed_dim)
embed_dim must be even.
"""
assert embed_dim % 2 == 0
half_dim = embed_dim // 2
omega = torch.arange(half_dim, dtype=torch.float32)
omega = 1.0 / (
10000 ** (omega / half_dim)
)
# (..., 1) * (half_dim,)
angles = pos[..., None] * omega
emb = torch.cat([
torch.sin(angles),
torch.cos(angles)
], dim=-1)
return emb
def get_2d_sincos_pos_embed(height, width, embed_dim):
"""
returns:
(height * width, embed_dim)
"""
assert embed_dim % 4 == 0
y = torch.arange(height, dtype=torch.float32)
x = torch.arange(width, dtype=torch.float32)
yy, xx = torch.meshgrid(
y,
x,
indexing="ij"
)
# Each coordinate receives half the total embedding.
x_emb = get_1d_sincos_pos_embed(
xx,
embed_dim // 2
)
y_emb = get_1d_sincos_pos_embed(
yy,
embed_dim // 2
)
pos_emb = torch.cat(
[x_emb, y_emb],
dim=-1
)
# H x W x D -> HW x D
pos_emb = pos_emb.reshape(
height * width,
embed_dim
)
return pos_emb
For:
pos_embed = get_2d_sincos_pos_embed(
height=14,
width=14,
embed_dim=768
)
the shape is:
[196, 768]
which can be added directly to:
patch_tokens.shape = [B, 196, 768]
by broadcasting:
patch_tokens = patch_tokens + pos_embed.unsqueeze(0)
9.1. Understanding the code mathematically
This line:
omega = torch.arange(half_dim)
creates something like:
Then:
omega = 1 / (10000 ** (omega / half_dim))
produces frequencies roughly like:
Not those exact values, but progressively smaller frequencies.
Then:
angles = pos[..., None] * omega
computes the outer product:
If
you get:
Finally:
sin(angles)
cos(angles)
produces the positional vector.
10. Why sine and cosine?
There are several useful properties.
No trainable parameters
The embeddings are deterministic:
So you don’t need a learned table like:
That means:
for sin/cos embeddings.
Smooth positions
Nearby coordinates produce nearby representations.
If:
and
then their low-frequency components are quite similar.
This reflects spatial continuity.
Multiple spatial scales
High frequencies distinguish neighboring patches.
Low frequencies represent broader location.
You can loosely think of the dimensions as encoding:
- exact/local position
- medium-scale position
- coarse/global position
simultaneously.
Relative position can be inferred
A particularly elegant property comes from trigonometric identities.
For example:
Because both sine and cosine are included, linear combinations of embeddings can expose information related to:
So attention layers can relatively easily learn spatial displacement relationships.
11. Why encode x and y separately?
Consider these two patches:
and
They have the exact same vertical position:
Therefore their -embedding is identical:
Only their horizontal embedding changes.
Likewise:
and
share their -embedding but differ in their -embedding.
This factorization gives the Transformer a clean representation of image geometry.
12. Learned ViT embeddings vs fixed 2D sin/cos
The original ViT commonly uses a learned positional embedding:
Each patch index gets its own learned vector:
2D sin/cos instead defines:
There are important differences:
| Learned | 2D sin/cos | |
|---|---|---|
| Trainable | Yes | No |
| Parameters | 0 | |
| Explicit 2D structure | Not necessarily | Yes |
| Arbitrary resolution | Harder | Easier |
| Must interpolate | Often | Usually not |
| Common in vanilla ViT | Yes | Sometimes |
| Common in MAE-style models | Very common | Yes |
For example, Meta’s MAE architecture famously uses fixed 2D sinusoidal positional embeddings.
13. One major advantage: changing image resolution
Suppose the network was trained with:
patches but at inference you want:
With learned embeddings, your learned matrix might be:
But now you need:
You generally have to interpolate the learned positional grid.
With sinusoidal embeddings, you can simply compute:
for:
No learned table needs resizing.
That’s one reason fixed embeddings are attractive for architectures that may process varying resolutions.
Limitation
That advantage is real, but it is not unlimited resolution extrapolation.
A fixed 2D sin/cos embedding solves the representation problem:
can be computed for coordinates that were never present during training.
But it does not automatically solve the generalization problem:
Those are different things.
Suppose training used a patch grid, so the model only ever saw
At inference, if you give it a grid, sin/cos PE can generate embeddings for
Mathematically, there is no problem:
are perfectly well-defined.
But the transformer weights were optimized only under the distribution of positional codes corresponding to the original region. So positions such as or relationships like
may be completely out-of-distribution.
That creates several limitations.
First, absolute-position extrapolation can fail. During training, some sin/cos channels only occupied a limited portion of their sinusoidal cycle. At larger coordinates, the model may encounter combinations of phase values it never saw. Although the embedding function is smooth, the downstream MLPs and attention layers are not guaranteed to interpret those novel combinations correctly.
Second, relative distances also extrapolate. If the largest horizontal separation during training was
then inference on a grid introduces distances up to
The sinusoidal representation does have useful algebraic structure—for example,
so relative displacement can in principle be recovered from the embeddings. But “can be represented” is different from “the model learned to use it correctly for that distance.”
Third, changing resolution changes more than the PE. It changes the attention regime. A grid has 196 patch tokens. A grid has 784. Each token now attends over a much larger set, and objects can occupy very different numbers of tokens. The model’s learned attention patterns may therefore no longer correspond to the spatial scales encountered during training.
A useful distinction is:
There isn’t usually a sharp expansion limit like “1.4× works and 1.5× fails.” It is more of a distribution-shift curve. Small extrapolations often work better than large ones because the new geometry remains closer to what the model saw.
For example, a model trained on patches might plausibly tolerate
better than
The first introduces moderately larger coordinates and distances. The second creates dramatically different absolute positions, relative distances, token counts, object scales, and global attention patterns.
There is another subtle issue: whether the coordinates are absolute patch indices or normalized coordinates.
With ordinary PE,
a patch at the far-right edge has coordinate during training, but during inference. Thus “right edge of image” is encoded differently.
If instead coordinates are normalized, for example
then the right edge is always approximately
That improves scale invariance in one sense, but it changes the meaning of distances. A one-patch displacement at different resolutions no longer has the same coordinate difference. So there is a tradeoff between preserving absolute patch spacing and preserving relative image location.
This is one reason modern vision architectures often go beyond basic absolute 2D sin/cos PE. Relative position bias, RoPE-style encodings, decomposed relative position embeddings, coordinate normalization, resolution augmentation during training, and multi-scale pretraining all try to make spatial generalization more robust.
So I would phrase the original advantage carefully:
Fixed 2D sin/cos PE removes the need to interpolate a learned positional table when image resolution changes. It lets the architecture represent arbitrary new patch coordinates analytically. However, it does not guarantee that the transformer has learned how to interpret positions, distances, or attention configurations outside the range encountered during training.
In fact, there are two separate extrapolation axes:
and
Sin/cos PE helps substantially with the first at the encoding level, but it does not eliminate either problem at the learned-model level.
14. What about the CLS token?
A ViT often has:
Your 2D positional embedding naturally applies only to the patch tokens.
A common approach is:
So:
For a grid and :
then prepend:
giving:
For example:
cls_pos = torch.zeros(1, 768)
pos_embed = torch.cat([
cls_pos,
pos_embed
], dim=0)
Now:
pos_embed.shape
= [197, 768]
15. An intuitive mental model
Think of the embedding dimensions as a collection of waves painted over the image.
One dimension might vary quickly horizontally:
Another varies slowly horizontally:
Another varies quickly vertically:
Another varies slowly vertically:
So every patch is effectively asking:
What values do all these horizontal and vertical waves have at my location?
That combination forms a unique positional signature.
You can visualize it roughly as:
x direction
───────────────────>
y patch [horizontal waves,
| vertical waves]
|
|
v
Rather than assigning each position an arbitrary ID, we’re assigning it coordinates within a family of periodic basis functions.
16. A useful signal-processing interpretation
There is also a deeper way to understand sinusoidal positional embeddings.
They are essentially representing position in a Fourier-like basis.
Instead of describing a coordinate directly as:
we describe it through responses to many frequencies:
This resembles Fourier feature mappings.
For 2D:
So the transformer receives a high-dimensional representation of the spatial coordinate.
17. The complete ViT pipeline
Suppose:
Patchification
Using patches:
Each patch contains:
raw values.
Patch projection
A learned linear layer maps each patch:
Suppose:
Then:
Generate spatial coordinates
Patch positions:
Generate sinusoidal embedding
Add
Optional CLS token
Transformer
Then:
18. The most important thing to remember
If you want one formula that captures the idea, it is:
where each expression actually represents a vector evaluated over many frequencies:
So for a ViT token:
The patch embedding represents visual content; the sin/cos vector represents spatial coordinates.
One subtle point that causes a lot of confusion: the four pieces are not usually literally one scalar each. For a ViT, you effectively have 192 sine dimensions for , 192 cosine dimensions for , 192 sine dimensions for , and 192 cosine dimensions for . That is where the full 768-dimensional positional vector comes from.