Technical note

The anatomy of 2D positional embedding

An in-depth analysis into 2D positional embedding.

  • Deep Learning
  • ViT

In a Vision Transformer, positional embeddings answer a simple question:

A patch embedding tells the transformer what is in a patch. A positional embedding tells it where that patch came from.

For images, the natural coordinate system is 2D, so each patch has a row coordinate yy and a column coordinate xx. A 2D sinusoidal positional embedding encodes both coordinates using fixed sine/cosine functions.

1. Start from the ViT patch grid

Suppose the image is

H×WH \times W

and the patch size is

P×P.P \times P.

Then the patch grid has

Hp=HP,Wp=WPH_p = \frac{H}{P}, \qquad W_p = \frac{W}{P}

patches along height and width.

For a standard ViT example:

224×224,P=16224 \times 224,\qquad P=16

so

Hp=Wp=14.H_p=W_p=14.

You therefore have a 14×1414\times14 grid:

196 patches.196 \text{ patches}.

Each patch gets embedded into a vector of dimension DD, giving

X∈R196×D.X\in \mathbb{R}^{196\times D}.

Without positional information, the transformer effectively sees these as an unordered set of 196 tokens.


2. First understand 1D sinusoidal embeddings

The original Transformer uses:

PE(pos,2i)=sin⁡(pos100002i/D)PE(pos,2i) = \sin \left( \frac{pos}{10000^{2i/D}} \right)

and

PE(pos,2i+1)=cos⁡(pos100002i/D).PE(pos,2i+1) = \cos \left( \frac{pos}{10000^{2i/D}} \right).

Here:

  • pospos = token position
  • ii = frequency index
  • DD = embedding dimension

A more intuitive notation is

ωi=1100002i/D,\omega_i = \frac{1}{10000^{2i/D}},

so that

PE(pos)=[sin⁡(posω0),cos⁡(posω0),sin⁡(posω1),cos⁡(posω1),… ].PE(pos) = [ \sin(pos\omega_0), \cos(pos\omega_0), \sin(pos\omega_1), \cos(pos\omega_1), \dots ].

The important idea is that every position is represented at many frequencies.

Some dimensions vary rapidly:

sin⁡(pos)\sin(pos)

while others vary very slowly:

sin⁡(0.001 pos).\sin(0.001\,pos).

This gives the model both fine-grained and coarse positional information.


3. Why 1D isn’t quite right for images

You could flatten a 14×1414\times14 image grid into:

0,1,2,…,195.0,1,2,\dots,195.

For example:

(0,0)→0(0,0)\rightarrow0 (0,1)→1(0,1)\rightarrow1 …\dots (1,0)→14.(1,0)\rightarrow14.

But this destroys some of the explicit 2D structure.

For example, patches

(0,13)(0,13)

and

(1,0)(1,0)

are adjacent in the flattened sequence:

13, 14,13,\ 14,

even though spatially one is at the far right of one row and the other is at the far left of the next.

A 2D embedding instead represents

(x,y)(x,y)

directly.


4. The basic idea of 2D sin/cos embeddings

Suppose the ViT hidden dimension is DD.

A common design is:

PE(x,y)=[PEx(x)  ∣∣  PEy(y)]PE(x,y) = [ PE_x(x) \; || \; PE_y(y) ]

where ∣∣|| means concatenation.

Half of the dimensions encode the xx-coordinate:

D2\frac{D}{2}

and half encode the yy-coordinate:

D2.\frac{D}{2}.

Each half itself contains sine/cosine pairs.

So effectively:

D=D4+D4+D4+D4\boxed{ D = \frac D4 + \frac D4 + \frac D4 + \frac D4 }

corresponding to:

[sin⁡(x),cos⁡(x),sin⁡(y),cos⁡(y)][ \sin(x), \cos(x), \sin(y), \cos(y) ]

across many frequencies.

This is why implementations often require:

D mod 4=0.D \bmod 4 = 0.

5. Mathematical definition

Let

d=D4.d = \frac{D}{4}.

Construct dd frequencies:

ωi=110000i/d,i=0,…,d−1.\omega_i = \frac{1}{10000^{i/d}}, \qquad i=0,\dots,d-1.

There are slightly different conventions in implementations, but the idea is identical.

For horizontal coordinate xx:

Ex(x)=[sin⁡(xω0),…,sin⁡(xωd−1),cos⁡(xω0),…,cos⁡(xωd−1)].E_x(x) = [ \sin(x\omega_0), \dots, \sin(x\omega_{d-1}), \cos(x\omega_0), \dots, \cos(x\omega_{d-1}) ].

This has dimension

2d=D2.2d=\frac D2.

Similarly:

Ey(y)=[sin⁡(yω0),…,sin⁡(yωd−1),cos⁡(yω0),…,cos⁡(yωd−1)].E_y(y) = [ \sin(y\omega_0), \dots, \sin(y\omega_{d-1}), \cos(y\omega_0), \dots, \cos(y\omega_{d-1}) ].

Then:

PE(x,y)=Ex(x)  ∣∣  Ey(y)\boxed{ PE(x,y) = E_x(x)\;||\;E_y(y) }

has total dimension:

D2+D2=D.\frac D2+\frac D2=D.

6. Small concrete example

Imagine a tiny patch grid:

3×4.3\times4.

The coordinates are:

(0,0)(0,1)(0,2)(0,3)(1,0)(1,1)(1,2)(1,3)(2,0)(2,1)(2,2)(2,3)\begin{matrix} (0,0)&(0,1)&(0,2)&(0,3)\\ (1,0)&(1,1)&(1,2)&(1,3)\\ (2,0)&(2,1)&(2,2)&(2,3) \end{matrix}

depending on whether you write coordinates as (y,x)(y,x) or (x,y)(x,y).

Let’s use:

(x,y).(x,y).

Suppose:

D=8.D=8.

Then:

D/2=4D/2=4

dimensions encode xx, and 4 encode yy.

For simplicity, suppose we only use two frequencies:

ω0=1,ω1=0.1.\omega_0=1,\qquad \omega_1=0.1.

Then the xx-embedding could be:

Ex(x)=[sin⁡(x),cos⁡(x),sin⁡(0.1x),cos⁡(0.1x)].E_x(x) = [ \sin(x), \cos(x), \sin(0.1x), \cos(0.1x) ].

And similarly:

Ey(y)=[sin⁡(y),cos⁡(y),sin⁡(0.1y),cos⁡(0.1y)].E_y(y) = [ \sin(y), \cos(y), \sin(0.1y), \cos(0.1y) ].

For patch:

(x,y)=(2,1),(x,y)=(2,1),

we get:

PE(2,1)=[sin⁡2,cos⁡2,sin⁡0.2,cos⁡0.2,sin⁡1,cos⁡1,sin⁡0.1,cos⁡0.1].PE(2,1) = [ \sin2, \cos2, \sin0.2, \cos0.2, \sin1, \cos1, \sin0.1, \cos0.1 ].

Numerically, roughly:

PE(2,1)≈[0.909,−0.416,0.199,0.980,0.841,0.540,0.100,0.995].PE(2,1) \approx [ 0.909, -0.416, 0.199, 0.980, 0.841, 0.540, 0.100, 0.995 ].

That vector uniquely characterizes the patch’s location across multiple spatial frequencies.


7. What the full position tensor looks like

For an Hp×WpH_p\times W_p grid, you create:

PE∈RHp×Wp×D.PE\in \mathbb{R}^{H_p\times W_p\times D}.

For a 14×1414\times14 ViT with D=768D=768:

PE∈R14×14×768.PE\in \mathbb{R}^{14\times14\times768}.

You then flatten the spatial dimensions:

PE→R196×768.PE \rightarrow \mathbb{R}^{196\times768}.

Your patch embeddings also have shape:

X∈R196×768.X\in\mathbb{R}^{196\times768}.

Then simply:

X′=X+PE\boxed{ X' = X + PE }

before feeding the tokens into the Transformer.


8. Constructing the coordinate grids

Conceptually, you first create:

x=[012301230123]x= \begin{bmatrix} 0&1&2&3\\ 0&1&2&3\\ 0&1&2&3 \end{bmatrix}

and

y=[000011112222].y= \begin{bmatrix} 0&0&0&0\\ 1&1&1&1\\ 2&2&2&2 \end{bmatrix}.

For a 3×43\times4 grid.

Then encode every scalar in xx with sin/cos and every scalar in yy with sin/cos.

This produces something conceptually like:

Ex∈R3×4×D/2E_x \in \mathbb{R}^{3\times4\times D/2}

and

Ey∈R3×4×D/2.E_y \in \mathbb{R}^{3\times4\times D/2}.

Concatenate:

PE=concat⁡(Ex,Ey)PE = \operatorname{concat}(E_x,E_y)

giving:

PE∈R3×4×D.PE \in \mathbb{R}^{3\times4\times D}.

9. A PyTorch implementation

Here’s a relatively transparent implementation:

import torch


def get_1d_sincos_pos_embed(pos, embed_dim):
    """
    pos: (...,) coordinate tensor
    returns: (..., embed_dim)

    embed_dim must be even.
    """
    assert embed_dim % 2 == 0

    half_dim = embed_dim // 2

    omega = torch.arange(half_dim, dtype=torch.float32)

    omega = 1.0 / (
        10000 ** (omega / half_dim)
    )

    # (..., 1) * (half_dim,)
    angles = pos[..., None] * omega

    emb = torch.cat([
        torch.sin(angles),
        torch.cos(angles)
    ], dim=-1)

    return emb


def get_2d_sincos_pos_embed(height, width, embed_dim):
    """
    returns:
        (height * width, embed_dim)
    """
    assert embed_dim % 4 == 0

    y = torch.arange(height, dtype=torch.float32)
    x = torch.arange(width, dtype=torch.float32)

    yy, xx = torch.meshgrid(
        y,
        x,
        indexing="ij"
    )

    # Each coordinate receives half the total embedding.
    x_emb = get_1d_sincos_pos_embed(
        xx,
        embed_dim // 2
    )

    y_emb = get_1d_sincos_pos_embed(
        yy,
        embed_dim // 2
    )

    pos_emb = torch.cat(
        [x_emb, y_emb],
        dim=-1
    )

    # H x W x D -> HW x D
    pos_emb = pos_emb.reshape(
        height * width,
        embed_dim
    )

    return pos_emb

For:

pos_embed = get_2d_sincos_pos_embed(
    height=14,
    width=14,
    embed_dim=768
)

the shape is:

[196, 768]

which can be added directly to:

patch_tokens.shape = [B, 196, 768]

by broadcasting:

patch_tokens = patch_tokens + pos_embed.unsqueeze(0)

9.1. Understanding the code mathematically

This line:

omega = torch.arange(half_dim)

creates something like:

[0,1,2,… ].[0,1,2,\dots].

Then:

omega = 1 / (10000 ** (omega / half_dim))

produces frequencies roughly like:

[1,0.8,0.6,…,0.001,0.0001].[ 1, 0.8, 0.6, \dots, 0.001, 0.0001 ].

Not those exact values, but progressively smaller frequencies.

Then:

angles = pos[..., None] * omega

computes the outer product:

pos×ωi.pos\times\omega_i.

If

pos=3,pos=3,

you get:

[3ω0,3ω1,3ω2,… ].[ 3\omega_0, 3\omega_1, 3\omega_2, \dots ].

Finally:

sin(angles)
cos(angles)

produces the positional vector.


10. Why sine and cosine?

There are several useful properties.

No trainable parameters

The embeddings are deterministic:

PE=f(x,y).PE=f(x,y).

So you don’t need a learned table like:

P∈R196×D.P\in\mathbb R^{196\times D}.

That means:

0 learned positional parameters\boxed{\text{0 learned positional parameters}}

for sin/cos embeddings.

Smooth positions

Nearby coordinates produce nearby representations.

If:

x=4x=4

and

x=5,x=5,

then their low-frequency components are quite similar.

This reflects spatial continuity.

Multiple spatial scales

High frequencies distinguish neighboring patches.

Low frequencies represent broader location.

You can loosely think of the dimensions as encoding:

  • exact/local position
  • medium-scale position
  • coarse/global position

simultaneously.

Relative position can be inferred

A particularly elegant property comes from trigonometric identities.

For example:

sin⁡(a−b)=sin⁡acos⁡b−cos⁡asin⁡b.\sin(a-b) = \sin a\cos b-\cos a\sin b.

Because both sine and cosine are included, linear combinations of embeddings can expose information related to:

x1−x2.x_1-x_2.

So attention layers can relatively easily learn spatial displacement relationships.


11. Why encode x and y separately?

Consider these two patches:

A=(2,5)A=(2,5)

and

B=(2,8).B=(2,8).

They have the exact same vertical position:

yA=yB=2.y_A=y_B=2.

Therefore their yy-embedding is identical:

Ey(2)=Ey(2).E_y(2)=E_y(2).

Only their horizontal embedding changes.

Likewise:

A=(2,5)A=(2,5)

and

C=(7,5)C=(7,5)

share their xx-embedding but differ in their yy-embedding.

This factorization gives the Transformer a clean representation of image geometry.


12. Learned ViT embeddings vs fixed 2D sin/cos

The original ViT commonly uses a learned positional embedding:

P∈RN×D.P\in\mathbb R^{N\times D}.

Each patch index gets its own learned vector:

P0,P1,…,PN−1.P_0,P_1,\dots,P_{N-1}.

2D sin/cos instead defines:

Px,y=f(x,y).P_{x,y}=f(x,y).

There are important differences:

Learned2D sin/cos
TrainableYesNo
ParametersN×DN\times D0
Explicit 2D structureNot necessarilyYes
Arbitrary resolutionHarderEasier
Must interpolateOftenUsually not
Common in vanilla ViTYesSometimes
Common in MAE-style modelsVery commonYes

For example, Meta’s MAE architecture famously uses fixed 2D sinusoidal positional embeddings.


13. One major advantage: changing image resolution

Suppose the network was trained with:

14×1414\times14

patches but at inference you want:

16×16.16\times16.

With learned embeddings, your learned matrix might be:

196×768.196\times768.

But now you need:

256×768.256\times768.

You generally have to interpolate the learned positional grid.

With sinusoidal embeddings, you can simply compute:

PE(x,y)PE(x,y)

for:

x,y=0,…,15.x,y=0,\dots,15.

No learned table needs resizing.

That’s one reason fixed embeddings are attractive for architectures that may process varying resolutions.

Limitation

That advantage is real, but it is not unlimited resolution extrapolation.

A fixed 2D sin/cos embedding solves the representation problem:

PE(x,y)PE(x,y)

can be computed for coordinates that were never present during training.

But it does not automatically solve the generalization problem:

Can the trained transformer correctly interpret those unseen coordinates and spatial relationships?\text{Can the trained transformer correctly interpret those unseen coordinates and spatial relationships?}

Those are different things.

Suppose training used a 14×1414\times14 patch grid, so the model only ever saw

x,y∈[0,13].x,y\in[0,13].

At inference, if you give it a 28×2828\times28 grid, sin/cos PE can generate embeddings for

x,y∈[0,27].x,y\in[0,27].

Mathematically, there is no problem:

PE(14,0),PE(15,0),…,PE(27,27)PE(14,0), PE(15,0), \ldots, PE(27,27)

are perfectly well-defined.

But the transformer weights were optimized only under the distribution of positional codes corresponding to the original region. So positions such as x=24x=24 or relationships like

Δx=20\Delta x = 20

may be completely out-of-distribution.

That creates several limitations.

First, absolute-position extrapolation can fail. During training, some sin/cos channels only occupied a limited portion of their sinusoidal cycle. At larger coordinates, the model may encounter combinations of phase values it never saw. Although the embedding function is smooth, the downstream MLPs and attention layers are not guaranteed to interpret those novel combinations correctly.

Second, relative distances also extrapolate. If the largest horizontal separation during training was

13,13,

then inference on a 28×2828\times28 grid introduces distances up to

27.27.

The sinusoidal representation does have useful algebraic structure—for example,

cos⁡(a−b)=cos⁡acos⁡b+sin⁡asin⁡b,\cos(a-b)=\cos a\cos b+\sin a\sin b,

so relative displacement can in principle be recovered from the embeddings. But “can be represented” is different from “the model learned to use it correctly for that distance.”

Third, changing resolution changes more than the PE. It changes the attention regime. A 14×1414\times14 grid has 196 patch tokens. A 28×2828\times28 grid has 784. Each token now attends over a much larger set, and objects can occupy very different numbers of tokens. The model’s learned attention patterns may therefore no longer correspond to the spatial scales encountered during training.

A useful distinction is:

sin/cos PE has unbounded mathematical support, but bounded learned generalization\boxed{ \text{sin/cos PE has unbounded mathematical support, but bounded learned generalization} }

There isn’t usually a sharp expansion limit like “1.4× works and 1.5× fails.” It is more of a distribution-shift curve. Small extrapolations often work better than large ones because the new geometry remains closer to what the model saw.

For example, a model trained on 14×1414\times14 patches might plausibly tolerate

14×14→16×1614\times14 \rightarrow 16\times16

better than

14×14→56×56.14\times14 \rightarrow 56\times56.

The first introduces moderately larger coordinates and distances. The second creates dramatically different absolute positions, relative distances, token counts, object scales, and global attention patterns.

There is another subtle issue: whether the coordinates are absolute patch indices or normalized coordinates.

With ordinary PE,

PE(x,y),PE(x,y),

a patch at the far-right edge has coordinate x=13x=13 during 14×1414\times14 training, but x=27x=27 during 28×2828\times28 inference. Thus “right edge of image” is encoded differently.

If instead coordinates are normalized, for example

x~=xWp−1,y~=yHp−1,\tilde{x}=\frac{x}{W_p-1}, \qquad \tilde{y}=\frac{y}{H_p-1},

then the right edge is always approximately

x~=1.\tilde{x}=1.

That improves scale invariance in one sense, but it changes the meaning of distances. A one-patch displacement at different resolutions no longer has the same coordinate difference. So there is a tradeoff between preserving absolute patch spacing and preserving relative image location.

This is one reason modern vision architectures often go beyond basic absolute 2D sin/cos PE. Relative position bias, RoPE-style encodings, decomposed relative position embeddings, coordinate normalization, resolution augmentation during training, and multi-scale pretraining all try to make spatial generalization more robust.

So I would phrase the original advantage carefully:

Fixed 2D sin/cos PE removes the need to interpolate a learned positional table when image resolution changes. It lets the architecture represent arbitrary new patch coordinates analytically. However, it does not guarantee that the transformer has learned how to interpret positions, distances, or attention configurations outside the range encountered during training.

In fact, there are two separate extrapolation axes:

coordinate extrapolation\boxed{\text{coordinate extrapolation}}

and

visual-scale / token-count extrapolation.\boxed{\text{visual-scale / token-count extrapolation}}.

Sin/cos PE helps substantially with the first at the encoding level, but it does not eliminate either problem at the learned-model level.


14. What about the CLS token?

A ViT often has:

[CLS,p1,p2,…,pN].[\text{CLS},p_1,p_2,\dots,p_N].

Your 2D positional embedding naturally applies only to the patch tokens.

A common approach is:

PCLS=0.P_{\text{CLS}}=0.

So:

PE=[0,PE(0,0),PE(0,1),… ].PE = [ \mathbf 0, PE(0,0), PE(0,1), \dots ].

For a 14×1414\times14 grid and D=768D=768:

PEpatch∈R196×768,PE_{\text{patch}} \in \mathbb R^{196\times768},

then prepend:

0∈R1×768,0\in\mathbb R^{1\times768},

giving:

PEfull∈R197×768.PE_{\text{full}} \in \mathbb R^{197\times768}.

For example:

cls_pos = torch.zeros(1, 768)

pos_embed = torch.cat([
    cls_pos,
    pos_embed
], dim=0)

Now:

pos_embed.shape
= [197, 768]

15. An intuitive mental model

Think of the embedding dimensions as a collection of waves painted over the image.

One dimension might vary quickly horizontally:

sin⁡(x).\sin(x).

Another varies slowly horizontally:

sin⁡(0.01x).\sin(0.01x).

Another varies quickly vertically:

sin⁡(y).\sin(y).

Another varies slowly vertically:

sin⁡(0.01y).\sin(0.01y).

So every patch is effectively asking:

What values do all these horizontal and vertical waves have at my location?

That combination forms a unique positional signature.

You can visualize it roughly as:

             x direction
        ───────────────────>

 y     patch       [horizontal waves,
 |                  vertical waves]
 |
 |
 v

Rather than assigning each position an arbitrary ID, we’re assigning it coordinates within a family of periodic basis functions.


16. A useful signal-processing interpretation

There is also a deeper way to understand sinusoidal positional embeddings.

They are essentially representing position in a Fourier-like basis.

Instead of describing a coordinate directly as:

x=7,x=7,

we describe it through responses to many frequencies:

[sin⁡(ω1x),cos⁡(ω1x),sin⁡(ω2x),cos⁡(ω2x),… ].[ \sin(\omega_1x), \cos(\omega_1x), \sin(\omega_2x), \cos(\omega_2x), \dots ].

This resembles Fourier feature mappings.

For 2D:

(x,y)→[ϕ(x),ϕ(y)].(x,y) \rightarrow [ \phi(x),\phi(y) ].

So the transformer receives a high-dimensional representation of the spatial coordinate.


17. The complete ViT pipeline

Suppose:

I∈RB×3×224×224.I\in \mathbb R^{B\times3\times224\times224}.

Patchification

Using 16×1616\times16 patches:

I→196 patches.I \rightarrow 196\text{ patches}.

Each patch contains:

16×16×3=76816\times16\times3=768

raw values.

Patch projection

A learned linear layer maps each patch:

R768→RD.\mathbb R^{768} \rightarrow \mathbb R^D.

Suppose:

D=768.D=768.

Then:

X∈RB×196×768.X\in \mathbb R^{B\times196\times768}.

Generate spatial coordinates

Patch positions:

(x,y)∈{0,…,13}2.(x,y) \in \{0,\dots,13\}^2.

Generate sinusoidal embedding

PE∈R196×768.PE\in \mathbb R^{196\times768}.

Add

Z=X+PE.Z=X+PE.

Optional CLS token

Z→[CLS,z1,…,z196].Z \rightarrow [ CLS,z_1,\dots,z_{196} ].

Transformer

Then:

Z→Transformer blocks.Z \rightarrow \text{Transformer blocks}.

18. The most important thing to remember

If you want one formula that captures the idea, it is:

PE(x,y)=[sin⁡(xω),cos⁡(xω),sin⁡(yω),cos⁡(yω)]\boxed{ PE(x,y) = [ \sin(x\omega), \cos(x\omega), \sin(y\omega), \cos(y\omega) ] }

where each expression actually represents a vector evaluated over many frequencies:

ω0,ω1,…,ωD/4−1.\omega_0,\omega_1,\dots,\omega_{D/4-1}.

So for a ViT token:

zx,y=PatchEmbed(Ix,y)+PE(x,y)\boxed{ z_{x,y} = \text{PatchEmbed}(I_{x,y}) + PE(x,y) }

The patch embedding represents visual content; the sin/cos vector represents spatial coordinates.

One subtle point that causes a lot of confusion: the four pieces are not usually literally one scalar each. For a D=768D=768 ViT, you effectively have 192 sine dimensions for xx, 192 cosine dimensions for xx, 192 sine dimensions for yy, and 192 cosine dimensions for yy. That is where the full 768-dimensional positional vector comes from.