Technical note

Positional Embedding Intermediate II

A introduction to PE interpolation/extrapolation.

  • Deep Learning
  • Transformer

PE interpolation is useful, but it is not a harmless detail. It is a spatial-prior transformation. Sometimes it is the right engineering trick; sometimes it silently changes the model’s geometry.

1. What PE interpolation actually does

Assume a ViT was trained on images of size

H0×W0H_0 \times W_0

with patch size

P×PP \times P

so its patch grid is

G0h=H0/P,G0w=W0/PG_0^h = H_0/P,\quad G_0^w = W_0/P

and its learned absolute positional embedding has shape

Epos∈R(G0hG0w+1)×dE_{\text{pos}} \in \mathbb{R}^{(G_0^hG_0^w+1)\times d}

where the extra +1+1 is usually the class token.

Now suppose we use a new image size

H1×W1H_1 \times W_1

with the same patch size PP. The new grid is

G1h=H1/P,G1w=W1/PG_1^h = H_1/P,\quad G_1^w = W_1/P

The old positional table no longer has the right number of patch positions. So the usual trick is:

Epos[1:]→reshape to [G0h,G0w,d]E_{\text{pos}}[1:] \rightarrow \text{reshape to } [G_0^h,G_0^w,d] →2D interpolate to [G1h,G1w,d]\rightarrow \text{2D interpolate to } [G_1^h,G_1^w,d] →flatten back to [G1hG1w,d]\rightarrow \text{flatten back to } [G_1^hG_1^w,d]

The class-token PE is usually kept unchanged.

This is exactly the workaround used in ViT-style models: ViT can handle arbitrary sequence length up to memory limits, but the pretrained position embeddings may no longer be meaningful when the image resolution changes, so the authors perform 2D interpolation of the pretrained position embeddings according to their original image locations. They also note that patch extraction and this PE interpolation are almost the only places where 2D image structure is manually injected into vanilla ViT.


2. The hidden assumption: learned PE is a smooth spatial field

PE interpolation treats the learned position table as if it were samples from a smooth function:

Epos(r,c)E_{\text{pos}}(r,c)

where r,cr,c are 2D image coordinates.

So interpolation assumes:

nearby positions should have nearby PE vectors\text{nearby positions should have nearby PE vectors}

This is often reasonable. The original ViT paper observed that learned position embeddings tend to encode image distance: closer patches often have more similar position embeddings, patches in the same row or column can have similar embeddings, and sinusoidal-like structures sometimes appear for larger grids.

But this is still an assumption, not a guarantee. If the learned PE table contains high-frequency, task-specific, or dataset-specific spatial patterns, interpolation may blur, stretch, or distort them.


3. Main impacts of PE interpolation

Impact 1: It enables checkpoint reuse at a new resolution

This is the positive impact. Without interpolation, a learned absolute PE table trained on a 14×1414 \times 14 patch grid cannot be directly added to a 24×2424 \times 24 patch grid.

So PE interpolation lets us reuse a pretrained ViT at higher or lower image resolution. This is why libraries expose switches such as interpolate_pos_encoding; Hugging Face’s ViT documentation notes that fine-tuning at higher resolution is possible by setting this option, which interpolates the pretrained position embeddings. (Hugging Face)


Impact 2: It changes the model’s spatial prior

In additive absolute PE, the token becomes

zr,c=xr,c+Epos(r,c)z_{r,c}=x_{r,c}+E_{\text{pos}}(r,c)

After interpolation, the content token xr,cx_{r,c} may be produced by the same patch encoder, but the position vector is different from anything the model saw during pretraining.

Then:

qr,c=(xr,c+Eposinterp(r,c))WQq_{r,c}=(x_{r,c}+E_{\text{pos}}^{interp}(r,c))W_Q kr,c=(xr,c+Eposinterp(r,c))WKk_{r,c}=(x_{r,c}+E_{\text{pos}}^{interp}(r,c))W_K

So interpolation affects attention through query/key/value projections. It does not merely “fill missing positions.” It changes the geometry of attention.


Impact 3: It rescales position

This is subtle and important.

Suppose the model was trained on a 14×1414 \times 14 grid and I use it on a 28×2828 \times 28 grid. Interpolation usually treats the old PE table as covering the whole image from top to bottom and left to right.

So the center of the new 28×2828 \times 28 image receives a PE similar to the center of the old 14×1414 \times 14 image.

That means interpolation usually preserves normalized image coordinates, not absolute patch-index distance.

This is good when the task is scale-normalized classification:

“Is this a dog?”\text{“Is this a dog?”}

It may be less good when the task depends on exact patch-level geometry:

“Where exactly is the boundary?”\text{“Where exactly is the boundary?”} “How far is object A from object B in pixels?”\text{“How far is object A from object B in pixels?”} “Which cell, tile, or document region is this?”\text{“Which cell, tile, or document region is this?”}

Impact 4: It can blur or distort learned positional patterns

Upsampling stretches the PE field. Downsampling compresses it.

With bicubic or bilinear interpolation, the new PE table is a smoothed version of the old table. That can be helpful if the learned PE is smooth, but harmful if the model learned sharp position-specific features.

For example, if a model learned special behavior for borders, corners, or certain image zones, interpolation may weaken or distort those cues.


Impact 5: It is not the same as image resizing

There are two different operations:

Image resizing:change the image to the training resolution\textbf{Image resizing:} \quad \text{change the image to the training resolution} PE interpolation:keep the new image resolution and resize the PE table\textbf{PE interpolation:} \quad \text{keep the new image resolution and resize the PE table}

If resizing a 384×384384 \times 384 image down to 224×224224 \times 224, keep the original number of tokens, but lose image detail.

If keeping 384×384384 \times 384 and interpolate PE, preserve more visual detail, but increase the number of tokens and change the positional prior.

So PE interpolation is not just a convenience. It changes both computation and representation.


4. Is PE interpolation always effective?

No.

It is often effective for moderate resolution changes, especially when followed by fine-tuning. This is why it became common in ViT-style image classification. The original ViT paper even says it is often beneficial to fine-tune at higher resolution than pretraining, while keeping patch size fixed and interpolating positional embeddings.

But it is not guaranteed. Recent vision PE work explicitly states that both absolute positional embedding and relative position bias can work well at fixed resolution but struggle with resolution changes, and that this can degrade performance in multi-resolution recognition, object detection, and segmentation.

So the correct statement is:

PE interpolation is a useful approximation, not a universal solution.\boxed{ \text{PE interpolation is a useful approximation, not a universal solution.} }

5. When it works best

PE interpolation can work best when all of these are true:

ConditionWhy it helps
The PE is learned absolute PEThis is the main case where table-size mismatch occurs
The resolution change is moderateInterpolation stays close to the training grid
The aspect ratio is similarThe spatial prior is not severely warped
The task is image-level classificationExact pixel-level geometry is less critical
Fine-tune at the target resolutionThe model can adapt to the interpolated PE
The patch size is unchangedOnly the grid changes, not the patch encoder

Example: taking a ViT pretrained at 224×224224 \times 224, using patch size 16, and fine-tuning at 384×384384 \times 384. That changes the grid from 14×1414 \times 14 to 24×2424 \times 24. PE interpolation is usually a reasonable starting point.


6. When it becomes risky

PE interpolation becomes risky when the new resolution is far from training resolution.

For example:

224×224→1024×1024224 \times 224 \rightarrow 1024 \times 1024

This creates many positions whose PE values are extrapolated by smooth interpolation from a much coarser table. The model may process the longer sequence, but its positional behavior is no longer the same as what it learned.

It is also risky when the aspect ratio changes:

224×224→512×1024224 \times 224 \rightarrow 512 \times 1024

because the PE field is stretched differently along height and width.

It is especially risky when the task requires dense spatial precision. Swin Transformer’s ablation is very relevant here: relative position bias improved classification, object detection, and segmentation compared with no PE and with absolute PE; the authors also found that absolute position embedding slightly improved ImageNet classification but hurt COCO detection and ADE20K semantic segmentation.

That result does not prove that all absolute PE interpolation is bad for dense tasks, but it does show the broader point:

classification tolerates absolute position artifacts more than dense prediction does.\text{classification tolerates absolute position artifacts more than dense prediction does.}

7. Can its impact be ignored?

Usually, no.

For a learned absolute PE model, if the input grid changes, the PE table is part of the model’s input. Changing it changes the first-layer token representation and therefore every later attention computation.

So its impact should not be ignored in:

  • object detection
  • segmentation
  • keypoint detection
  • depth estimation
  • OCR/layout analysis
  • medical imaging
  • remote sensing
  • robotics or calibrated vision

It can sometimes be treated as a minor engineering detail for image classification if the resolution change is small and the model is fine-tuned. But even then, it is better to report it, because it affects reproducibility.

A good experimental comparison should include:

  • training resolution
  • inference/fine-tuning resolution
  • patch size
  • PE interpolation method: bicubic, bilinear, nearest, etc.
  • whether PE was frozen or fine-tuned

8. Important distinction: image-size change vs patch-size change

Changing image size while keeping patch size fixed is one problem.

Changing patch size is a harder problem.

If changing from patch size 16 to patch size 8, both the number of tokens and the patch embedding kernel change. PE interpolation alone does not solve that.

FlexiViT is useful evidence here. The paper notes that standard pretrained ViTs are not flexible when evaluated at different patch sizes; even after resizing patch embedding weights and position embeddings, performance rapidly degrades as inference-time patch size moves away from the training patch size. FlexiViT addresses this by training with variable patch sizes. (CVF Open Access)

So:

PE interpolation can adapt the position table, but it cannot magically adapt the whole model.\boxed{ \text{PE interpolation can adapt the position table, but it cannot magically adapt the whole model.} }

9. PE interpolation for relative position bias is also not always safe

In windowed models like Swin, the issue is often not global image size, but window size.

Original Swin uses a learned relative position bias table for a fixed window size. If changing the window size, it may also interpolate the relative bias table. But Swin V2 found that directly transferring to larger image/window resolutions with bicubic interpolation could significantly reduce accuracy, so it introduced a log-spaced continuous position bias to transfer more smoothly across window sizes.

So even for relative PE:

interpolation is not automatically reliable.\text{interpolation is not automatically reliable.}

10. In what image tasks should PE interpolation not be used?

I would phrase this carefully:

PE interpolation should not be used blindly as the only resolution-adaptation method when the task depends on precise geometry, scale, or dense spatial output.

Avoid naive learned-APE interpolation for dense prediction

For tasks like:

  • semantic segmentation
  • instance segmentation
  • object detection
  • keypoint detection
  • depth estimation
  • optical flow

I should prefer one of these:

  • relative position bias
  • conditional/convolutional PE
  • 2D RoPE
  • continuous position bias
  • multi-resolution training
  • fine-tuning at the target resolution

Naively interpolating a classifier’s learned absolute PE and attaching a dense prediction head can work as a baseline, but it should not be trusted without validation.


Avoid it when absolute physical coordinates matter

For medical images, satellite images, microscopy, robotics, and calibrated camera systems, positions may have physical meaning.

For example:

1 patch=0.5 mm\text{1 patch} = 0.5\text{ mm}

or

1 pixel=specific geographic area\text{1 pixel} = \text{specific geographic area}

Interpolating PE according to normalized image coordinates may destroy that physical interpretation. In these cases, use explicit coordinate encodings in physical units, train at the target resolution, or use a PE mechanism designed for scale/coordinate consistency.


Avoid it as a substitute for multi-resolution training

If your deployment input sizes vary widely, do not rely on one PE table plus interpolation.

Use multi-resolution training or a position mechanism designed for variable resolution. CAPE, for example, was proposed partly because absolute positional embeddings are simple but have generalization issues under longer or changed input lengths, while relative positions are more robust but more complex. (OpenReview)


Avoid it when the model does not use learned absolute PE

If the model uses:

2D sinusoidal PE\text{2D sinusoidal PE} RoPE\text{RoPE} relative position bias\text{relative position bias} continuous position bias\text{continuous position bias} convolutional/conditional PE\text{convolutional/conditional PE}

then there may be no learned absolute table to interpolate. In those cases, it is usually better to recompute the PE at the new coordinates rather than interpolate an old table.


11. Suggested practical rule

Use this rule:

For classification, PE interpolation is usually acceptable.\boxed{ \text{For classification, PE interpolation is usually acceptable.} } For dense spatial tasks, treat PE interpolation as a risky approximation.\boxed{ \text{For dense spatial tasks, treat PE interpolation as a risky approximation.} } For large resolution/aspect-ratio changes, validate it experimentally.\boxed{ \text{For large resolution/aspect-ratio changes, validate it experimentally.} } For variable-resolution deployment, prefer relative, rotary, continuous, or multi-resolution-trained PE.\boxed{ \text{For variable-resolution deployment, prefer relative, rotary, continuous, or multi-resolution-trained PE.} }

The main idea is:

PE interpolation is not merely resizing a parameter. It is resizing the model’s learned spatial coordinate system.

That coordinate-system change can be benign for image-level recognition, but it can be quite important for tasks where the exact location, boundary, scale, or geometry matters.