Technical note
Positional Embedding Intermediate II
A introduction to PE interpolation/extrapolation.
PE interpolation is useful, but it is not a harmless detail. It is a spatial-prior transformation. Sometimes it is the right engineering trick; sometimes it silently changes the model’s geometry.
1. What PE interpolation actually does
Assume a ViT was trained on images of size
with patch size
so its patch grid is
and its learned absolute positional embedding has shape
where the extra is usually the class token.
Now suppose we use a new image size
with the same patch size . The new grid is
The old positional table no longer has the right number of patch positions. So the usual trick is:
The class-token PE is usually kept unchanged.
This is exactly the workaround used in ViT-style models: ViT can handle arbitrary sequence length up to memory limits, but the pretrained position embeddings may no longer be meaningful when the image resolution changes, so the authors perform 2D interpolation of the pretrained position embeddings according to their original image locations. They also note that patch extraction and this PE interpolation are almost the only places where 2D image structure is manually injected into vanilla ViT.
2. The hidden assumption: learned PE is a smooth spatial field
PE interpolation treats the learned position table as if it were samples from a smooth function:
where are 2D image coordinates.
So interpolation assumes:
This is often reasonable. The original ViT paper observed that learned position embeddings tend to encode image distance: closer patches often have more similar position embeddings, patches in the same row or column can have similar embeddings, and sinusoidal-like structures sometimes appear for larger grids.
But this is still an assumption, not a guarantee. If the learned PE table contains high-frequency, task-specific, or dataset-specific spatial patterns, interpolation may blur, stretch, or distort them.
3. Main impacts of PE interpolation
Impact 1: It enables checkpoint reuse at a new resolution
This is the positive impact. Without interpolation, a learned absolute PE table trained on a patch grid cannot be directly added to a patch grid.
So PE interpolation lets us reuse a pretrained ViT at higher or lower image resolution. This is why libraries expose switches such as interpolate_pos_encoding; Hugging Face’s ViT documentation notes that fine-tuning at higher resolution is possible by setting this option, which interpolates the pretrained position embeddings. (Hugging Face)
Impact 2: It changes the model’s spatial prior
In additive absolute PE, the token becomes
After interpolation, the content token may be produced by the same patch encoder, but the position vector is different from anything the model saw during pretraining.
Then:
So interpolation affects attention through query/key/value projections. It does not merely “fill missing positions.” It changes the geometry of attention.
Impact 3: It rescales position
This is subtle and important.
Suppose the model was trained on a grid and I use it on a grid. Interpolation usually treats the old PE table as covering the whole image from top to bottom and left to right.
So the center of the new image receives a PE similar to the center of the old image.
That means interpolation usually preserves normalized image coordinates, not absolute patch-index distance.
This is good when the task is scale-normalized classification:
It may be less good when the task depends on exact patch-level geometry:
Impact 4: It can blur or distort learned positional patterns
Upsampling stretches the PE field. Downsampling compresses it.
With bicubic or bilinear interpolation, the new PE table is a smoothed version of the old table. That can be helpful if the learned PE is smooth, but harmful if the model learned sharp position-specific features.
For example, if a model learned special behavior for borders, corners, or certain image zones, interpolation may weaken or distort those cues.
Impact 5: It is not the same as image resizing
There are two different operations:
If resizing a image down to , keep the original number of tokens, but lose image detail.
If keeping and interpolate PE, preserve more visual detail, but increase the number of tokens and change the positional prior.
So PE interpolation is not just a convenience. It changes both computation and representation.
4. Is PE interpolation always effective?
No.
It is often effective for moderate resolution changes, especially when followed by fine-tuning. This is why it became common in ViT-style image classification. The original ViT paper even says it is often beneficial to fine-tune at higher resolution than pretraining, while keeping patch size fixed and interpolating positional embeddings.
But it is not guaranteed. Recent vision PE work explicitly states that both absolute positional embedding and relative position bias can work well at fixed resolution but struggle with resolution changes, and that this can degrade performance in multi-resolution recognition, object detection, and segmentation.
So the correct statement is:
5. When it works best
PE interpolation can work best when all of these are true:
| Condition | Why it helps |
|---|---|
| The PE is learned absolute PE | This is the main case where table-size mismatch occurs |
| The resolution change is moderate | Interpolation stays close to the training grid |
| The aspect ratio is similar | The spatial prior is not severely warped |
| The task is image-level classification | Exact pixel-level geometry is less critical |
| Fine-tune at the target resolution | The model can adapt to the interpolated PE |
| The patch size is unchanged | Only the grid changes, not the patch encoder |
Example: taking a ViT pretrained at , using patch size 16, and fine-tuning at . That changes the grid from to . PE interpolation is usually a reasonable starting point.
6. When it becomes risky
PE interpolation becomes risky when the new resolution is far from training resolution.
For example:
This creates many positions whose PE values are extrapolated by smooth interpolation from a much coarser table. The model may process the longer sequence, but its positional behavior is no longer the same as what it learned.
It is also risky when the aspect ratio changes:
because the PE field is stretched differently along height and width.
It is especially risky when the task requires dense spatial precision. Swin Transformer’s ablation is very relevant here: relative position bias improved classification, object detection, and segmentation compared with no PE and with absolute PE; the authors also found that absolute position embedding slightly improved ImageNet classification but hurt COCO detection and ADE20K semantic segmentation.
That result does not prove that all absolute PE interpolation is bad for dense tasks, but it does show the broader point:
7. Can its impact be ignored?
Usually, no.
For a learned absolute PE model, if the input grid changes, the PE table is part of the model’s input. Changing it changes the first-layer token representation and therefore every later attention computation.
So its impact should not be ignored in:
- object detection
- segmentation
- keypoint detection
- depth estimation
- OCR/layout analysis
- medical imaging
- remote sensing
- robotics or calibrated vision
It can sometimes be treated as a minor engineering detail for image classification if the resolution change is small and the model is fine-tuned. But even then, it is better to report it, because it affects reproducibility.
A good experimental comparison should include:
- training resolution
- inference/fine-tuning resolution
- patch size
- PE interpolation method: bicubic, bilinear, nearest, etc.
- whether PE was frozen or fine-tuned
8. Important distinction: image-size change vs patch-size change
Changing image size while keeping patch size fixed is one problem.
Changing patch size is a harder problem.
If changing from patch size 16 to patch size 8, both the number of tokens and the patch embedding kernel change. PE interpolation alone does not solve that.
FlexiViT is useful evidence here. The paper notes that standard pretrained ViTs are not flexible when evaluated at different patch sizes; even after resizing patch embedding weights and position embeddings, performance rapidly degrades as inference-time patch size moves away from the training patch size. FlexiViT addresses this by training with variable patch sizes. (CVF Open Access)
So:
9. PE interpolation for relative position bias is also not always safe
In windowed models like Swin, the issue is often not global image size, but window size.
Original Swin uses a learned relative position bias table for a fixed window size. If changing the window size, it may also interpolate the relative bias table. But Swin V2 found that directly transferring to larger image/window resolutions with bicubic interpolation could significantly reduce accuracy, so it introduced a log-spaced continuous position bias to transfer more smoothly across window sizes.
So even for relative PE:
10. In what image tasks should PE interpolation not be used?
I would phrase this carefully:
PE interpolation should not be used blindly as the only resolution-adaptation method when the task depends on precise geometry, scale, or dense spatial output.
Avoid naive learned-APE interpolation for dense prediction
For tasks like:
- semantic segmentation
- instance segmentation
- object detection
- keypoint detection
- depth estimation
- optical flow
I should prefer one of these:
- relative position bias
- conditional/convolutional PE
- 2D RoPE
- continuous position bias
- multi-resolution training
- fine-tuning at the target resolution
Naively interpolating a classifier’s learned absolute PE and attaching a dense prediction head can work as a baseline, but it should not be trusted without validation.
Avoid it when absolute physical coordinates matter
For medical images, satellite images, microscopy, robotics, and calibrated camera systems, positions may have physical meaning.
For example:
or
Interpolating PE according to normalized image coordinates may destroy that physical interpretation. In these cases, use explicit coordinate encodings in physical units, train at the target resolution, or use a PE mechanism designed for scale/coordinate consistency.
Avoid it as a substitute for multi-resolution training
If your deployment input sizes vary widely, do not rely on one PE table plus interpolation.
Use multi-resolution training or a position mechanism designed for variable resolution. CAPE, for example, was proposed partly because absolute positional embeddings are simple but have generalization issues under longer or changed input lengths, while relative positions are more robust but more complex. (OpenReview)
Avoid it when the model does not use learned absolute PE
If the model uses:
then there may be no learned absolute table to interpolate. In those cases, it is usually better to recompute the PE at the new coordinates rather than interpolate an old table.
11. Suggested practical rule
Use this rule:
The main idea is:
PE interpolation is not merely resizing a parameter. It is resizing the model’s learned spatial coordinate system.
That coordinate-system change can be benign for image-level recognition, but it can be quite important for tasks where the exact location, boundary, scale, or geometry matters.