Technical note
LayerNorm in ViT Models
An introduction to LayerNorm used in ViT and CLIP Models.
A typical ViT applies LayerNorm to the CLS embedding before projection:
x = transformer(x) # [B, N+1, D]
x = x[:, 0, :] # [B, D]
x = layer_norm(x) # [B, D]
x = x @ proj
LayerNorm operates independently on each sample across its feature dimensions.
Suppose
Each batch element contains one CLS embedding:
LayerNorm computes the mean and variance of those values:
and
Each feature is then normalized:
Finally, LayerNorm applies learned parameters :
LayerNorm does not normalize across the batch dimension . It normalizes across the embedding dimension , independently for each CLS token.
For example, let
and
The two rows are normalized independently. Their means are
Ignoring , , and for simplicity, both rows become approximately
When is negligible, the normalized values are invariant to a positive uniform rescaling and a uniform offset of an embedding. The learned and parameters can subsequently apply a separate scale and bias to each feature.
This illustrates an important property of LayerNorm: it removes the overall offset and scale of each embedding, while preserving the relative pattern among its dimensions.
PyTorch behavior
CLIP and ViT implementations commonly define the layer as
ln_post = nn.LayerNorm(D)
For an input with shape [B, D], PyTorch treats the final dimension as the normalized shape. Conceptually, the operation is
for b in range(B):
mean = x[b].mean()
var = x[b].var(unbiased=False)
x_norm[b] = (x[b] - mean) / sqrt(var + eps)
output[b] = gamma * x_norm[b] + beta
Both gamma and beta have shape [D].
In the above code snippet, gamma * x_norm[b], * means element-wise multiplication. All three vectors have shape [D]:
Each feature has its own learned scale and bias . Matrix multiplication would use @ instead.
Why is this useful before the CLIP projection?
After the transformer, the CLS token may have arbitrary magnitude and offset:
LayerNorm normalizes this representation and applies its learned affine transformation before projection:
If
then
LayerNorm(D) does not normalize dimension globally across the tensor. It independently normalizes each vector in the final dimension:
normalize this way →
x = [
[ d1, d2, d3, ... dD ], ← image 1 CLS token
[ d1, d2, d3, ... dD ], ← image 2 CLS token
[ d1, d2, d3, ... dD ], ← image 3 CLS token
...
]
LayerNorm before CLS extraction
Applying the same LayerNorm(D) before extracting the CLS token
x = transformer(x) # [B, N+1, D]
x = layer_norm(x) # [B, N+1, D]
independently normalizes every token’s -dimensional embedding. A tensor with shape
therefore contains separate vectors to normalize. For a standard LayerNorm(D),
layer_norm(x)[:, 0, :]
and
layer_norm(x[:, 0, :])
produce the same normalized CLS embeddings. This equivalence helps clarify LayerNorm placement within ViT architectures.