Technical note
YOLO v9 Series: CSPNet -> ELAN -> GELAN
A deep study into YOLO v9's model architecture.
is as a sequence of increasingly broader gradient-path design ideas.
- CSPNet: cross stage partial network (arXiv)
- ELAN: Efficient Layer Aggregation Network (arXiv)
- GELAN: Generalized Efficient Layer Aggregation Network (arXiv) (GitHub)
CSPNet asks: How should I split and merge features inside a stage so I do less redundant computation and get more diverse gradients?
ELAN asks: As I make a network deeper, how do I keep gradient paths efficient enough that the network still trains well?
GELAN asks: Can I generalize those ideas into a reusable aggregation framework that works with different computational blocks and scales efficiently?
They are primarily CNN architecture design strategies, although they became especially visible through YOLO.
What problem they are all addressing:
Consider a conventional deep CNN:
Input
│
Conv
│
Block
│
Block
│
Block
│
Block
│
Output
There are two paths. The obvious one is the forward feature path:
But training depends on the reverse path:
CSPNet argues that architecture design should consider not only forward feature extraction, but also the structure, diversity, and length of backward gradient paths. CSPNet addresses this mainly at the stage level, while ELAN extends gradient-path planning to the network level.
I. CSPNet: Cross Stage Partial Network
CSPNet paper identified three practical issues in conventional CNNs:
- duplicated gradient information,
- excessive computation / computational bottlenecks,
- excessive memory traffic.
The core observation came particularly from DenseNet-like architectures.
DenseNet intentionally reuses features heavily:
That is very effective for feature reuse, but it also means many layers receive highly overlapping gradient information.
The CSPNet authors call this duplicate gradient information.
Suppose we have an input feature tensor:
Instead of sending all channels through an expensive block, CSPNet splits the channels:
Typically, think roughly:
Then:
┌────────────────────────┐
│ │
X ─────── split ────┤ Xa → heavy block ─────┤
│ ├── concat → fusion → output
└ Xb → bypass ───────────┘
Or mathematically:
Only part of the channels go through the expensive transformation .
The other channels effectively establish a cross-stage path.

The original paper describes CSP as splitting a stage’s base feature map into two parts, sending one through the computational block and connecting the other more directly to the end of the stage.
Why?
There are two benefits.
First: lower computation
Imagine a conventional block operating on all channels.
For a standard convolution:
Reduce the channels entering a deep block from to approximately , its expensive internal convolutions become substantially cheaper.
For example, with equal input/output width:
becomes approximately:
It is NOT that the whole CSP stage is therefore four times cheaper, because:
- split/fusion convolutions,
- the bypass,
- concatenation,
- output projection.
But the expensive inner block processes fewer channels. Empirically, the CSPNet paper reports roughly 10–20% reduction in computation after applying CSP to architectures such as ResNet, ResNeXt and DenseNet, while maintaining or improving accuracy.
Second: more diverse gradient paths
Consider a non-CSP block:
X
│
Block1
│
Block2
│
Block3
│
Y
Every feature essentially goes through similar transformations.
CSP creates two routes:
┌─ deep transformation ──────┐
X ──────┤ ├─ Y
└─ shorter partial route ────┘
So gradients propagate through different paths.
This reduces the tendency for every layer in the stage to repeatedly learn from extremely similar gradient signals. The authors describe CSP as increasing gradient combinations while avoiding duplicate gradient information.
CSP is a general CNN stage-design method. It can be used at anywhere that have repeated convolutional blocks.
II. ELAN: Efficient Layer Aggregation Network
CSP solves stage-level issues very effectively. But when keep making the model deeper, another question appears:
Even if each stage is individually efficient, what happens to gradient propagation across the entire network?
This motivates ELAN.
ELAN appeared in the work leading to YOLOv7. The deeper theoretical treatment describes it as a network-level gradient-path design.
Its central problem is:
As you scale a network deeper, the effective gradient path through some layers can become progressively worse, causing convergence degradation.
This is subtly different from vanishing gradients in the classic sense.
They are particularly concerned with the structure of:
- shortest gradient paths,
- longest gradient paths,
- aggregation paths,
- transition layers.
Why ordinary stacking can become problematic?
Consider repeatedly stacked blocks, each transition layer becomes another operation through which gradients need to propagate. As the architecture scales deeper, even the shortest gradient path for earlier layers can get longer. The ELAN work argues that excessively lengthening these paths can hurt trainability.
A simplified ELAN block looks like this:
Input
│
├──── Conv ─────────────────────────┐
│ │
└──── Conv → Conv → Conv → Conv ────┤
│
concatenate ◄─────┘
│
1×1 Conv
│
Output
A more faithful conceptual representation is:
branch 1 ────────────────┐
Input ─ split ┤ │
branch 2 → L1 ───────────┤
│ │
L2 ────────────┤
│ │
L3 ────────────┤
↓
concatenate
│
output
The network collects outputs from different depths. This is layer aggregation.
ELAN carefully restructures aggregation to control gradient path lengths. The authors describe ELAN as combining ideas from VoVNet and CSPNet while using a strategy called stack in computational block to preserve efficient gradient propagation as depth increases.
Why aggregation helps?
Suppose:
Instead of only using:
ELAN aggregates:
That gives the output direct access to representations generated at different transformation depths.
Backward propagation also gains multiple paths:
Loss
│
├────────→ X1
├────────→ X2
├────────→ X3
└────────→ X4
rather than requiring every gradient to traverse all four operations.
This sounds like DenseNet. But they’re different.
DenseNet approximately does:
So every layer receives all previous features. This gives strong feature reuse but can become costly because feature width grows.
ELAN instead aggregates strategically at the block output, rather than feeding every previous output into every future layer.
ELAN tries to retain aggregation benefits without DenseNet’s rapidly growing connection/computation complexity.
CSP vs ELAN
A useful mental model is:
| Property | CSPNet | ELAN |
|---|---|---|
| Primary design level | Stage | Whole network |
| Main concern | Duplicate gradient information | Gradient-path length during scaling |
| Major mechanism | Split + partial processing + merge | Multi-depth layer aggregation |
| Typical benefit | Lower compute / memory + gradient diversity | Better trainability of deep networks |
| Feature reuse | Partial | Aggregated across depths |
| Main historical use | CSPDarknet, YOLOv4 etc. | YOLOv7 |
| General-purpose? | Yes | Yes |
But these two ideas are complementary, not competitors.
III. GELAN: Generalized Efficient Layer Aggregation Network
YOLOv9 introduced GELAN. The YOLOv9 authors describe GELAN as a new lightweight architecture based on gradient path planning, capable of achieving strong parameter utilization even using ordinary convolutions.
ELAN had a strong gradient path design, but it was relatively tied to a specific kind of computational structure.
The authors wanted something more general:
Can we preserve ELAN-style gradient-path properties while allowing arbitrary computational blocks to be plugged into the architecture?
For example, instead of requiring:
ELAN
└── standard Conv blocks
you might want:
GELAN
├── Conv
├── RepConv
├── CSP block
├── Bottleneck
└── potentially another compatible block
That is the “Generalized” part.
GELAN architecture concept
A simplified GELAN-style block can be viewed like this:
Input
│
├──────────── partial branch ───────────────────┐
│ │
└→ transform → block → block → block ───────────┤
│ │ │ │ │
└─────────┴───────┴───────┴───────────────┘
│
concatenate
│
fuse
│
Output
Notice two ideas (CSP and ELAN) at once:
- CSP idea: Only a portion of feature flow needs to experience all expensive computation.
- ELAN idea: Intermediate outputs from different depths are aggregated.
So roughly something like:
with
then
That equation captures much of the intuition.
YOLOv9 block: RepNCSPELAN4
In the YOLOv9 repository, I frequently see a block named: RepNCSPELAN4.
That name tells a lot:
Rep
+
NCSP
+
ELAN
It combines:
- reparameterized convolution ideas,
- CSP-like partial structure,
- ELAN-style aggregation.
The repository explicitly contains both CSP and RepNCSP building blocks, which are used inside the larger YOLOv9 architecture.
A simplified decomposition is:
Input
│
1×1 Conv
│
split
├─────────────────────────────┐
│ │
└→ RepNCSP → Conv → RepNCSP ──┤
│
concatenate
│
1×1 Conv
│
output
The exact YOLOv9 implementation has model-specific details, but this gives the core structural idea.
The main improvement is generality and parameter utilization.
ELAN:
Here is an efficient layer-aggregation architecture with favorable gradient paths.
GELAN:
Let’s preserve that aggregation/gradient-path logic while allowing the internal computational block to be replaced or redesigned.
That makes it easier to adapt the architecture to:
- model size,
- hardware,
- computational operator,
- latency constraints.
GELAN and computation load
GELAN is not necessarily the lowest-FLOP structure.
For example:
Depthwise convolution
can have dramatically fewer theoretical FLOPs than conventional convolution.
A standard convolution costs approximately:
A depthwise + pointwise convolution costs approximately:
For sufficiently large , this is much cheaper theoretically. Depthwise convolution has low theoretical FLOPs, but on some hardware it can be:
- memory-bandwidth bound,
- less efficient at utilizing GPU compute units,
- less friendly to certain accelerators.
GELAN achieves better parameter utilization even though it relies largely on normal convolution. Regular convolution can still be attractive because standard convolutions are extremely optimized on GPUs.
The progression in one picture

Think of a conventional CNN:
A → B → C → D
CSPNet = stage-level design
Split the feature stream:
┌→ B → C ─┐
A ─────┤ ├→ D
└─────────┘
Goal:
ELAN = network-level design
Aggregate features from different computational depths:
A → B → C → D
│ │ │
└───┴───┴──→ concat → E
Goal:
GELAN = generalized architectural framework
Combine partial paths and multi-depth aggregation:
┌──────────────────────────┐
A → split → B → C → D │
│ │ │ │
└────┴───┴─────────────┤
↓
concat
↓
output
with allowed to be more general computational blocks.
Goal:
Their computation characteristics side-by-side
| CSPNet | ELAN | GELAN | |
|---|---|---|---|
| Expensive channels partially bypassed? | Yes | Sometimes | Yes / generalized |
| Multi-depth aggregation? | Limited | Yes | Yes |
| Primary compute objective | Reduce redundant work | Efficient deep scaling | High parameter utilization |
| Memory traffic concern | Strong | Moderate | Strong/practical |
| Standard convolution friendly | Yes | Yes | Very much so |
| Typical FLOP effect | Usually reduces | Depends on configuration | Depends on configuration |
| Main training benefit | Gradient diversity | Short/controlled gradient paths | Both |
| Hardware-conscious | Yes | Yes | Yes |