Technical note

YOLO v9 Series: CSPNet -> ELAN -> GELAN

A deep study into YOLO v9's model architecture.

  • Deep Learning
  • YOLO

CSPNet→ELAN→GELAN\boxed{\text{CSPNet} \rightarrow \text{ELAN} \rightarrow \text{GELAN}} is as a sequence of increasingly broader gradient-path design ideas.

  1. CSPNet: cross stage partial network (arXiv)
  2. ELAN: Efficient Layer Aggregation Network (arXiv)
  3. GELAN: Generalized Efficient Layer Aggregation Network (arXiv) (GitHub)

CSPNet asks: How should I split and merge features inside a stage so I do less redundant computation and get more diverse gradients?

ELAN asks: As I make a network deeper, how do I keep gradient paths efficient enough that the network still trains well?

GELAN asks: Can I generalize those ideas into a reusable aggregation framework that works with different computational blocks and scales efficiently?

They are primarily CNN architecture design strategies, although they became especially visible through YOLO.

What problem they are all addressing:

Consider a conventional deep CNN:

Input
  │
 Conv
  │
 Block
  │
 Block
  │
 Block
  │
 Block
  │
Output

There are two paths. The obvious one is the forward feature path:

x0→x1→x2→⋯→xnx_0 \rightarrow x_1 \rightarrow x_2 \rightarrow \cdots \rightarrow x_n

But training depends on the reverse path:

∂L∂xn→∂L∂xn−1→⋯→∂L∂x0.\frac{\partial L}{\partial x_n} \rightarrow \frac{\partial L}{\partial x_{n-1}} \rightarrow \cdots \rightarrow \frac{\partial L}{\partial x_0}.

CSPNet argues that architecture design should consider not only forward feature extraction, but also the structure, diversity, and length of backward gradient paths. CSPNet addresses this mainly at the stage level, while ELAN extends gradient-path planning to the network level.

I. CSPNet: Cross Stage Partial Network

CSPNet paper identified three practical issues in conventional CNNs:

  • duplicated gradient information,
  • excessive computation / computational bottlenecks,
  • excessive memory traffic.

The core observation came particularly from DenseNet-like architectures.

DenseNet intentionally reuses features heavily: xl=Hl([x0,x1,…,xl−1]).x_l = H_l([x_0,x_1,\ldots,x_{l-1}]).

That is very effective for feature reuse, but it also means many layers receive highly overlapping gradient information.

The CSPNet authors call this duplicate gradient information.

Suppose we have an input feature tensor: X∈RH×W×C.X\in\mathbb{R}^{H\times W\times C}.

Instead of sending all CC channels through an expensive block, CSPNet splits the channels: X=[Xa,Xb].X=[X_a,X_b].

Typically, think roughly: Ca≈Cb≈C/2.C_a\approx C_b\approx C/2.

Then:

                    ┌────────────────────────┐
                    │                        │
X ─────── split ────┤ Xa → heavy block  ─────┤
                    │                        ├── concat → fusion → output
                    └ Xb → bypass ───────────┘

Or mathematically:

Ya=F(Xa)Y_a = F(X_a)

Y=G([Ya,Xb]).Y = G([Y_a,X_b]).

Only part of the channels go through the expensive transformation FF.

The other channels effectively establish a cross-stage path.

The original paper describes CSP as splitting a stage’s base feature map into two parts, sending one through the computational block and connecting the other more directly to the end of the stage.

Why?

There are two benefits.

First: lower computation

Imagine a conventional block operating on all CC channels.

For a standard convolution: FLOPs∝HWCinCoutK2.\text{FLOPs}\propto H W C_{in}C_{out}K^2.

Reduce the channels entering a deep block from CC to approximately C/2C/2, its expensive internal convolutions become substantially cheaper.

For example, with equal input/output width: C2C^2

becomes approximately: (C2)2=C24.\left(\frac C2\right)^2=\frac{C^2}{4}.

It is NOT that the whole CSP stage is therefore four times cheaper, because:

  • split/fusion convolutions,
  • the bypass,
  • concatenation,
  • output projection.

But the expensive inner block processes fewer channels. Empirically, the CSPNet paper reports roughly 10–20% reduction in computation after applying CSP to architectures such as ResNet, ResNeXt and DenseNet, while maintaining or improving accuracy.

Second: more diverse gradient paths

Consider a non-CSP block:

X
│
Block1
│
Block2
│
Block3
│
Y

Every feature essentially goes through similar transformations.

CSP creates two routes:

        ┌─ deep transformation ──────┐
X ──────┤                            ├─ Y
        └─ shorter partial route ────┘

So gradients propagate through different paths.

This reduces the tendency for every layer in the stage to repeatedly learn from extremely similar gradient signals. The authors describe CSP as increasing gradient combinations while avoiding duplicate gradient information.

CSP is a general CNN stage-design method. It can be used at anywhere that have repeated convolutional blocks.

II. ELAN: Efficient Layer Aggregation Network

CSP solves stage-level issues very effectively. But when keep making the model deeper, another question appears:

Even if each stage is individually efficient, what happens to gradient propagation across the entire network?

This motivates ELAN.

ELAN appeared in the work leading to YOLOv7. The deeper theoretical treatment describes it as a network-level gradient-path design.

Its central problem is:

As you scale a network deeper, the effective gradient path through some layers can become progressively worse, causing convergence degradation.

This is subtly different from vanishing gradients in the classic sense.

They are particularly concerned with the structure of:

  • shortest gradient paths,
  • longest gradient paths,
  • aggregation paths,
  • transition layers.

Why ordinary stacking can become problematic?

Consider repeatedly stacked blocks, each transition layer becomes another operation through which gradients need to propagate. As the architecture scales deeper, even the shortest gradient path for earlier layers can get longer. The ELAN work argues that excessively lengthening these paths can hurt trainability.

A simplified ELAN block looks like this:

Input
  │
  ├──── Conv ─────────────────────────┐
  │                                   │
  └──── Conv → Conv → Conv → Conv ────┤
                                      │
                    concatenate ◄─────┘
                            │
                         1×1 Conv
                            │
                          Output

A more faithful conceptual representation is:

               branch 1 ────────────────┐
Input ─ split ┤                         │
               branch 2 → L1 ───────────┤
                         │              │
                         L2 ────────────┤
                         │              │
                         L3 ────────────┤
                                        ↓
                                   concatenate
                                        │
                                     output

The network collects outputs from different depths. This is layer aggregation.

ELAN carefully restructures aggregation to control gradient path lengths. The authors describe ELAN as combining ideas from VoVNet and CSPNet while using a strategy called stack in computational block to preserve efficient gradient propagation as depth increases.

Why aggregation helps?

Suppose:

X0→X1→X2→X3→X4.X_0 \rightarrow X_1 \rightarrow X_2 \rightarrow X_3 \rightarrow X_4.

Instead of only using:

Y=X4,Y=X_4,

ELAN aggregates:

Y=G([X1,X2,X3,X4]).Y=G([X_1,X_2,X_3,X_4]).

That gives the output direct access to representations generated at different transformation depths.

Backward propagation also gains multiple paths:

Loss
 │
 ├────────→ X1
 ├────────→ X2
 ├────────→ X3
 └────────→ X4

rather than requiring every gradient to traverse all four operations.

This sounds like DenseNet. But they’re different.

DenseNet approximately does:

Xl=Hl([X0,X1,…,Xl−1]).X_l = H_l([X_0,X_1,\ldots,X_{l-1}]).

So every layer receives all previous features. This gives strong feature reuse but can become costly because feature width grows.

ELAN instead aggregates strategically at the block output, rather than feeding every previous output into every future layer.

ELAN tries to retain aggregation benefits without DenseNet’s rapidly growing connection/computation complexity.

CSP vs ELAN

A useful mental model is:

PropertyCSPNetELAN
Primary design levelStageWhole network
Main concernDuplicate gradient informationGradient-path length during scaling
Major mechanismSplit + partial processing + mergeMulti-depth layer aggregation
Typical benefitLower compute / memory + gradient diversityBetter trainability of deep networks
Feature reusePartialAggregated across depths
Main historical useCSPDarknet, YOLOv4 etc.YOLOv7
General-purpose?YesYes

But these two ideas are complementary, not competitors.

III. GELAN: Generalized Efficient Layer Aggregation Network

YOLOv9 introduced GELAN. The YOLOv9 authors describe GELAN as a new lightweight architecture based on gradient path planning, capable of achieving strong parameter utilization even using ordinary convolutions.

GELAN≈CSP principles+ELAN principles+flexible computational blocks\boxed{\text{GELAN} \approx \text{CSP principles} + \text{ELAN principles} + \text{flexible computational blocks}}

ELAN had a strong gradient path design, but it was relatively tied to a specific kind of computational structure.

The authors wanted something more general:

Can we preserve ELAN-style gradient-path properties while allowing arbitrary computational blocks to be plugged into the architecture?

For example, instead of requiring:

ELAN
  └── standard Conv blocks

you might want:

GELAN
  ├── Conv
  ├── RepConv
  ├── CSP block
  ├── Bottleneck
  └── potentially another compatible block

That is the “Generalized” part.

GELAN architecture concept

A simplified GELAN-style block can be viewed like this:

Input
  │
  ├──────────── partial branch ───────────────────┐
  │                                               │
  └→ transform → block → block → block ───────────┤
        │         │       │       │               │
        └─────────┴───────┴───────┴───────────────┘
                          │
                      concatenate
                          │
                        fuse
                          │
                        Output

Notice two ideas (CSP and ELAN) at once:

  1. CSP idea: Only a portion of feature flow needs to experience all expensive computation.
  2. ELAN idea: Intermediate outputs from different depths are aggregated.

So roughly something like:

X=[Xa,Xb]X=[X_a,X_b]

with

Z1=F1(Xa),Z_1=F_1(X_a), Z2=F2(Z1),Z_2=F_2(Z_1), Z3=F3(Z2),Z_3=F_3(Z_2),

then

Y=G([Xb,Z1,Z2,Z3]).Y=G([X_b,Z_1,Z_2,Z_3]).

That equation captures much of the intuition.


YOLOv9 block: RepNCSPELAN4

In the YOLOv9 repository, I frequently see a block named: RepNCSPELAN4.

That name tells a lot:

Rep
+
NCSP
+
ELAN

It combines:

  • reparameterized convolution ideas,
  • CSP-like partial structure,
  • ELAN-style aggregation.

The repository explicitly contains both CSP and RepNCSP building blocks, which are used inside the larger YOLOv9 architecture.

A simplified decomposition is:

Input
 │
1×1 Conv
 │
split
 ├─────────────────────────────┐
 │                             │
 └→ RepNCSP → Conv → RepNCSP ──┤
                               │
                           concatenate
                               │
                            1×1 Conv
                               │
                             output

The exact YOLOv9 implementation has model-specific details, but this gives the core structural idea.

The main improvement is generality and parameter utilization.

ELAN:

Here is an efficient layer-aggregation architecture with favorable gradient paths.

GELAN:

Let’s preserve that aggregation/gradient-path logic while allowing the internal computational block to be replaced or redesigned.

That makes it easier to adapt the architecture to:

  • model size,
  • hardware,
  • computational operator,
  • latency constraints.

GELAN and computation load

GELAN is not necessarily the lowest-FLOP structure.

For example:

Depthwise convolution

can have dramatically fewer theoretical FLOPs than conventional convolution.

A standard 3×33\times3 convolution costs approximately:

HW(9CinCout).HW(9C_{in}C_{out}).

A depthwise + pointwise convolution costs approximately:

HW(9C+C2).HW(9C + C^2).

For sufficiently large CC, this is much cheaper theoretically. Depthwise convolution has low theoretical FLOPs, but on some hardware it can be:

  • memory-bandwidth bound,
  • less efficient at utilizing GPU compute units,
  • less friendly to certain accelerators.

GELAN achieves better parameter utilization even though it relies largely on normal convolution. Regular convolution can still be attractive because standard convolutions are extremely optimized on GPUs.

The progression in one picture

Refer to caption

Think of a conventional CNN:

A → B → C → D

CSPNet = stage-level design

Split the feature stream:

       ┌→ B → C ─┐
A ─────┤         ├→ D
       └─────────┘

Goal:

less redundant computation+more diverse gradients.\text{less redundant computation} + \text{more diverse gradients}.


ELAN = network-level design

Aggregate features from different computational depths:

A → B → C → D
    │   │   │
    └───┴───┴──→ concat → E

Goal:

efficient gradient paths as depth increases.\text{efficient gradient paths as depth increases}.


GELAN = generalized architectural framework

Combine partial paths and multi-depth aggregation:

      ┌──────────────────────────┐
A → split → B → C → D            │
          │    │   │             │
          └────┴───┴─────────────┤
                                 ↓
                              concat
                                 ↓
                               output

with B,C,DB,C,D allowed to be more general computational blocks.

Goal:

gradient-path efficiency+parameter efficiency+architectural flexibility.\text{gradient-path efficiency} + \text{parameter efficiency} + \text{architectural flexibility}.


Their computation characteristics side-by-side

CSPNetELANGELAN
Expensive channels partially bypassed?YesSometimesYes / generalized
Multi-depth aggregation?LimitedYesYes
Primary compute objectiveReduce redundant workEfficient deep scalingHigh parameter utilization
Memory traffic concernStrongModerateStrong/practical
Standard convolution friendlyYesYesVery much so
Typical FLOP effectUsually reducesDepends on configurationDepends on configuration
Main training benefitGradient diversityShort/controlled gradient pathsBoth
Hardware-consciousYesYesYes