Technical note

YOLO v9 Series: Deep Studies

A deep study into YOLO v9.

  • Deep Learning
  • YOLO

YOLOv9 showed that a detector can be improved not only by designing better forward features, but by explicitly designing how useful information reaches the loss and how reliable gradient information flows backward through the network.

Compared with prior YOLO generations, the main achievement of YOLOv9 was not simply “a better backbone” or another set of training tricks. Its central contribution was a new way to think about information loss and gradient quality during training, coupled with an architecture designed around efficient gradient flow.

The two key ideas are:

  1. Programmable Gradient Information (PGI) — a new training framework intended to preserve useful target information and provide more reliable gradients.
  2. GELAN (Generalized Efficient Layer Aggregation Network) — a new network architecture derived from ELAN/CSP concepts that improves parameter utilization and gives a strong accuracy/efficiency tradeoff.

Together, these produced YOLOv9. (arXiv)

1. PGI is probably the most important conceptual contribution

Previous YOLO versions mainly improved things such as backbone architecture, feature aggregation, label assignment, loss functions, auxiliary heads, reparameterization, and anchor-free detection.

Earlier YOLO development was often approximately:

How do we build a better feature extractor / neck / prediction head?

YOLOv9 asks another question:

Does the loss function still have enough information to generate the correct gradient for earlier layers?

So, YOLOv9 instead explicitly focuses on what the authors call the information bottleneck problem.

As features propagate through:

X→f1(X)→f2(f1(X))→⋯→Y^X \rightarrow f_1(X) \rightarrow f_2(f_1(X)) \rightarrow \cdots \rightarrow \hat Y

some information from the original image is inevitably discarded.

That is normally desirable to some degree: the network should compress irrelevant information. But the problem is that it can also discard information that is important for predicting YY.

The authors describe this approximately through the information-bottleneck relationship:

I(X,X)≥I(Y,X)≥I(Y,fθ(X))≥⋯≥I(Y,Y^)I(X,X) \ge I(Y,X) \ge I(Y,f_\theta(X)) \ge \cdots \ge I(Y,\hat Y)

The important quantity is not necessarily preserving all information in XX; it is preserving the part relevant to the target:

I(Y,X)I(Y,X)

If that information disappears during forward propagation, the loss is calculated using incomplete features, and consequently the back-propagated gradient itself can become misleading. (arXiv)

This is a subtle but important shift in perspective.

That leads to PGI.


2. Programmable Gradient Information

PGI adds an auxiliary training structure specifically designed to generate better gradient information.

Conceptually:

                        Main branch
Input ───────────────────────────────→ Prediction
   │                                      │
   │                                      ↓
   │                                    Loss
   │
   └──── Auxiliary reversible branch ────┘
                   ↑
          better gradient information

PGI has three main components:

  • the main branch
  • an auxiliary reversible branch
  • multi-level auxiliary information

The critical point is that the auxiliary reversible branch exists primarily during training.

At inference:

Input → Main YOLOv9 network → detections

The auxiliary branch can be removed.

Therefore PGI can improve training without adding corresponding inference cost.

Why the reversible branch?

Reversible architectures are one way of retaining information through a deep network. But making the entire inference network reversible would impose substantial computational overhead.

YOLOv9 instead makes the reversible path an auxiliary training branch.

So rather than forcing the deployed network to preserve everything, PGI essentially says:

During training, provide the network with enough additional information to generate good gradients, and then discard that machinery at inference.

The paper reports that directly incorporating reversible structures into inference could increase inference time substantially, while PGI avoids this penalty.


3. PGI improves on conventional deep supervision

This is particularly relevant if comparing YOLOv9 with YOLOv7, because YOLOv7 already used auxiliary heads/deep supervision.

Traditional deep supervision roughly looks like:

Backbone
   │
   ├──────── Auxiliary Head → Aux loss
   │
   ↓
 deeper layers
   │
   ↓
 Main Head → Main loss

This helps gradients reach earlier layers. But YOLOv9 points out a problem. Different feature pyramid levels specialize in objects of different scales:

P3 → small objects
P4 → medium objects
P5 → large objects

If an auxiliary head supervises one level independently, objects that belong to another level may effectively be treated as background. This can lead to what the paper calls error accumulation / broken information.

PGI therefore introduces multi-level auxiliary information, which combines gradient information from multiple prediction levels before propagating it back.

Conceptually:

 Small-object head ─┐
 Medium-object head ├──→ integration → auxiliary gradient
 Large-object head ─┘

instead of independently pushing potentially contradictory supervision into the network.

The authors’ ablation study is particularly interesting here. Conventional deep supervision actually hurt the smallest GELAN model, while PGI improved models across different scales. They conclude that PGI makes auxiliary supervision useful not just for very deep networks but also for shallower/lightweight ones.


4. Architectural contribution — GELAN

Generalized Efficient Layer Aggregation Network — GELAN

GELAN can roughly be viewed as an evolution of the architectural ideas behind:

CSPNet
   +
ELAN
   ↓
GELAN

ELAN was an important component of YOLOv7. GELAN generalizes the aggregation strategy so that different computational blocks can be inserted into the architecture while maintaining favorable gradient paths.

The paper describes it as allowing users to choose computational blocks according to the target inference hardware.

One of the interesting claims is that GELAN achieves excellent parameter utilization while relying on relatively ordinary convolution operations.

That is noteworthy because a common trend had been:

standard convolution
        ↓
depthwise separable convolution
        ↓
fewer parameters / FLOPs

YOLOv9 shows that clever network topology and gradient-path design can sometimes matter as much as replacing standard convolutions with theoretically cheaper operators.

The authors report that GELAN using conventional convolutions achieved better parameter utilization than several contemporary architectures using depth-wise convolution.


5. YOLOv9 is therefore really GELAN + PGI

A useful way to think about the framework is:

               YOLOv9
                  │
          ┌───────┴────────┐
          │                │
        GELAN             PGI
          │                │
     architecture       training
          │                │
  efficient feature    reliable
    aggregation        gradients

GELAN answers:

How should the network be constructed?

PGI answers:

How should information and gradients flow during training?


6. The performance improvements were substantial

Refer to the paper for detailed performance improvement summaries. Here I only keep highlight notes for quick digest.

On COCO, the paper reports the following example data:

ModelParamsFLOPsAP50:95
YOLOv736.9M104.7G51.2
YOLOv7 AF43.6M130.5G53.0
YOLOv8-X68.2M257.8G53.9
YOLOv9-C25.3M102.1G53.0
YOLOv9-E57.3M189.0G55.6

YOLOv9-C vs YOLOv7 AF

YOLOv7 AF:

43.6M,130.5G,53.0AP43.6M,\quad130.5G,\quad53.0 AP

YOLOv9-C:

25.3M,102.1G,53.0AP25.3M,\quad102.1G,\quad53.0 AP

  • 42% fewer parameters
  • 22% fewer FLOPs

YOLOv9-E vs YOLOv8-X

YOLOv8-X:

68.2M,257.8G,53.9AP68.2M,\quad257.8G,\quad53.9 AP

YOLOv9-E:

57.3M,189.0G,55.6AP57.3M,\quad189.0G,\quad55.6 AP

  • 16% fewer parameters
  • 27% fewer FLOPs
  • +1.7 AP

7. The ablation from YOLOv7 → YOLOv9 shows where the gains came from

The authors give a very useful progression:

ArchitectureParamsFLOPsAP
YOLOv736.9M104.7G51.2
+ anchor-free43.6M130.5G53.0
+ GELAN41.2M126.4G53.2
+ DHLC57.3M189.0G55.0
+ PGI57.3M189.0G55.6

It is also important to note that PGI adds roughly +0.6 AP+0.6 \text{ AP} without increasing the deployed model’s parameter count or inference FLOPs.


8. YOLOv9 vs the progression of previous YOLO versions

Now, I’d like to give a useful summary about progressive improvements along YOLO’s history for future reference.

GenerationMajor emphasis
YOLOv1Single-stage end-to-end detection
YOLOv2Better anchors, BN, multi-scale training
YOLOv3Multi-scale prediction + Darknet-53
YOLOv4Bag of freebies / specials, CSP, training improvements
YOLOv5Practical engineering/deployment ecosystem
YOLOv6Industrial efficiency and reparameterization
YOLOv7ELAN/E-ELAN, trainable bag-of-freebies, auxiliary heads
YOLOv8Anchor-free split head, C2f-style architecture, strong ecosystem
YOLOv9Information preservation + gradient quality through PGI; GELAN architecture

Roughly

YOLOv7 ideas→GELAN + PGI→YOLOv9\text{YOLOv7 ideas} \rightarrow \text{GELAN + PGI} \rightarrow \text{YOLOv9}

rather than:

Ultralytics YOLOv8→YOLOv9.\text{Ultralytics YOLOv8} \rightarrow \text{YOLOv9}.

They are different development lineages.

Advanced thinking: why trained network can discard important information

Even a well-trained network can discard target-relevant information because training only optimizes the final objective under the constraints of the architecture and optimization process; it does not guarantee that every intermediate representation preserves all information useful for the target.

well optimized≠information preserving≠well generalizing\boxed{ \text{well optimized} \neq \text{information preserving} \neq \text{well generalizing} }

This can be understood by two separate questions:

Did training converge well?\text{Did training converge well?}

Did the architecture preserve every useful signal while producing its representation?\text{Did the architecture preserve every useful signal while producing its representation?}

1. Many common network operations are intrinsically lossy

Suppose an early feature tensor is

X∈RH×W×C.X \in \mathbb{R}^{H\times W\times C}.

Then the network performs something like:

X→Conv→activation→stride/downsampling→X′.X \rightarrow \text{Conv} \rightarrow \text{activation} \rightarrow \text{stride/downsampling} \rightarrow X'.

If

X′∈RH/2×W/2×C′,X' \in \mathbb{R}^{H/2\times W/2\times C'},

there usually is no unique inverse

X′→X.X' \rightarrow X.

Multiple different input patterns can produce essentially the same X′X'.

A stride-2 convolution is an obvious example. Imagine a tiny object occupies only a few pixels:

original feature map

. . . . . .
. . X X . .
. . X X . .
. . . . . .

After repeated downsampling:

. X .
. . .

and later:

X

Some spatial details such as:

  • precise boundaries,
  • local texture,
  • small structures,
  • relative positions,

may simply cease to exist explicitly. Once lost,

f(X1)=f(X2)f(X_1)=f(X_2)

for two distinguishable inputs X1X_1 and X2X_2, and later layers cannot determine which input produced the representation.

No amount of additional training can reconstruct information that the architecture has made indistinguishable.

This is one of the motivations behind the YOLOv9 paper’s discussion of irreversible transformations and information bottlenecks.


2. Activation functions can also destroy information

Consider ReLU:

y=max⁡(0,x).y=\max(0,x).

Then

x=−1, −5, −100x=-1,\,-5,\,-100

all become

y=0.y=0.

So ReLU is many-to-one:

f(−1)=f(−5)=f(−100)=0.f(-1)=f(-5)=f(-100)=0.

The original values cannot be recovered.

Usually this is perfectly acceptable because those distinctions may not matter.

But imagine that subtle negative activation values encode some distinction that helps recognize a difficult object. Once collapsed to zero, downstream layers cannot use that distinction.

This is an architectural information loss, not necessarily evidence of bad optimization.


3. Channel compression forces the network to choose what to retain

Suppose a layer converts

256 channels→64 channels.256\text{ channels}\rightarrow64\text{ channels}.

The network has to compress a high-dimensional representation into a smaller one:

x∈R256→z∈R64.\mathbf{x}\in\mathbb{R}^{256} \rightarrow \mathbf{z}\in\mathbb{R}^{64}.

Unless the original information lies on a sufficiently low-dimensional manifold, this mapping cannot preserve everything.

The learned network therefore makes a practical tradeoff:

retain statistically useful information\text{retain statistically useful information}

and discard information that appears less useful.

The difficulty is that “less useful on average” is not the same as “never useful.”

For example, suppose 98% of cars in training images can be detected from:

  • overall shape,
  • wheels,
  • windows.

The network may learn that tiny mirror details contribute very little to average loss.

For 2% of difficult examples, though, those details might be exactly what distinguishes a car from some confusing background structure.

The globally optimal solution under finite capacity may still discard them.


4. The loss function itself does not tell the network exactly what information to retain

Suppose the target is object detection:

Y=(class,bbox).Y=(\text{class},\text{bbox}).

The network minimizes:

L=Lcls+Lbox+Lobj.\mathcal{L} = \mathcal{L}_{cls} + \mathcal{L}_{box} + \mathcal{L}_{obj}.

The optimizer does not directly say:

Preserve edge A, texture B, small-object feature C, and spatial relation D.

It simply says:

Change parameters in whichever direction reduces the loss.

There may be many representations that produce similar training loss.

For example:

X→Z1→YX \rightarrow Z_1 \rightarrow Y

and

X→Z2→YX \rightarrow Z_2 \rightarrow Y

might both achieve:

L≈0.1.\mathcal{L}\approx0.1.

But Z1Z_1 might retain much richer target-relevant information than Z2Z_2.

The optimizer has no explicit preference for Z1Z_1 unless the architecture or training framework encourages it.

That is essentially where YOLOv9’s argument starts: ordinary end-to-end optimization does not guarantee that the feature representation supplied to the objective contains all the useful information from the input. The authors argue that this can consequently produce biased or incomplete gradient signals.


5. Dataset statistics encourage shortcuts

Another important cause is that the network learns what is predictive, not necessarily what humans regard as the “correct” information.

Imagine:

Training images:

boat → almost always water
airplane → almost always sky
cow → almost always grass

The model can minimize its loss partly using:

P(Y∣background)P(Y\mid\text{background})

instead of learning exclusively:

P(Y∣object shape).P(Y\mid\text{object shape}).

So it might progressively discard fine object information because context is sufficient for most training examples.

A well-trained model could therefore learn:

input
  ↓
object details ─────────── partially discarded
background/context ─────── strongly retained
  ↓
correct training prediction

Then under distribution shift:

cow standing on concrete

the missing object-specific features suddenly matter.

The model was still “well trained” according to its optimization objective. This is usually called shortcut learning or learning spurious correlations.


6. Limited capacity forces tradeoffs

This becomes especially important for lightweight detectors.

Imagine the ideal representation needs: 10001000

effective degrees of freedom, but lightweight network only has capacity for: 300.300.

Even perfect optimization cannot fit all useful distinctions.

The network must approximate: f∗≈fθ.f^* \approx f_\theta.

It tends to spend capacity on signals that provide the largest average reduction in loss.

So features important only for:

  • small objects,
  • rare classes,
  • occluded objects,
  • unusual viewpoints,

may be sacrificed.

This helps explain why YOLOv9 emphasizes that the information-loss problem becomes especially significant in small models and difficult tasks. A large model might have enough redundant channels to preserve a lot of information accidentally. A tiny model cannot afford that redundancy.


7. Small objects are the clearest example

Suppose an object occupies:

8×88\times8

pixels.

After three stride-2 reductions:

8×8→4×4→2×2→1×1.8\times8 \rightarrow 4\times4 \rightarrow 2\times2 \rightarrow 1\times1.

At the deepest feature level the entire object could correspond to roughly one activation.

For a large object:

256×256→128×128→64×64→32×32,256\times256 \rightarrow 128\times128 \rightarrow 64\times64 \rightarrow 32\times32,

there is still extensive structure available.

So the same network can be very well optimized and still systematically lose information for small objects.

That is not contradictory at all. It is a consequence of the representation.


8. Why YOLOv9’s argument is interesting

A naïve view of deep learning would be:

If useful information is being lost, backpropagation should eventually teach the earlier layers not to lose it.

YOLOv9 essentially argues:

Not necessarily—because once that information has disappeared before the loss is calculated, the loss itself may lack enough information to generate the gradient needed to tell earlier layers that they discarded something important.

That’s a somewhat circular optimization problem.

Consider:

X→A→B→C→L.X \rightarrow \boxed{A} \rightarrow \boxed{B} \rightarrow \boxed{C} \rightarrow L.

Suppose BB removes feature qq, and qq was useful for the target.

Then:

C never sees q.C \text{ never sees } q.

Therefore:

L cannot fully assess how q would have improved the prediction.L \text{ cannot fully assess how }q\text{ would have improved the prediction}.

And consequently:

∂L∂B\frac{\partial L}{\partial B}

may not contain a sufficiently strong signal saying:

“Preserve qq.”

YOLOv9’s PGI tries to provide an additional information-rich training path so that the objective can generate more reliable gradients even when the normal main branch has compressed information. The authors specifically position the auxiliary reversible branch as a mechanism for retaining sufficiently complete information for target-task gradient calculation without imposing the reversible machinery at inference.