Technical note

Extracting Road-Navigation Features from Satellite Imagery with YOLOv9

A presentation of my YOLO application project in satellite imagery feature extractions.

  • Deep Learning
  • YOLO

SFES Group-1 detections over a dense urban road network

Extracting road-navigation features from overhead imagery is more demanding than simply locating roads. A useful system must distinguish individual markings and signs, preserve their shapes, and determine their precise positions—even in dense intersections where multiple features appear close together.

This article introduces the first version of my Satellite Feature Extraction Service (SFES) Group-1 model, a customized YOLOv9-based instance detection and segmentation model. It extracts nine types of road-navigation features together in a single forward pass.

For a deeper discussion of YOLOv9, including Programmable Gradient Information (PGI) and GELAN, please refer to my earlier articles in this series. Here, I focus on the application, model design, evaluation results, and qualitative examples. The underlying YOLOv9 implementation is available in the official YOLOv9 repository.

Why instance segmentation?

Object detection identifies and localizes an object with a bounding box. Semantic segmentation classifies pixels but does not distinguish separate objects of the same class. Instance segmentation combines both ideas: it detects each feature independently and predicts its pixel-level shape.

That distinction is important for road-navigation features. A bounding box may be sufficient to indicate that a crosswalk exists, but its segmented outline provides much richer information about its position, orientation, and spatial extent. This makes instance segmentation better suited to downstream tasks such as map construction, geometry extraction, and spatial analysis.

SFES Group-1 features

The version-1 model predicts the following dataset classes:

Feature categoryClass labels
Stop and yield controlsstop_line, yield_indication, yield_indication_triangle, yield_ahead
Crosswalkslogical_crosswalk, stripes_crosswalk
Road signsbicycle_sign, bus_sign
Directional markingsdirection_arrow

These classes vary considerably in appearance and scale. Crosswalks and directional arrows can occupy large, visually distinctive areas, while bicycle and bus signs may cover only a small number of pixels in a large satellite image.

System architecture

SFES Group-1 feature-extractor architecture

The extractor uses a customized YOLOv9-based architecture with three main stages:

  1. A multi-scale backbone converts the input image into feature maps at several resolutions.
  2. A feature-pyramid neck combines spatial detail with higher-level context.
  3. An enhanced prediction head produces both object detections and instance masks.

Processing all nine classes in one model avoids running a separate extractor for every feature type. The resulting detections provide the class, confidence, and approximate location of each feature, while the masks preserve the feature geometry needed by later mapping stages.

Model profile and runtime

MeasurementVersion-1 result
Parameters27.4 million
Computation144.6 GFLOPs
Input resolution1280 × 1280 pixels
Observed memory footprintLess than 3 GB
Observed inference timeApproximately 30 ms per image

The memory and latency measurements describe the original test environment; they should not be treated as hardware-independent benchmarks. Actual throughput will depend on the accelerator, software stack, batch size, and preprocessing pipeline.

Evaluation results

The following report summarizes the version-1 checkpoint on a validation set containing 1,168 images and 27,200 annotated instances.

Per-class evaluation results for the SFES Group-1 model

The aggregate bounding-box results were:

PrecisionRecallmAP@0.50mAP@0.50–0.95
0.8060.7780.8360.548

Aggregate mask precision reached 0.800, with mask recall at 0.758. Performance varied by feature class:

  • stripes_crosswalk produced the strongest bounding-box result, reaching 0.950 mAP@0.50 and 0.731 mAP@0.50–0.95.
  • direction_arrow also performed well, with 0.914 mAP@0.50 and mask precision/recall of 0.851/0.868.
  • logical_crosswalk reached 0.875 mAP@0.50 and 0.648 mAP@0.50–0.95.
  • bicycle_sign was the most difficult class, with 0.509 mAP@0.50 and 0.386 bounding-box recall. This class is therefore a clear priority for additional data review and model improvement.

Overall, the model demonstrates sufficient capacity to learn this multi-class extraction task. However, these results must be interpreted in the context of the dataset quality.

Dataset limitations

The training and validation annotations contained three recurring issues:

  • Missing annotations: many visible Group-1 features were not labeled. During evaluation, a correct prediction over an unlabeled feature can be counted as a false positive; during training, missing labels also provide inconsistent supervision.
  • Duplicate annotations: some individual objects were labeled more than once. This is especially damaging for instance segmentation because the model is expected to associate one mask with each distinct object.
  • Overlapping features: some feature annotations overlap spatially. Instance segmentation can represent overlapping objects, so this is less serious than missing or duplicate labels, but consistent boundaries and spatial placement remain important.

Cleaning these annotations should improve both precision and recall and will make future comparisons more reliable. The version-1 results should therefore be viewed as a baseline rather than the model’s performance ceiling.

Qualitative demonstrations

The following examples compare each model input with its prediction overlay across dense urban intersections, roundabouts, curved roads, heavy shadows, and bridge approaches. Click any image to open the full-resolution version and inspect the scene, masks, and labels.

Example 1 — Dense urban streets with navigation features distributed across several intersections

Input imageModel prediction
Input aerial image of a dense urban road networkSFES detections across a dense urban road network

Example 2 — A roundabout and adjacent road network surrounded by buildings, trees, and parked vehicles

Input imageModel prediction
Input aerial image near a roundabout and parking areaSFES detections near a roundabout and parking area

Example 3 — An urban corridor containing construction activity and strong building shadows

Input imageModel prediction
Input aerial image of an urban corridor with strong shadowsSFES detections along an urban corridor with strong shadows

Example 4 — Curved road segments and intersections near a large bus-storage area

Input imageModel prediction
Input aerial image of curved roads near a bus depotSFES detections on curved roads near a bus depot

Example 5 — A winding road through a mixed residential and urban landscape

Input imageModel prediction
Input aerial image of a winding residential roadSFES detections along a winding residential road

Example 6 — Road-navigation features at both approaches to a bridge in a riverside scene

Input imageModel prediction
Input aerial image of intersections near a bridgeSFES detections at intersections near a bridge

Conclusions and next steps

The SFES Group-1 version-1 model shows that a single YOLOv9-based instance segmentation network can extract a diverse set of road-navigation features from large overhead images. It delivers useful results at 1280 × 1280 resolution while retaining the object-level masks required for precise spatial analysis.

The evaluation also makes the next priorities clear:

  1. remove duplicate annotations and fill in missing labels;
  2. review class definitions and mask boundaries for consistency;
  3. expand and rebalance the more difficult classes, particularly bicycle_sign;
  4. retrain and compare the revised model against this version-1 baseline; and
  5. evaluate how accurately the predicted masks can be converted into map-ready geometry.

With cleaner data and targeted refinement, this model should provide a stronger foundation for automated road-feature extraction and downstream navigation-map production.