Technical note
Extracting Road-Navigation Features from Satellite Imagery with YOLOv9
A presentation of my YOLO application project in satellite imagery feature extractions.

Extracting road-navigation features from overhead imagery is more demanding than simply locating roads. A useful system must distinguish individual markings and signs, preserve their shapes, and determine their precise positions—even in dense intersections where multiple features appear close together.
This article introduces the first version of my Satellite Feature Extraction Service (SFES) Group-1 model, a customized YOLOv9-based instance detection and segmentation model. It extracts nine types of road-navigation features together in a single forward pass.
For a deeper discussion of YOLOv9, including Programmable Gradient Information (PGI) and GELAN, please refer to my earlier articles in this series. Here, I focus on the application, model design, evaluation results, and qualitative examples. The underlying YOLOv9 implementation is available in the official YOLOv9 repository.
Why instance segmentation?
Object detection identifies and localizes an object with a bounding box. Semantic segmentation classifies pixels but does not distinguish separate objects of the same class. Instance segmentation combines both ideas: it detects each feature independently and predicts its pixel-level shape.
That distinction is important for road-navigation features. A bounding box may be sufficient to indicate that a crosswalk exists, but its segmented outline provides much richer information about its position, orientation, and spatial extent. This makes instance segmentation better suited to downstream tasks such as map construction, geometry extraction, and spatial analysis.
SFES Group-1 features
The version-1 model predicts the following dataset classes:
| Feature category | Class labels |
|---|---|
| Stop and yield controls | stop_line, yield_indication, yield_indication_triangle, yield_ahead |
| Crosswalks | logical_crosswalk, stripes_crosswalk |
| Road signs | bicycle_sign, bus_sign |
| Directional markings | direction_arrow |
These classes vary considerably in appearance and scale. Crosswalks and directional arrows can occupy large, visually distinctive areas, while bicycle and bus signs may cover only a small number of pixels in a large satellite image.
System architecture

The extractor uses a customized YOLOv9-based architecture with three main stages:
- A multi-scale backbone converts the input image into feature maps at several resolutions.
- A feature-pyramid neck combines spatial detail with higher-level context.
- An enhanced prediction head produces both object detections and instance masks.
Processing all nine classes in one model avoids running a separate extractor for every feature type. The resulting detections provide the class, confidence, and approximate location of each feature, while the masks preserve the feature geometry needed by later mapping stages.
Model profile and runtime
| Measurement | Version-1 result |
|---|---|
| Parameters | 27.4 million |
| Computation | 144.6 GFLOPs |
| Input resolution | 1280 × 1280 pixels |
| Observed memory footprint | Less than 3 GB |
| Observed inference time | Approximately 30 ms per image |
The memory and latency measurements describe the original test environment; they should not be treated as hardware-independent benchmarks. Actual throughput will depend on the accelerator, software stack, batch size, and preprocessing pipeline.
Evaluation results
The following report summarizes the version-1 checkpoint on a validation set containing 1,168 images and 27,200 annotated instances.
The aggregate bounding-box results were:
| Precision | Recall | mAP@0.50 | mAP@0.50–0.95 |
|---|---|---|---|
| 0.806 | 0.778 | 0.836 | 0.548 |
Aggregate mask precision reached 0.800, with mask recall at 0.758. Performance varied by feature class:
stripes_crosswalkproduced the strongest bounding-box result, reaching 0.950 mAP@0.50 and 0.731 mAP@0.50–0.95.direction_arrowalso performed well, with 0.914 mAP@0.50 and mask precision/recall of 0.851/0.868.logical_crosswalkreached 0.875 mAP@0.50 and 0.648 mAP@0.50–0.95.bicycle_signwas the most difficult class, with 0.509 mAP@0.50 and 0.386 bounding-box recall. This class is therefore a clear priority for additional data review and model improvement.
Overall, the model demonstrates sufficient capacity to learn this multi-class extraction task. However, these results must be interpreted in the context of the dataset quality.
Dataset limitations
The training and validation annotations contained three recurring issues:
- Missing annotations: many visible Group-1 features were not labeled. During evaluation, a correct prediction over an unlabeled feature can be counted as a false positive; during training, missing labels also provide inconsistent supervision.
- Duplicate annotations: some individual objects were labeled more than once. This is especially damaging for instance segmentation because the model is expected to associate one mask with each distinct object.
- Overlapping features: some feature annotations overlap spatially. Instance segmentation can represent overlapping objects, so this is less serious than missing or duplicate labels, but consistent boundaries and spatial placement remain important.
Cleaning these annotations should improve both precision and recall and will make future comparisons more reliable. The version-1 results should therefore be viewed as a baseline rather than the model’s performance ceiling.
Qualitative demonstrations
The following examples compare each model input with its prediction overlay across dense urban intersections, roundabouts, curved roads, heavy shadows, and bridge approaches. Click any image to open the full-resolution version and inspect the scene, masks, and labels.
Example 1 — Dense urban streets with navigation features distributed across several intersections
| Input image | Model prediction |
|---|---|
![]() | ![]() |
Example 2 — A roundabout and adjacent road network surrounded by buildings, trees, and parked vehicles
| Input image | Model prediction |
|---|---|
![]() | ![]() |
Example 3 — An urban corridor containing construction activity and strong building shadows
| Input image | Model prediction |
|---|---|
![]() | ![]() |
Example 4 — Curved road segments and intersections near a large bus-storage area
| Input image | Model prediction |
|---|---|
![]() | ![]() |
Example 5 — A winding road through a mixed residential and urban landscape
| Input image | Model prediction |
|---|---|
![]() | ![]() |
Example 6 — Road-navigation features at both approaches to a bridge in a riverside scene
| Input image | Model prediction |
|---|---|
![]() | ![]() |
Conclusions and next steps
The SFES Group-1 version-1 model shows that a single YOLOv9-based instance segmentation network can extract a diverse set of road-navigation features from large overhead images. It delivers useful results at 1280 × 1280 resolution while retaining the object-level masks required for precise spatial analysis.
The evaluation also makes the next priorities clear:
- remove duplicate annotations and fill in missing labels;
- review class definitions and mask boundaries for consistency;
- expand and rebalance the more difficult classes, particularly
bicycle_sign; - retrain and compare the revised model against this version-1 baseline; and
- evaluate how accurately the predicted masks can be converted into map-ready geometry.
With cleaner data and targeted refinement, this model should provide a stronger foundation for automated road-feature extraction and downstream navigation-map production.












