Mohammed Mudassir Uddin

LAF-YOLOv10arXiv 2609.14560Computer vision · UAV imagerySeptember 2026

Small Object Detection in Drone Aerial Imagery with LAF-YOLOv10

A composability study of four architectural techniques

Four upgrades each help small-object detection on their own: partial convolution, attention-gated feature fusion, a stride-4 detection head and Wise-IoU regression. Combined inside YOLOv10n, they lose 7.8 mAP points. This study measures the combination instead of assuming the gains stack, and traces most of the loss to one interaction: the head swap, made on a backbone whose pretrained weights only partly transferred.

24.0%
mAP@0.5 on VisDrone (3 seeds, ±0.4), against 31.8% for plain YOLOv10n
−2.5pts
interaction penalty: the stack loses 5.5 points where its parts predict 3.0
73of 150
backbone tensors carried over from the COCO-pretrained checkpoint
−10.0pts
on UAVDT zero-shot, matching the vehicle-class pattern

The findingThe parts predict −3.0; the stack measures −5.5

Drone footage is hard for general-purpose detectors. A car can be under 20 pixels across in a 1080p frame, more than half of VisDrone’s objects are smaller than 32×32 pixels, and onboard compute is small. Recent drone detectors respond by stacking techniques that each have published gains, on the assumption that the gains carry over. This paper tests that assumption directly on YOLOv10n.

Adding the modules one at a time: the P2/−P5 step costs twice what it costs alone

Figure 1. mAP@0.5 on VisDrone as each module is added to the same reimplemented YOLOv10n baseline (seed 42, 300 epochs). Bars show the measured change at each step. The dashed track and open circles show what adding each module’s stand-alone effect would predict. The full model, with YOLOv10’s attention block and NMS-free dual head, reaches 24.0%.

What is newTesting the combination instead of assuming it

Common practiceThis study
AssumptionTechniques that help alone will help togetherTested directly on one model, with three training seeds
EvidenceOne aggregate mAP numberIndependent and additive ablations, TIDE error types, per-class results, zero-shot transfer, held-out and test-dev splits, and an audit of the weight transfer
OutcomeA new state-of-the-art rowA measured failure, located in one interaction, plus the single experiment that would confirm the cause

The four modulesOne change for each structural gap

Each module answers a known weakness of YOLOv10n on aerial images. None is new on its own; the contribution is putting all four in one model and measuring what happens.

BACKBONE: PC-C2f REPLACES EVERY C2f BLOCK Input640×640×3 Stemstride 2 PC-C2f160×160 · 32 PC-C2f80×80 · 64 PC-C2f40×40 · 128 PC-C2f20×20 · 256 SPPFk = 5 SE SE SE NECK: AG-FPN SE-gated laterals, DySample upsampling AG-FPN160×160 AG-FPN80×80 AG-FPN40×40 top-down path: DySample ↑2 HEADS P2 added, P5 removed P2 headstride 4 P3 headstride 8 P4 headstride 16 P5 headstride 32 Wise-IoU v3replaces CIoU training loss on every head
Figure 2. LAF-YOLOv10. The four changes to YOLOv10n are highlighted: PC-C2f blocks in the backbone, the AG-FPN neck (SE gates plus DySample), the P2 head that replaces P5, and the Wise-IoU v3 loss. The rest is stock YOLOv10n.

PC-C2f: spatial convolution on a quarter of the channels

Small objects light up only a few feature maps, so a full 3×3 convolution over every channel spends most of its FLOPs on little. Partial convolution applies the 3×3 kernel to the first quarter of the channels, passes the rest through, and mixes them with a 1×1 convolution:

𝐘=Conv1×1(Concat(Conv3×3(𝐗p), 𝐗u))\mathbf{Y} = \mathrm{Conv}_{1\times 1}\big(\mathrm{Concat}(\mathrm{Conv}_{3\times 3}(\mathbf{X}_p),\ \mathbf{X}_u)\big)

That removes about 75% of the spatial convolution cost; with the 1×1 projection counted, the saving per block is closer to 55–60%. The split also changes tensor shapes throughout the backbone, which matters later.

AG-FPN: gate each level before fusing it

Plain concatenation treats every channel as equally useful, so background clutter rides along into the fused map. A squeeze-and-excitation gate rescales each lateral input before fusion (squeeze ratio 16), and DySample replaces nearest-neighbour upsampling with content-dependent sampling offsets for 0.06 extra GFLOPs:

𝜶i=σ(W2 δ(W1 GAP(𝐅i))),𝐅i′=𝜶i⊗𝐅i\boldsymbol{\alpha}_i = \sigma\big(W_2\,\delta(W_1\,\mathrm{GAP}(\mathbf{F}_i))\big),\qquad \mathbf{F}_i' = \boldsymbol{\alpha}_i \otimes \mathbf{F}_i

P2 head in, P5 head out; Wise-IoU v3 as the box loss

The P5 head (stride 32) exists for objects larger than about 256 pixels, and fewer than 0.8% of VisDrone boxes are that large. Removing it saves about 0.3M parameters and 1.2 GFLOPs, which pays for a P2 head (stride 4, a 160×160 grid) aimed at the 6–16-pixel range where the P3 head loses recall. Wise-IoU v3 replaces CIoU and damps the gradient from boxes whose loss is far above the running average, which in crowded aerial scenes are often ambiguous annotations.

(a) PC-C2f block input X, C channels Xₚ · C/4 Xᵤ · 3C/4, passed through untouched Conv 3×3 ≈75% fewer 3×3 FLOPs concatenate, then Conv 1×1 output Y, C channels (b) Grid cell per head against object size P2 · 4 px P3 · 8 P4 · 16 P5 · 32, removed VisDrone training boxes by size (343,205 boxes) under 8 px 5% 8–16 px 17% 16–32 px 32% 32 px and up 46% over 256 px, the P5 head’s range: under 0.8%
Figure 3. (a) The PC-C2f block. (b) Why the head swap looked safe. Grid cells are drawn to scale with one another, and object sizes are counted from the training-split annotations. P5’s territory is almost empty by count, but Section 4 shows the larger vehicles still depended on it.

Where the accuracy wentBelow every compared detector, and not uniformly

LAF-YOLOv10 was trained three times (seeds 42, 123 and 256; 300 epochs; SGD; COCO-pretrained initialization) on VisDrone-DET2019. That is 6,471 training, 548 validation and 1,610 test-dev images across ten classes. The other rows in Figure 4 are taken from their own papers rather than retrained here.

24.0% mAP@0.5, 7.8 points below the unmodified baseline

Figure 4. mAP@0.5 on the VisDrone-DET2019 validation set. LAF-YOLOv10 is the mean of three seeds (whisker: one s.d.); other values are reported by their authors. LAF-YOLOv10 has 2.14M parameters but the highest compute in the comparison (13.55 GFLOPs), from the 160×160 P2 maps and the training-only branch of the NMS-free head.

Large vehicles lose 10–15 points; pedestrians and cars lose 1–3

Figure 5. Per-class mAP@0.5, YOLOv10n against LAF-YOLOv10 (three-seed mean), sorted by the change. The classes that collapse are the medium-to-large vehicles nearest in scale to what the removed P5 head covered.

What the modules were built to fix did improve; missed detections swamped it

Figure 6. Change in each TIDE error type, LAF-YOLOv10 minus YOLOv10n, in points of mAP@0.5 lost to that error. Background false positives nearly vanish (8.1 to 0.04) and localization improves, as AG-FPN and Wise-IoU intend. Missed detections (+15.4) and classification errors (+7.7) grow far more.

Real detections from the best, middle and worst of the validation set

Aerial view of a highway through forest. The cars on the road have green ground-truth boxes, each matched by a red prediction labelled car. Best

Sparse highway traffic. Every car is found.

Oblique aerial view of a crowded outdoor sports court. Dozens of pedestrians are boxed in green; most have overlapping red predictions labelled pedestrian. Typical

A crowded court. Most people are found; distant ones are missed.

Aerial view of a wide road with trees. Two large trucks in the foreground have green ground-truth boxes, but the red predictions cover only small parts of them and label one as a car. Worst

Two trucks in the foreground are missed or labelled as cars, the failure Figure 5 predicts.

Figure 7. Validation images sampled from the bottom, middle and top of a per-image precision-plus-recall ranking, not hand-picked. Green: ground truth. Red: LAF-YOLOv10 predictions with class and confidence.
Table 1. Each module added alone to the reimplemented baseline, and added in sequence (seed 42, 300 epochs, mAP@0.5). The alone effects sum to −3.0; the sequence measures −5.5.
ModuleAdded aloneAdded in sequence
mAPΔGFLOPsmAPΔ
Baseline31.5—6.2831.5—
PC-C2f29.5−2.05.4129.5−2.0
AG-FPN32.5+1.06.3230.5+1.0
P2 head, P5 removed29.0−2.58.0825.5−5.0
Wise-IoU v332.0+0.56.2826.0+0.5
Table 2. Beyond the validation set. The held-out slice (10% of the training images, locked before training) and test-dev were each evaluated once, after every design decision was frozen. UAVDT is zero-shot: trained on VisDrone only, tested on car, bus and truck.
EvaluationImagesmAP@0.5mAP@[.5:.95]YOLOv10n mAP@0.5
VisDrone validation54824.0±0.412.3±0.231.8
VisDrone held-out slice64723.511.9—
VisDrone test-dev1,61022.511.4—
UAVDT test, zero-shot15,06937.819.847.8

The held-out and test-dev results sit close to validation, so the deficit is not validation-set overfitting. The larger UAVDT gap (−10.0) is compositional. UAVDT contains only cars, buses and trucks, and restricting VisDrone’s own results to those classes gives a 10.5-point gap.

Table 3. Each design choice against its alternatives, in the full configuration (seed 42). The differences are within the ±0.4 run-to-run spread, so no single choice explains the deficit.
ComponentVariants, mAP@0.5
Attention gate in AG-FPNSE 24.0 · Coordinate Attention 23.9 · CBAM 23.7
Box regression lossWise-IoU v3 24.0 · SIoU 23.8 · EIoU 23.7 · CIoU 23.6 (TIDE localization error 2.6 to 3.1, same order)
Channels given the 3×3 convfirst C/4 24.0 ± 0.4 · a fixed random C/4 23.8 ± 0.5

The first-quarter and random-quarter results being indistinguishable matters for interpretation. Channel order after initialization means nothing, so PC-C2f works as a regularizer that limits which channels see spatial context, not as a principled selection of useful features.

Training converges smoothly; the gap is not undertraining

Figure 8. Validation mAP@0.5 every 5 epochs for LAF-YOLOv10, seed 42, read from the training log. It rises steeply for 100 epochs and flattens from about epoch 250. A three-seed run takes 17.4 GPU-hours, against 4.2 hours for one YOLOv10n run.

The likely causeHalf the backbone started from random weights

LAF-YOLOv10 is initialized by copying a COCO-pretrained YOLOv10n checkpoint tensor by tensor, wherever the shapes match. PC-C2f changes the channel layout throughout the backbone, so only 73 of 150 backbone tensors match. The other 77 start from random initialization and have to be learned from 6,471 images, orders of magnitude fewer than COCO.

Three results fit this account. PC-C2f alone costs 2.0 points while cutting GFLOPs by 14% as intended, so its compute saving works while its accuracy does not. The P2/−P5 swap costs twice as much on top of that weakened backbone (−5.0) as it does alone (−2.5), because a fine-grained head needs reliable low-level features. And the classes hit hardest are the large vehicles the removed P5 head used to serve.

The deciding experiment

The account is consistent with the data but not yet isolated. The test is a controlled run that transfers weights by name into PC-C2f’s reshaped backbone, or trains longer from random initialization, and checks whether the interaction penalty disappears. A P5-retained variant would also separate adding P2 from removing P5.

LimitationsWhat this study does not show

  • The pretrained-transfer explanation is supported by the pattern of results, not by a controlled run.
  • P2 was added and P5 removed in one step, so their effects are not separated.
  • Comparison rows come from each method’s own paper, not a shared training recipe; their speeds were measured on their own hardware.
  • Generalization is tested on one other dataset (UAVDT). AI-TOD would be a natural third.
  • Hyperparameters follow YOLOv10 defaults, and variance is reported over three seeds without formal significance tests.

CitationCite this work

@article{nayeem2026lafyolov10,
  title   = {Small Object Detection in Drone Aerial Imagery with LAF-YOLOv10},
  author  = {Nayeem, Quratulain and Taranum, Fahmina
             and Uddin, Mohammed Mudassir},
  journal = {arXiv preprint arXiv:2609.14560},
  year    = {2026}
}