dronefreak/uavdt-yolo26m

Model

YOLO26m Finetuned on UAVDT

0

8 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

YOLO26m Finetuned on UAVDT

Fine-tuned YOLO26m object detector on the UAVDT benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.

YOLO26m detections on two UAVDT test clips


Task Framework Base Model
mAP@50 mAP@50:95 Params
License Source

Usage

Install Dependencies

pip install ultralytics huggingface_hub

Load Model from Hugging Face

from huggingface_hub import hf_hub_download
from ultralytics import YOLO

weights = hf_hub_download(
    repo_id="dronefreak/uavdt-yolo26m",
    filename="best.pt"
)

model = YOLO(weights)

Run Inference

results = model.predict(
    source="image.jpg",
    conf=0.25
)

results[0].show()

Performance

Evaluated on the UAVDT test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).

MetricScore (%)
mAP@5033.43
mAP@50-9519.56
Precision38.14
Recall39.84
F1 Score38.97
Parameters21.9M
FLOPs75.4B (at 640 px)

UAVDT Model Zoo

Every model DetectionBench has trained and evaluated on UAVDT so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.

ModelmAP@50mAP@50-95PrecisionRecall
YOLO26m33.4319.5638.1439.84
RF-DETR Medium33.2820.5473.0370.03
YOLO26s32.9819.6143.8640.38
RF-DETR Nano32.7820.3173.666.98
YOLO26x32.6519.2241.8538.19
YOLO26l32.6418.7540.1736.45
RF-DETR Small32.6220.2173.8371.63
YOLOv9s31.8218.7139.8338.12
YOLOv8m31.4218.840.2737.79
YOLO11x31.0518.3137.436.38
YOLO11m30.4717.7137.737.01
YOLOv8x30.4717.6639.6136.16
YOLOv10m30.1217.3340.1335.68
YOLOv9m29.4316.9735.9235.7
YOLOv9t29.4217.0335.7536.47
YOLOv10x29.3817.1537.2935.15
YOLOv10l29.1616.5436.935.6
YOLOv9c29.1616.4635.3834.35
YOLO11s29.117.1634.3237.31
YOLO26n28.8816.7933.1435.66
YOLOv8l28.8617.2738.3332.86
YOLOv10s28.8516.4836.5333.16
YOLO11l28.6417.1634.7534.02
YOLO11n28.5616.338.0432.26
YOLOv9e28.116.635.5132.54
YOLOv8n27.815.3435.4233.61
YOLOv10n27.1715.1633.331.21
YOLOv8s27.1215.3334.6531.87

Per-Class Performance

ClassmAP@50mAP@50-95
car73.7540.42
truck12.678.12
bus13.8810.14

Normalized Confusion Matrix


Dataset

This model was trained on UAVDT. For the full dataset description, provenance, license, and citation, see the dataset card:

https://huggingface.co/datasets/dronefreak/UAVDT

Classes

  • car
  • truck
  • bus

Training Configuration

SettingValue
DatasetUAVDT
FrameworkUltralytics YOLO
Training ToolkitDetectionBench
Epochs (configured max)30
Epochs (actually trained)11
Early Stopping Patience8
Batch Sizeauto (Ultralytics AutoBatch)
Image Size1024
OptimizerAdamW
Initial Learning Rate0.0005
Seed0

Repository Contents

best.pt
results.csv
args.yaml
BoxPR_curve.png
BoxF1_curve.png
BoxP_curve.png
BoxR_curve.png
confusion_matrix.png
confusion_matrix_normalized.png
val_batch0_pred.jpg
uavdt_yolo26m_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.md


Training Framework

This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.

Features include:

  • A dataset-adapter registry for converting real-world datasets into a canonical format
  • Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
  • Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
  • One-command reproducibility via versioned Hydra configs

If you find this model useful, please consider starring the repository.


Known Limitations

  • Severe class imbalance: car (94.6%) dominates the annotated boxes, while truck (3.1%) and bus (2.3%) are rare -- per-class accuracy on the minority classes is measured on comparatively few examples, and every model here scores far lower on them than on car.
  • Very small objects: the median box covers only 0.14% of the image area (mean 0.26%), so this is a hard small-object regime and absolute mAP values are low for every architecture; the numbers are best read as a relative comparison between models, not as a production-quality detector.
  • Video-derived, highly correlated frames: the ~40.7k labelled images come from 50 video sequences, so consecutive frames are near-duplicates. UAVDT's 50 tracking-only sequences have no detection labels and are excluded. The validation split is carved out of the training sequences by sequence (not by frame) to avoid leakage, but effective diversity is far lower than the image count suggests.
  • Different density per split: instances per image are 15.7 (train), 28.0 (valid) and 22.7 (test), because the splits contain different sequences -- validation metrics are not directly predictive of test metrics.
  • Research-use-only data: UAVDT is distributed "for research purpose only" with no redistribution grant, so the dataset is not mirrored here -- obtain it from the official source (see the Dataset section above) and check its terms before any use beyond research.

Citation

If you use this model in your research, please consider citing the dataset and the model architecture:

@InProceedings{du2018unmanned,
  title={The Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking},
  author={Du, Dawei and Qi, Yuankai and Yu, Hongyang and Yang, Yifan and Duan, Kaiwen and Li, Guorong and Zhang, Weigang and Huang, Qingming and Tian, Qi},
  booktitle={Proceedings of the European Conference on Computer Vision (ECCV)},
  year={2018}
}
@article{jocher2026yolo26,
  title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models},
  author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat},
  journal={arXiv preprint arXiv:2606.03748},
  year={2026}
}
aerial-imagery
computer-vision
detectionbench
drone
model-index
object-detection
pytorch
small-object-detection
traffic-surveillance
ultralytics
vehicle-detection

dronefreak/uavdt-yolo26m

Model

YOLO26m Finetuned on UAVDT

0

8 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

YOLO26m Finetuned on UAVDT

Fine-tuned YOLO26m object detector on the UAVDT benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.

YOLO26m detections on two UAVDT test clips


Task Framework Base Model
mAP@50 mAP@50:95 Params
License Source

Usage

Install Dependencies

pip install ultralytics huggingface_hub

Load Model from Hugging Face

from huggingface_hub import hf_hub_download
from ultralytics import YOLO

weights = hf_hub_download(
    repo_id="dronefreak/uavdt-yolo26m",
    filename="best.pt"
)

model = YOLO(weights)

Run Inference

results = model.predict(
    source="image.jpg",
    conf=0.25
)

results[0].show()

Performance

Evaluated on the UAVDT test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).

MetricScore (%)
mAP@5033.43
mAP@50-9519.56
Precision38.14
Recall39.84
F1 Score38.97
Parameters21.9M
FLOPs75.4B (at 640 px)

UAVDT Model Zoo

Every model DetectionBench has trained and evaluated on UAVDT so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.

ModelmAP@50mAP@50-95PrecisionRecall
YOLO26m33.4319.5638.1439.84
RF-DETR Medium33.2820.5473.0370.03
YOLO26s32.9819.6143.8640.38
RF-DETR Nano32.7820.3173.666.98
YOLO26x32.6519.2241.8538.19
YOLO26l32.6418.7540.1736.45
RF-DETR Small32.6220.2173.8371.63
YOLOv9s31.8218.7139.8338.12
YOLOv8m31.4218.840.2737.79
YOLO11x31.0518.3137.436.38
YOLO11m30.4717.7137.737.01
YOLOv8x30.4717.6639.6136.16
YOLOv10m30.1217.3340.1335.68
YOLOv9m29.4316.9735.9235.7
YOLOv9t29.4217.0335.7536.47
YOLOv10x29.3817.1537.2935.15
YOLOv10l29.1616.5436.935.6
YOLOv9c29.1616.4635.3834.35
YOLO11s29.117.1634.3237.31
YOLO26n28.8816.7933.1435.66
YOLOv8l28.8617.2738.3332.86
YOLOv10s28.8516.4836.5333.16
YOLO11l28.6417.1634.7534.02
YOLO11n28.5616.338.0432.26
YOLOv9e28.116.635.5132.54
YOLOv8n27.815.3435.4233.61
YOLOv10n27.1715.1633.331.21
YOLOv8s27.1215.3334.6531.87

Per-Class Performance

ClassmAP@50mAP@50-95
car73.7540.42
truck12.678.12
bus13.8810.14

Normalized Confusion Matrix


Dataset

This model was trained on UAVDT. For the full dataset description, provenance, license, and citation, see the dataset card:

https://huggingface.co/datasets/dronefreak/UAVDT

Classes

  • car
  • truck
  • bus

Training Configuration

SettingValue
DatasetUAVDT
FrameworkUltralytics YOLO
Training ToolkitDetectionBench
Epochs (configured max)30
Epochs (actually trained)11
Early Stopping Patience8
Batch Sizeauto (Ultralytics AutoBatch)
Image Size1024
OptimizerAdamW
Initial Learning Rate0.0005
Seed0

Repository Contents

best.pt
results.csv
args.yaml
BoxPR_curve.png
BoxF1_curve.png
BoxP_curve.png
BoxR_curve.png
confusion_matrix.png
confusion_matrix_normalized.png
val_batch0_pred.jpg
uavdt_yolo26m_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.md


Training Framework

This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.

Features include:

  • A dataset-adapter registry for converting real-world datasets into a canonical format
  • Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
  • Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
  • One-command reproducibility via versioned Hydra configs

If you find this model useful, please consider starring the repository.


Known Limitations

  • Severe class imbalance: car (94.6%) dominates the annotated boxes, while truck (3.1%) and bus (2.3%) are rare -- per-class accuracy on the minority classes is measured on comparatively few examples, and every model here scores far lower on them than on car.
  • Very small objects: the median box covers only 0.14% of the image area (mean 0.26%), so this is a hard small-object regime and absolute mAP values are low for every architecture; the numbers are best read as a relative comparison between models, not as a production-quality detector.
  • Video-derived, highly correlated frames: the ~40.7k labelled images come from 50 video sequences, so consecutive frames are near-duplicates. UAVDT's 50 tracking-only sequences have no detection labels and are excluded. The validation split is carved out of the training sequences by sequence (not by frame) to avoid leakage, but effective diversity is far lower than the image count suggests.
  • Different density per split: instances per image are 15.7 (train), 28.0 (valid) and 22.7 (test), because the splits contain different sequences -- validation metrics are not directly predictive of test metrics.
  • Research-use-only data: UAVDT is distributed "for research purpose only" with no redistribution grant, so the dataset is not mirrored here -- obtain it from the official source (see the Dataset section above) and check its terms before any use beyond research.

Citation

If you use this model in your research, please consider citing the dataset and the model architecture:

@InProceedings{du2018unmanned,
  title={The Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking},
  author={Du, Dawei and Qi, Yuankai and Yu, Hongyang and Yang, Yifan and Duan, Kaiwen and Li, Guorong and Zhang, Weigang and Huang, Qingming and Tian, Qi},
  booktitle={Proceedings of the European Conference on Computer Vision (ECCV)},
  year={2018}
}
@article{jocher2026yolo26,
  title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models},
  author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat},
  journal={arXiv preprint arXiv:2606.03748},
  year={2026}
}
aerial-imagery
computer-vision
detectionbench
drone
model-index
object-detection
pytorch
small-object-detection
traffic-surveillance
ultralytics
vehicle-detection