dronefreak/visdrone-yolo26m

Model

YOLO26m Finetuned on VisDrone-DET

4

6 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

YOLO26m Finetuned on VisDrone-DET

Fine-tuned YOLO26m object detector on the VisDrone-DET benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.

YOLO26m detections on two VisDrone-DET test clips


Task Framework Base Model
mAP@50 mAP@50:95 Params
License Source

Usage

Install Dependencies

pip install ultralytics huggingface_hub

Load Model from Hugging Face

from huggingface_hub import hf_hub_download
from ultralytics import YOLO

weights = hf_hub_download(
    repo_id="dronefreak/visdrone-yolo26m",
    filename="best.pt"
)

model = YOLO(weights)

Run Inference

results = model.predict(
    source="image.jpg",
    conf=0.25
)

results[0].show()

Performance

Evaluated on the VisDrone-DET test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).

MetricScore (%)
mAP@5049.11
mAP@50-9529.24
Precision60.21
Recall49.76
F1 Score54.49
Parameters21.9M
FLOPs75.4B (at 640 px)

VisDrone-DET Model Zoo

Every model DetectionBench has trained and evaluated on VisDrone-DET so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.

ModelmAP@50mAP@50-95PrecisionRecall
YOLO26m49.1129.2460.2149.76
YOLO11m48.0328.5659.3448.3
YOLOv9m47.6228.5259.2648.4
YOLOv10m46.7327.7259.2347.35
YOLOv8m45.4726.9457.8646.39
YOLOv9s45.3826.9556.9546.19
YOLO26s44.8726.4356.4445.41
YOLOv10s44.5126.2855.9745.61
YOLO11s43.6425.8854.4845.01
YOLOv8s43.4725.7756.0944.51
YOLOv9t40.6723.7352.8441.78
YOLO26n39.922.9551.1442.08
YOLOv10n39.823.0851.2241.5
YOLOv8n39.6923.0352.0241.36
RF-DETR Medium39.6221.9970.6246.76
YOLO11n39.5223.051.4941.02
RF-DETR Small39.1321.6864.3849.47
RF-DETR Nano37.9220.8869.0246.21

Per-Class Performance

ClassmAP@50mAP@50-95
pedestrian51.5422.38
people33.3712.8
bicycle26.0712.27
car84.4955.5
van52.1136.4
truck59.740.56
tricycle33.9819.88
awning-tricycle27.8218.02
bus69.5551.08
motor52.4423.56
others0.00.0

Normalized Confusion Matrix


Dataset

This model was trained on VisDrone-DET. For the full dataset description, provenance, license, and citation, see the dataset card:

https://huggingface.co/datasets/Voxel51/VisDrone2019-DET

Classes

  • pedestrian
  • people
  • bicycle
  • car
  • van
  • truck
  • tricycle
  • awning-tricycle
  • bus
  • motor
  • others

Training Configuration

SettingValue
DatasetVisDrone-DET
FrameworkUltralytics YOLO
Training ToolkitDetectionBench
Epochs (configured max)100
Epochs (actually trained)75
Early Stopping Patience25
Batch Size4
Image Size1280
OptimizerSGD
Initial Learning Rate0.01
Seed0

Repository Contents

best.pt
results.csv
args.yaml
BoxPR_curve.png
BoxF1_curve.png
BoxP_curve.png
BoxR_curve.png
confusion_matrix.png
confusion_matrix_normalized.png
val_batch0_pred.jpg
visdrone_yolo26m_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.md


Training Framework

This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.

Features include:

  • A dataset-adapter registry for converting real-world datasets into a canonical format
  • Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
  • Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
  • One-command reproducibility via versioned Hydra configs

If you find this model useful, please consider starring the repository.


Known Limitations

  • Severe class imbalance: car (42.21%) and pedestrian (23.12%) account for two-thirds of all annotated boxes in the training set, while awning-tricycle (0.95%) and tricycle (1.40%) are rare -- the others class has zero annotated instances in the training set entirely and is effectively unusable (always 0 AP).
  • Extreme small-object density: ~53 annotated boxes per image on average, with roughly 69% of boxes covering under 0.1% of the image area -- consistent with VisDrone's aerial small-object detection challenge (objects captured from significant altitude).
  • The original authors license VisDrone under CC BY-NC-SA 3.0 -- non-commercial research use only (see the dataset's homepage); this applies to any model trained on it, not only the raw images.
  • These RF-DETR checkpoints were trained/evaluated directly through DetectionBench. The YOLO/RT-DETR rows in the External VisDrone Model Zoo comparison below were trained via a separate companion codebase, not reproduced inside DetectionBench -- see that collection for their own training details and caveats.

Citation

If you use this model in your research, please consider citing the dataset and the model architecture:

@article{zhu2018vision,
  title={Vision meets drones: A challenge},
  author={Zhu, Pengfei and Wen, Longyin and Bian, Xiao and Ling, Haibin and Hu, Qinghua},
  journal={arXiv preprint arXiv:1804.07437},
  year={2018}
}
@article{jocher2026yolo26,
  title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models},
  author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat},
  journal={arXiv preprint arXiv:2606.03748},
  year={2026}
}
aerial-imagery
computer-vision
detectionbench
detr
drone
model-index
object-detection
pytorch
roboflow
transformer
ultralytics
visdrone

dronefreak/visdrone-yolo26m

Model

YOLO26m Finetuned on VisDrone-DET

4

6 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

YOLO26m Finetuned on VisDrone-DET

Fine-tuned YOLO26m object detector on the VisDrone-DET benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.

YOLO26m detections on two VisDrone-DET test clips


Task Framework Base Model
mAP@50 mAP@50:95 Params
License Source

Usage

Install Dependencies

pip install ultralytics huggingface_hub

Load Model from Hugging Face

from huggingface_hub import hf_hub_download
from ultralytics import YOLO

weights = hf_hub_download(
    repo_id="dronefreak/visdrone-yolo26m",
    filename="best.pt"
)

model = YOLO(weights)

Run Inference

results = model.predict(
    source="image.jpg",
    conf=0.25
)

results[0].show()

Performance

Evaluated on the VisDrone-DET test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).

MetricScore (%)
mAP@5049.11
mAP@50-9529.24
Precision60.21
Recall49.76
F1 Score54.49
Parameters21.9M
FLOPs75.4B (at 640 px)

VisDrone-DET Model Zoo

Every model DetectionBench has trained and evaluated on VisDrone-DET so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.

ModelmAP@50mAP@50-95PrecisionRecall
YOLO26m49.1129.2460.2149.76
YOLO11m48.0328.5659.3448.3
YOLOv9m47.6228.5259.2648.4
YOLOv10m46.7327.7259.2347.35
YOLOv8m45.4726.9457.8646.39
YOLOv9s45.3826.9556.9546.19
YOLO26s44.8726.4356.4445.41
YOLOv10s44.5126.2855.9745.61
YOLO11s43.6425.8854.4845.01
YOLOv8s43.4725.7756.0944.51
YOLOv9t40.6723.7352.8441.78
YOLO26n39.922.9551.1442.08
YOLOv10n39.823.0851.2241.5
YOLOv8n39.6923.0352.0241.36
RF-DETR Medium39.6221.9970.6246.76
YOLO11n39.5223.051.4941.02
RF-DETR Small39.1321.6864.3849.47
RF-DETR Nano37.9220.8869.0246.21

Per-Class Performance

ClassmAP@50mAP@50-95
pedestrian51.5422.38
people33.3712.8
bicycle26.0712.27
car84.4955.5
van52.1136.4
truck59.740.56
tricycle33.9819.88
awning-tricycle27.8218.02
bus69.5551.08
motor52.4423.56
others0.00.0

Normalized Confusion Matrix


Dataset

This model was trained on VisDrone-DET. For the full dataset description, provenance, license, and citation, see the dataset card:

https://huggingface.co/datasets/Voxel51/VisDrone2019-DET

Classes

  • pedestrian
  • people
  • bicycle
  • car
  • van
  • truck
  • tricycle
  • awning-tricycle
  • bus
  • motor
  • others

Training Configuration

SettingValue
DatasetVisDrone-DET
FrameworkUltralytics YOLO
Training ToolkitDetectionBench
Epochs (configured max)100
Epochs (actually trained)75
Early Stopping Patience25
Batch Size4
Image Size1280
OptimizerSGD
Initial Learning Rate0.01
Seed0

Repository Contents

best.pt
results.csv
args.yaml
BoxPR_curve.png
BoxF1_curve.png
BoxP_curve.png
BoxR_curve.png
confusion_matrix.png
confusion_matrix_normalized.png
val_batch0_pred.jpg
visdrone_yolo26m_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.md


Training Framework

This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.

Features include:

  • A dataset-adapter registry for converting real-world datasets into a canonical format
  • Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
  • Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
  • One-command reproducibility via versioned Hydra configs

If you find this model useful, please consider starring the repository.


Known Limitations

  • Severe class imbalance: car (42.21%) and pedestrian (23.12%) account for two-thirds of all annotated boxes in the training set, while awning-tricycle (0.95%) and tricycle (1.40%) are rare -- the others class has zero annotated instances in the training set entirely and is effectively unusable (always 0 AP).
  • Extreme small-object density: ~53 annotated boxes per image on average, with roughly 69% of boxes covering under 0.1% of the image area -- consistent with VisDrone's aerial small-object detection challenge (objects captured from significant altitude).
  • The original authors license VisDrone under CC BY-NC-SA 3.0 -- non-commercial research use only (see the dataset's homepage); this applies to any model trained on it, not only the raw images.
  • These RF-DETR checkpoints were trained/evaluated directly through DetectionBench. The YOLO/RT-DETR rows in the External VisDrone Model Zoo comparison below were trained via a separate companion codebase, not reproduced inside DetectionBench -- see that collection for their own training details and caveats.

Citation

If you use this model in your research, please consider citing the dataset and the model architecture:

@article{zhu2018vision,
  title={Vision meets drones: A challenge},
  author={Zhu, Pengfei and Wen, Longyin and Bian, Xiao and Ling, Haibin and Hu, Qinghua},
  journal={arXiv preprint arXiv:1804.07437},
  year={2018}
}
@article{jocher2026yolo26,
  title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models},
  author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat},
  journal={arXiv preprint arXiv:2606.03748},
  year={2026}
}
aerial-imagery
computer-vision
detectionbench
detr
drone
model-index
object-detection
pytorch
roboflow
transformer
ultralytics
visdrone