Fine-tuned YOLO26m object detector on the VisDrone-DET benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.

pip install ultralytics huggingface_hub
from huggingface_hub import hf_hub_download
from ultralytics import YOLO
weights = hf_hub_download(
repo_id="dronefreak/visdrone-yolo26m",
filename="best.pt"
)
model = YOLO(weights)
results = model.predict(
source="image.jpg",
conf=0.25
)
results[0].show()
Evaluated on the VisDrone-DET test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).
| Metric | Score (%) |
|---|---|
| mAP@50 | 49.11 |
| mAP@50-95 | 29.24 |
| Precision | 60.21 |
| Recall | 49.76 |
| F1 Score | 54.49 |
| Parameters | 21.9M |
| FLOPs | 75.4B (at 640 px) |
Every model DetectionBench has trained and evaluated on VisDrone-DET so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.
| Model | mAP@50 | mAP@50-95 | Precision | Recall |
|---|---|---|---|---|
| YOLO26m | 49.11 | 29.24 | 60.21 | 49.76 |
| YOLO11m | 48.03 | 28.56 | 59.34 | 48.3 |
| YOLOv9m | 47.62 | 28.52 | 59.26 | 48.4 |
| YOLOv10m | 46.73 | 27.72 | 59.23 | 47.35 |
| YOLOv8m | 45.47 | 26.94 | 57.86 | 46.39 |
| YOLOv9s | 45.38 | 26.95 | 56.95 | 46.19 |
| YOLO26s | 44.87 | 26.43 | 56.44 | 45.41 |
| YOLOv10s | 44.51 | 26.28 | 55.97 | 45.61 |
| YOLO11s | 43.64 | 25.88 | 54.48 | 45.01 |
| YOLOv8s | 43.47 | 25.77 | 56.09 | 44.51 |
| YOLOv9t | 40.67 | 23.73 | 52.84 | 41.78 |
| YOLO26n | 39.9 | 22.95 | 51.14 | 42.08 |
| YOLOv10n | 39.8 | 23.08 | 51.22 | 41.5 |
| YOLOv8n | 39.69 | 23.03 | 52.02 | 41.36 |
| RF-DETR Medium | 39.62 | 21.99 | 70.62 | 46.76 |
| YOLO11n | 39.52 | 23.0 | 51.49 | 41.02 |
| RF-DETR Small | 39.13 | 21.68 | 64.38 | 49.47 |
| RF-DETR Nano | 37.92 | 20.88 | 69.02 | 46.21 |
| Class | mAP@50 | mAP@50-95 |
|---|---|---|
| pedestrian | 51.54 | 22.38 |
| people | 33.37 | 12.8 |
| bicycle | 26.07 | 12.27 |
| car | 84.49 | 55.5 |
| van | 52.11 | 36.4 |
| truck | 59.7 | 40.56 |
| tricycle | 33.98 | 19.88 |
| awning-tricycle | 27.82 | 18.02 |
| bus | 69.55 | 51.08 |
| motor | 52.44 | 23.56 |
| others | 0.0 | 0.0 |

This model was trained on VisDrone-DET. For the full dataset description, provenance, license, and citation, see the dataset card:
https://huggingface.co/datasets/Voxel51/VisDrone2019-DET
| Setting | Value |
|---|---|
| Dataset | VisDrone-DET |
| Framework | Ultralytics YOLO |
| Training Toolkit | DetectionBench |
| Epochs (configured max) | 100 |
| Epochs (actually trained) | 75 |
| Early Stopping Patience | 25 |
| Batch Size | 4 |
| Image Size | 1280 |
| Optimizer | SGD |
| Initial Learning Rate | 0.01 |
| Seed | 0 |
best.pt
results.csv
args.yaml
BoxPR_curve.png
BoxF1_curve.png
BoxP_curve.png
BoxR_curve.png
confusion_matrix.png
confusion_matrix_normalized.png
val_batch0_pred.jpg
visdrone_yolo26m_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.md
This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.
Features include:
If you find this model useful, please consider starring the repository.
car (42.21%) and pedestrian (23.12%) account for two-thirds of all annotated boxes in the training set, while awning-tricycle (0.95%) and tricycle (1.40%) are rare -- the others class has zero annotated instances in the training set entirely and is effectively unusable (always 0 AP).If you use this model in your research, please consider citing the dataset and the model architecture:
@article{zhu2018vision,
title={Vision meets drones: A challenge},
author={Zhu, Pengfei and Wen, Longyin and Bian, Xiao and Ling, Haibin and Hu, Qinghua},
journal={arXiv preprint arXiv:1804.07437},
year={2018}
}
@article{jocher2026yolo26,
title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models},
author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat},
journal={arXiv preprint arXiv:2606.03748},
year={2026}
}
Fine-tuned YOLO26m object detector on the VisDrone-DET benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.

pip install ultralytics huggingface_hub
from huggingface_hub import hf_hub_download
from ultralytics import YOLO
weights = hf_hub_download(
repo_id="dronefreak/visdrone-yolo26m",
filename="best.pt"
)
model = YOLO(weights)
results = model.predict(
source="image.jpg",
conf=0.25
)
results[0].show()
Evaluated on the VisDrone-DET test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).
| Metric | Score (%) |
|---|---|
| mAP@50 | 49.11 |
| mAP@50-95 | 29.24 |
| Precision | 60.21 |
| Recall | 49.76 |
| F1 Score | 54.49 |
| Parameters | 21.9M |
| FLOPs | 75.4B (at 640 px) |
Every model DetectionBench has trained and evaluated on VisDrone-DET so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.
| Model | mAP@50 | mAP@50-95 | Precision | Recall |
|---|---|---|---|---|
| YOLO26m | 49.11 | 29.24 | 60.21 | 49.76 |
| YOLO11m | 48.03 | 28.56 | 59.34 | 48.3 |
| YOLOv9m | 47.62 | 28.52 | 59.26 | 48.4 |
| YOLOv10m | 46.73 | 27.72 | 59.23 | 47.35 |
| YOLOv8m | 45.47 | 26.94 | 57.86 | 46.39 |
| YOLOv9s | 45.38 | 26.95 | 56.95 | 46.19 |
| YOLO26s | 44.87 | 26.43 | 56.44 | 45.41 |
| YOLOv10s | 44.51 | 26.28 | 55.97 | 45.61 |
| YOLO11s | 43.64 | 25.88 | 54.48 | 45.01 |
| YOLOv8s | 43.47 | 25.77 | 56.09 | 44.51 |
| YOLOv9t | 40.67 | 23.73 | 52.84 | 41.78 |
| YOLO26n | 39.9 | 22.95 | 51.14 | 42.08 |
| YOLOv10n | 39.8 | 23.08 | 51.22 | 41.5 |
| YOLOv8n | 39.69 | 23.03 | 52.02 | 41.36 |
| RF-DETR Medium | 39.62 | 21.99 | 70.62 | 46.76 |
| YOLO11n | 39.52 | 23.0 | 51.49 | 41.02 |
| RF-DETR Small | 39.13 | 21.68 | 64.38 | 49.47 |
| RF-DETR Nano | 37.92 | 20.88 | 69.02 | 46.21 |
| Class | mAP@50 | mAP@50-95 |
|---|---|---|
| pedestrian | 51.54 | 22.38 |
| people | 33.37 | 12.8 |
| bicycle | 26.07 | 12.27 |
| car | 84.49 | 55.5 |
| van | 52.11 | 36.4 |
| truck | 59.7 | 40.56 |
| tricycle | 33.98 | 19.88 |
| awning-tricycle | 27.82 | 18.02 |
| bus | 69.55 | 51.08 |
| motor | 52.44 | 23.56 |
| others | 0.0 | 0.0 |

This model was trained on VisDrone-DET. For the full dataset description, provenance, license, and citation, see the dataset card:
https://huggingface.co/datasets/Voxel51/VisDrone2019-DET
| Setting | Value |
|---|---|
| Dataset | VisDrone-DET |
| Framework | Ultralytics YOLO |
| Training Toolkit | DetectionBench |
| Epochs (configured max) | 100 |
| Epochs (actually trained) | 75 |
| Early Stopping Patience | 25 |
| Batch Size | 4 |
| Image Size | 1280 |
| Optimizer | SGD |
| Initial Learning Rate | 0.01 |
| Seed | 0 |
best.pt
results.csv
args.yaml
BoxPR_curve.png
BoxF1_curve.png
BoxP_curve.png
BoxR_curve.png
confusion_matrix.png
confusion_matrix_normalized.png
val_batch0_pred.jpg
visdrone_yolo26m_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.md
This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.
Features include:
If you find this model useful, please consider starring the repository.
car (42.21%) and pedestrian (23.12%) account for two-thirds of all annotated boxes in the training set, while awning-tricycle (0.95%) and tricycle (1.40%) are rare -- the others class has zero annotated instances in the training set entirely and is effectively unusable (always 0 AP).If you use this model in your research, please consider citing the dataset and the model architecture:
@article{zhu2018vision,
title={Vision meets drones: A challenge},
author={Zhu, Pengfei and Wen, Longyin and Bian, Xiao and Ling, Haibin and Hu, Qinghua},
journal={arXiv preprint arXiv:1804.07437},
year={2018}
}
@article{jocher2026yolo26,
title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models},
author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat},
journal={arXiv preprint arXiv:2606.03748},
year={2026}
}