dronefreak/visdrone-yolov10n
YOLOv10n Finetuned on VisDrone-DET
Fine-tuned YOLOv10n object detector on the VisDrone-DET benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.
<!-- Demo banner: side-by-side video of this checkpoint's detections on two VisDrone-DET test clips. Media lives under assets/ in this repo. The <video> renders on the Hugging Face model page (absolute resolve/ URL); on GitHub the nested <img> poster is shown instead. --> <p align="center"><video controls autoplay loop muted playsinline width="900" poster="https://huggingface.co/dronefreak/visdrone-yolov10n/resolve/main/assets/demobannerposter.jpg" src="https://huggingface.co/dronefreak/visdrone-yolov10n/resolve/main/assets/demobanner.mp4"><img src="https://huggingface.co/dronefreak/visdrone-yolov10n/resolve/main/assets/demobanner_poster.jpg" alt="YOLOv10n detections on two VisDrone-DET test clips" width="900"></video></p>
<br>
<!-- ROW 1: Identity & Tech Stack --> <div style="display: flex; justify-content: center; align-items: center; gap: 8px; margin-bottom: 8px; flex-wrap: wrap;"> <img src="https://img.shields.io/badge/Task-ObjectDetection-blue?style=flat-square" alt="Task"> <img src="https://img.shields.io/badge/Framework-UltralyticsYOLO-0aa1a7?style=flat-square" alt="Framework"> <img src="https://img.shields.io/badge/Base_Model-YOLOv10n-purple?style=flat-square" alt="Base Model"> </div>
<!-- ROW 2: Performance Metrics --> <div style="display: flex; justify-content: center; align-items: center; gap: 8px; margin-bottom: 8px; flex-wrap: wrap;"> <img src="https://img.shields.io/badge/mAP@50-39.8%25-success?style=flat-square" alt="mAP@50"> <img src="https://img.shields.io/badge/mAP@50:95-23.08%25-orange?style=flat-square" alt="mAP@50:95"> <img src="https://img.shields.io/badge/Params-2.8M-lightgrey?style=flat-square" alt="Params"> </div>
<!-- ROW 3: Metadata --> <div style="display: flex; justify-content: center; align-items: center; gap: 8px; margin-bottom: 24px; flex-wrap: wrap;"> <img src="https://img.shields.io/badge/License-AGPL--3.0-lightgrey?style=flat-square" alt="License"> <a href="https://github.com/dronefreak/DetectionBench"><img src="https://img.shields.io/badge/Source-DetectionBench-black?style=flat-square" alt="Source"></a> </div>
Performance
Evaluation Protocol
Metrics reported in this model card are computed on the VisDrone-DET test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).
VisDrone-DET Model Zoo
Every model DetectionBench has trained and evaluated on VisDrone-DET so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.
Earlier VisDrone-DET Results (Companion Codebase)
The rows below are earlier VisDrone2019-DET runs (test split) from a separate companion codebase (VisDrone-dataset-python-toolkit), not reproduced inside DetectionBench and kept here for context and history. Where a model also appears in the Model Zoo table above, that row is the current DetectionBench run and supersedes the one here -- for example the earlier YOLOv9t run used 640 px inputs and 300 epochs, while the current DetectionBench YOLO runs use 1280 px. The RF-DETR rows are earlier DetectionBench-trained runs at RF-DETR's default input sizes (384/512/576 px).
Source: https://huggingface.co/collections/dronefreak/visdrone-object-detection-model-zoo
Per-Class Performance
Evaluation Visualizations
Precision-Recall Curve
F1 Curve
Confusion Matrix
Normalized Confusion Matrix
Dataset
This model was trained on VisDrone-DET. For the full dataset description, provenance, license, and citation, see the dataset card:
https://huggingface.co/datasets/Voxel51/VisDrone2019-DET
Classes
- pedestrian
- people
- bicycle
- car
- van
- truck
- tricycle
- awning-tricycle
- bus
- motor
- others ---
Usage
Install Dependencies
pip install ultralytics huggingface_hubLoad Model from Hugging Face
from huggingface_hub import hf_hub_download
from ultralytics import YOLO
weights = hf_hub_download(
repo_id="dronefreak/visdrone-yolov10n",
filename="best.pt"
)
model = YOLO(weights)Run Inference
results = model.predict(
source="image.jpg",
conf=0.25
)
results[0].show()Training Configuration
Repository Contents
best.pt
results.csv
args.yaml
BoxPR_curve.png
BoxF1_curve.png
BoxP_curve.png
BoxR_curve.png
confusion_matrix.png
confusion_matrix_normalized.png
val_batch0_pred.jpg
visdrone_yolov10n_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.mdRelated Resources
- VisDrone-DET dataset card on Hugging Face
- DetectionBench -- reproducible benchmarks for modern object detectors on real-world datasets
Training Framework
This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.
Features include:
- A dataset-adapter registry for converting real-world datasets into a canonical format
- Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
- Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
- One-command reproducibility via versioned Hydra configs
If you find this model useful, please consider starring the repository.
Known Limitations
- Severe class imbalance:
car(42.21%) andpedestrian(23.12%) account for two-thirds of all annotated boxes in the training set, whileawning-tricycle(0.95%) andtricycle(1.40%) are rare -- theothersclass has zero annotated instances in the training set entirely and is effectively unusable (always 0 AP). - Extreme small-object density: ~53 annotated boxes per image on average, with roughly 69% of boxes covering under 0.1% of the image area -- consistent with VisDrone's aerial small-object detection challenge (objects captured from significant altitude).
- The original authors license VisDrone under CC BY-NC-SA 3.0 -- non-commercial research use only (see the dataset's homepage); this applies to any model trained on it, not only the raw images.
- These RF-DETR checkpoints were trained/evaluated directly through DetectionBench. The YOLO/RT-DETR rows in the External VisDrone Model Zoo comparison below were trained via a separate companion codebase, not reproduced inside DetectionBench -- see that collection for their own training details and caveats. ---
Citation
If you use this model in your research, please consider citing:
- The VisDrone-DET dataset (see below)
- The original YOLOv10n architecture (see below)
- The other model architectures shown in the Model Zoo/External Comparison tables above, if you reference their results
- DetectionBench, the training/evaluation framework used to produce this checkpoint
@article{zhu2018vision,
title={Vision meets drones: A challenge},
author={Zhu, Pengfei and Wen, Longyin and Bian, Xiao and Ling, Haibin and Hu, Qinghua},
journal={arXiv preprint arXiv:1804.07437},
year={2018}
}@article{wang2024yolov10,
title={YOLOv10: Real-Time End-to-End Object Detection},
author={Wang, Ao and Chen, Hui and Liu, Lihao and Chen, Kai and Lin, Zijia and Han, Jungong and Ding, Guiguang},
journal={arXiv preprint arXiv:2405.14458},
year={2024}
}Other architectures compared against on VisDrone-DET in this model card:
RF-DETR
@inproceedings{robinson2026rfdetr,
title = {RF-DETR: Real-Time Detection Transformer},
author = {Robinson, Isaac and Robicheaux, Peter and Popov, Matvei and Ramanan, Deva and Peri, Neehar},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026},
url = {https://arxiv.org/abs/2511.09554}
}
@article{oquab2023dinov2,
title={DINOv2: Learning Robust Visual Features without Supervision},
author={Oquab, Maxime and Darcet, Timoth{\'e}e and Moutakanni, Theo and Vo, Huy and Szafraniec, Marc and Khalidov, Vasil and Fernandez, Pierre and Haziza, Daniel and Massa, Francisco and El-Nouby, Alaaeldin and others},
journal={arXiv preprint arXiv:2304.07193},
year={2023}
}YOLOv11
No official YOLO11 research paper has been published by Ultralytics; the most commonly cited independent architectural analysis is used instead:
@article{khanam2024yolov11,
title={YOLOv11: An Overview of the Key Architectural Enhancements},
author={Khanam, Rahima and Hussain, Muhammad},
journal={arXiv preprint arXiv:2410.17725},
year={2024}
}YOLOv26
@article{jocher2026yolo26,
title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models},
author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat},
journal={arXiv preprint arXiv:2606.03748},
year={2026}
}YOLOv8
No official YOLOv8 research paper has been published by Ultralytics; this is their own recommended software citation instead:
@software{jocher2023yolov8,
author = {Glenn Jocher and Ayush Chaurasia and Jing Qiu},
title = {Ultralytics YOLOv8},
version = {8.0.0},
year = {2023},
url = {https://github.com/ultralytics/ultralytics},
license = {AGPL-3.0}
}YOLOv9
@article{wang2024yolov9,
title={YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information},
author={Wang, Chien-Yao and Yeh, I-Hau and Liao, Hong-Yuan Mark},
journal={arXiv preprint arXiv:2402.13616},
year={2024}
}@software{Saksena_DetectionBench_2026,
author = {Saksena, Saumya Kumaar},
title = {DetectionBench: Reproducible Benchmarks for Modern Object Detectors on Real-World Datasets},
url = {https://github.com/dronefreak/DetectionBench},
year = {2026}
}