RF-DETR Medium Finetuned on VisDrone-DET
Fine-tuned RF-DETR Medium object detector on the VisDrone-DET benchmark dataset, trained and evaluated as part of DetectionBench -- a framework for reproducibly benchmarking modern object detectors with identical training recipes and evaluation metrics across multiple real-world datasets.
Usage
Install Dependencies
pip install rfdetr huggingface_hub
Load Model from Hugging Face
from huggingface_hub import hf_hub_download
import rfdetr
weights = hf_hub_download(
repo_id="dronefreak/visdrone-rfdetr-medium",
filename="checkpoint_best_total.pth"
)
model = rfdetr.RFDETRMedium(pretrain_weights=weights)
Run Inference
detections = model.predict("image.jpg", threshold=0.25)
Performance
Evaluated on the VisDrone-DET test split, using DetectionBench's standard evaluation pipeline (detectionbench-evaluate).
| Metric | Score (%) |
|---|---|
| mAP@50 | 39.62 |
| mAP@50-95 | 21.99 |
| Precision | 70.62 |
| Recall | 46.76 |
| F1 Score | 56.27 |
| Parameters | 33.7M |
| FLOPs | N/A (not published upstream) |
VisDrone-DET Model Zoo
Every model DetectionBench has trained and evaluated on VisDrone-DET so far, for full transparency -- see DetectionBench for the smaller, curated comparison set used on the project README.
| Model | mAP@50 | mAP@50-95 | Precision | Recall |
|---|---|---|---|---|
| YOLO26m | 49.11 | 29.24 | 60.21 | 49.76 |
| YOLO11m | 48.03 | 28.56 | 59.34 | 48.3 |
| YOLOv9m | 47.62 | 28.52 | 59.26 | 48.4 |
| YOLOv10m | 46.73 | 27.72 | 59.23 | 47.35 |
| YOLOv8m | 45.47 | 26.94 | 57.86 | 46.39 |
| YOLOv9s | 45.38 | 26.95 | 56.95 | 46.19 |
| YOLO26s | 44.87 | 26.43 | 56.44 | 45.41 |
| YOLOv10s | 44.51 | 26.28 | 55.97 | 45.61 |
| YOLO11s | 43.64 | 25.88 | 54.48 | 45.01 |
| YOLOv8s | 43.47 | 25.77 | 56.09 | 44.51 |
| YOLOv9t | 40.67 | 23.73 | 52.84 | 41.78 |
| YOLO26n | 39.9 | 22.95 | 51.14 | 42.08 |
| YOLOv10n | 39.8 | 23.08 | 51.22 | 41.5 |
| YOLOv8n | 39.69 | 23.03 | 52.02 | 41.36 |
| RF-DETR Medium | 39.62 | 21.99 | 70.62 | 46.76 |
| YOLO11n | 39.52 | 23.0 | 51.49 | 41.02 |
| RF-DETR Small | 39.13 | 21.68 | 64.38 | 49.47 |
| RF-DETR Nano | 37.92 | 20.88 | 69.02 | 46.21 |
Per-Class Performance
| Class | mAP@50 | mAP@50-95 |
|---|---|---|
| pedestrian | 31.03 | 12.36 |
| people | 25.29 | 9.08 |
| bicycle | 18.6 | 7.41 |
| car | 75.62 | 46.26 |
| van | 42.95 | 26.97 |
| truck | 50.0 | 32.35 |
| tricycle | 27.39 | 13.83 |
| awning-tricycle | 23.19 | 13.06 |
| bus | 64.29 | 43.83 |
| motor | 37.89 | 14.7 |
| others | 0.0 | 0.0 |
This model was evaluated with Supervision's detection metrics, which report mAP/Precision/Recall directly but don't produce a confusion-matrix plot the way Ultralytics' validator does.
Dataset
This model was trained on VisDrone-DET. For the full dataset description, provenance, license, and citation, see the dataset card:
https://huggingface.co/datasets/Voxel51/VisDrone2019-DET
Classes
- pedestrian
- people
- bicycle
- car
- van
- truck
- tricycle
- awning-tricycle
- bus
- motor
- others
Training Configuration
| Setting | Value |
|---|---|
| Dataset | VisDrone-DET |
| Framework | RF-DETR |
| Training Toolkit | DetectionBench |
| Epochs (configured max) | 100 |
| Epochs (actually trained) | 48 |
| Early Stopping Patience | 20 |
| Batch Size | 12 |
| Resolution | 640 |
| Optimizer | adamw |
| Learning Rate | 0.0001 |
| Seed | 42 |
Repository Contents
checkpoint_best_total.pth
metrics.csv
config.json
visdrone_rfdetr-medium_showcase.jpg
assets/demo_banner.mp4
assets/demo_banner_poster.jpg
README.md
Related Resources
- VisDrone-DET dataset card on Hugging Face
- DetectionBench -- reproducible benchmarks for modern object detectors on real-world datasets
Training Framework
This model was trained using DetectionBench, an open-source framework for benchmarking object detectors across multiple real-world datasets with a common pipeline.
Features include:
- A dataset-adapter registry for converting real-world datasets into a canonical format
- Identical training/evaluation recipes across model families (Ultralytics YOLO/RT-DETR, RF-DETR)
- Hardware profiling (latency, FPS, VRAM, parameters, FLOPs)
- One-command reproducibility via versioned Hydra configs
If you find this model useful, please consider starring the repository.
Known Limitations
- Severe class imbalance:
car(42.21%) andpedestrian(23.12%) account for two-thirds of all annotated boxes in the training set, whileawning-tricycle(0.95%) andtricycle(1.40%) are rare -- theothersclass has zero annotated instances in the training set entirely and is effectively unusable (always 0 AP). - Extreme small-object density: ~53 annotated boxes per image on average, with roughly 69% of boxes covering under 0.1% of the image area -- consistent with VisDrone's aerial small-object detection challenge (objects captured from significant altitude).
- The original authors license VisDrone under CC BY-NC-SA 3.0 -- non-commercial research use only (see the dataset's homepage); this applies to any model trained on it, not only the raw images.
- These RF-DETR checkpoints were trained/evaluated directly through DetectionBench. The YOLO/RT-DETR rows in the External VisDrone Model Zoo comparison below were trained via a separate companion codebase, not reproduced inside DetectionBench -- see that collection for their own training details and caveats.
Citation
If you use this model in your research, please consider citing the dataset and the model architecture:
@article{zhu2018vision,
title={Vision meets drones: A challenge},
author={Zhu, Pengfei and Wen, Longyin and Bian, Xiao and Ling, Haibin and Hu, Qinghua},
journal={arXiv preprint arXiv:1804.07437},
year={2018}
}
@inproceedings{robinson2026rfdetr,
title = {RF-DETR: Real-Time Detection Transformer},
author = {Robinson, Isaac and Robicheaux, Peter and Popov, Matvei and Ramanan, Deva and Peri, Neehar},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026},
url = {https://arxiv.org/abs/2511.09554}
}
@article{oquab2023dinov2,
title={DINOv2: Learning Robust Visual Features without Supervision},
author={Oquab, Maxime and Darcet, Timoth{\'e}e and Moutakanni, Theo and Vo, Huy and Szafraniec, Marc and Khalidov, Vasil and Fernandez, Pierre and Haziza, Daniel and Massa, Francisco and El-Nouby, Alaaeldin and others},
journal={arXiv preprint arXiv:2304.07193},
year={2023}
}
- Downloads last month
- 66
Model tree for dronefreak/visdrone-rfdetr-medium
Base model
Roboflow/rf-detr-mediumDataset used to train dronefreak/visdrone-rfdetr-medium
Collection including dronefreak/visdrone-rfdetr-medium
Papers for dronefreak/visdrone-rfdetr-medium
RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
DINOv2: Learning Robust Visual Features without Supervision
Vision Meets Drones: A Challenge
Evaluation results
- mAP@50 (test split) on VisDrone-DETDetectionBench39.620
- mAP@50-95 (test split) on VisDrone-DETDetectionBench21.990
- Precision (test split) on VisDrone-DETDetectionBench70.620
- Recall (test split) on VisDrone-DETDetectionBench46.760