DroneMegaDetector (Transformers Version)

This repository hosts the Hugging Face transformers-compatible version of DroneMegaDetector, an object detection model optimized for conservation drone imagery. The model has been converted from its original format to seamlessly integrate with the Hugging Face ecosystem, allowing for easy loading, inference, and deployment using standard AutoImageProcessor and AutoModelForObjectDetection classes.

Original GitHub Repository: ConservationDronesAI/DroneMegaDetector

Intended Use

  • Primary Use Case: Detecting objects of interest (e.g., wildlife, humans, vehicles) in aerial imagery captured by drones for conservation, ecology, and anti-poaching efforts.
  • Input: High-resolution RGB images.
  • Output: Bounding boxes, confidence scores, and class labels for detected objects.

Installation

Ensure you have the required dependencies installed to run the model. You can install the exact environment tested for this model using the repository's requirements.txt:

pip install torch==2.14.0 torchvision==0.29.0 transformers==5.18.0 Pillow==12.3.0

Quickstart & Inference

You can run inference using the standard Hugging Face pipeline.

The following example code demonstrates how to load the model, process a source image, filter for detections with a confidence score of >= 0.3, and save an annotated output image with bounding boxes.

import os
import sys
import torch
from PIL import Image, ImageDraw, ImageFont
from transformers import AutoImageProcessor, AutoModelForObjectDetection

# Hugging Face repository
repo_id = "ConservationDrones/DroneMegaDetector"

# Load the demo image (defaults to "example.png" if no argument is provided)
image_path = sys.argv[1] if len(sys.argv) > 1 else "example.png"
image = Image.open(image_path).convert("RGB")

# Load the processor and model from Hugging Face
processor = AutoImageProcessor.from_pretrained(repo_id)
model = AutoModelForObjectDetection.from_pretrained(repo_id).eval()

# Optimize CPU threading for local inference
torch.set_num_threads(4)

# Preprocess the image and run inference
encoded = processor(images=image, return_tensors="pt")

with torch.inference_mode():
    outputs = model(**encoded)

# Keep scores above 0.3 and map boxes back to the original image dimensions
results = processor.post_process_object_detection(
    outputs,
    threshold=0.3,
    target_sizes=[(image.height, image.width)],
)[0]

# Print detection results
print(f"{image_path}: {len(results['scores'])} detections (score >= 0.3)")
for score, label, box in zip(results["scores"], results["labels"], results["boxes"]):
    print(
        model.config.id2label[label.item()],
        f"{score.item():.3f}",
        [round(v, 1) for v in box.tolist()],
    )

# Draw the detections on the original image
draw = ImageDraw.Draw(image)
font = ImageFont.load_default(size=18)

for score, label, box in zip(results["scores"], results["labels"], results["boxes"]):
    bounds = box.tolist()
    text = f"{model.config.id2label[label.item()]} {score.item():.2f}"
    
    # Draw bounding box
    draw.rectangle(bounds, outline="red", width=3)
    
    # Position and draw text label
    text_x = min(max(2, bounds[0]), image.width - draw.textlength(text, font=font) - 2)
    draw.text(
        (text_x, max(0, bounds[1] - 20)),
        text,
        font=font,
        fill="white",
        stroke_width=2,
        stroke_fill="black",
    )

# Save the annotated image
output_path = "example_annotated.jpg"
image.save(output_path, quality=95)
print(f"Saved {output_path}")

Limitations & Best Practices

  • Confidence Thresholding: The default demonstration filters out predictions below a 0.3 confidence score. You may need to tune this threshold up or down based on your specific drone altitude, camera resolution, and environment.

  • Inference Speed: For optimal performance on large batches of aerial images, running inference on a CUDA-enabled GPU is highly recommended, though CPU inference is supported (and configured to use 4 threads in the demo script).

  • Image Dimensions: The AutoImageProcessor handles automatic resizing and normalization. However, for extremely high-resolution drone panoramas, consider splitting the image into tiles before passing them to the model to avoid missing small objects (like distant animals).

Citation

If you use this model in your research or conservation project, please consider linking back to the original GitHub repository:

https://github.com/ConservationDronesAI/DroneMegaDetector


Downloads last month
37
Safetensors
Model size
35.3M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support