--- license: mit library_name: transformers tags: - object-detection - vision - conservation - aerial-imagery - drones - pytorch --- # DroneMegaDetector (Transformers Version) This repository hosts the Hugging Face `transformers`-compatible version of **DroneMegaDetector**, an object detection model optimized for conservation drone imagery. The model has been converted from its original format to seamlessly integrate with the Hugging Face ecosystem, allowing for easy loading, inference, and deployment using standard `AutoImageProcessor` and `AutoModelForObjectDetection` classes. **Original GitHub Repository:** [ConservationDronesAI/DroneMegaDetector](https://github.com/ConservationDronesAI/DroneMegaDetector?tab=readme-ov-file) ## Intended Use * **Primary Use Case:** Detecting objects of interest (e.g., wildlife, humans, vehicles) in aerial imagery captured by drones for conservation, ecology, and anti-poaching efforts. * **Input:** High-resolution RGB images. * **Output:** Bounding boxes, confidence scores, and class labels for detected objects. ## Installation Ensure you have the required dependencies installed to run the model. You can install the exact environment tested for this model using the repository's `requirements.txt`: ```bash pip install torch==2.14.0 torchvision==0.29.0 transformers==5.18.0 Pillow==12.3.0 ``` ## Quickstart & Inference You can run inference using the standard Hugging Face pipeline. The following example code demonstrates how to load the model, process a source image, filter for detections with a confidence score of `>= 0.3`, and save an annotated output image with bounding boxes. ```python import os import sys import torch from PIL import Image, ImageDraw, ImageFont from transformers import AutoImageProcessor, AutoModelForObjectDetection # Hugging Face repository repo_id = "ConservationDrones/DroneMegaDetector" # Load the demo image (defaults to "example.png" if no argument is provided) image_path = sys.argv[1] if len(sys.argv) > 1 else "example.png" image = Image.open(image_path).convert("RGB") # Load the processor and model from Hugging Face processor = AutoImageProcessor.from_pretrained(repo_id) model = AutoModelForObjectDetection.from_pretrained(repo_id).eval() # Optimize CPU threading for local inference torch.set_num_threads(4) # Preprocess the image and run inference encoded = processor(images=image, return_tensors="pt") with torch.inference_mode(): outputs = model(**encoded) # Keep scores above 0.3 and map boxes back to the original image dimensions results = processor.post_process_object_detection( outputs, threshold=0.3, target_sizes=[(image.height, image.width)], )[0] # Print detection results print(f"{image_path}: {len(results['scores'])} detections (score >= 0.3)") for score, label, box in zip(results["scores"], results["labels"], results["boxes"]): print( model.config.id2label[label.item()], f"{score.item():.3f}", [round(v, 1) for v in box.tolist()], ) # Draw the detections on the original image draw = ImageDraw.Draw(image) font = ImageFont.load_default(size=18) for score, label, box in zip(results["scores"], results["labels"], results["boxes"]): bounds = box.tolist() text = f"{model.config.id2label[label.item()]} {score.item():.2f}" # Draw bounding box draw.rectangle(bounds, outline="red", width=3) # Position and draw text label text_x = min(max(2, bounds[0]), image.width - draw.textlength(text, font=font) - 2) draw.text( (text_x, max(0, bounds[1] - 20)), text, font=font, fill="white", stroke_width=2, stroke_fill="black", ) # Save the annotated image output_path = "example_annotated.jpg" image.save(output_path, quality=95) print(f"Saved {output_path}") ``` ## Limitations & Best Practices * **Confidence Thresholding:** The default demonstration filters out predictions below a `0.3` confidence score. You may need to tune this threshold up or down based on your specific drone altitude, camera resolution, and environment. * **Inference Speed:** For optimal performance on large batches of aerial images, running inference on a CUDA-enabled GPU is highly recommended, though CPU inference is supported (and configured to use 4 threads in the demo script). * **Image Dimensions:** The `AutoImageProcessor` handles automatic resizing and normalization. However, for extremely high-resolution drone panoramas, consider splitting the image into tiles before passing them to the model to avoid missing small objects (like distant animals). ## Citation If you use this model in your research or conservation project, please consider linking back to the original GitHub repository: > `https://github.com/ConservationDronesAI/DroneMegaDetector` ``` ```