Object-Intelligence-Backend / docs /getting-started.md
muhammadpriv001's picture
Frontend 1.0.0
4346a4c
|
Raw History Blame Contribute Delete
1.64 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade

Getting Started with Object Intelligence Platform

Welcome to the Open-World Object Intelligence Platform. This platform unifies three distinct computer vision paradigms into a single modular architecture:

  1. Known Object Detection (RT-DETR): Ultra-fast, predictable detection for standard trained classes (COCO).
  2. Open-Vocabulary Discovery (YOLO-World): Zero-shot visual concept detection driven by text prompts.
  3. Specific Object Recognition (Visual Embeddings): Teachable feature embeddings to recognize individual physical objects (e.g. My Cup).

πŸš€ Quick Setup

1. Installation

Ensure Python 3.9+ is installed, then run:

pip install -r requirements.txt

2. Launching the Gradio Web Application & API

To start the Gradio interface locally or prepare for Hugging Face Spaces deployment:

python app.py

Open your browser at http://localhost:7860.


πŸ“¦ Python SDK Usage

from object_intelligence import ObjectDetector
import cv2

# Initialize unified detector
detector = ObjectDetector()

# Select mode: 'combined', 'known', 'open_vocabulary', 'specific'
detector.set_mode("combined")

# Enable Lock Mode if desired
detector.lock("My Cup")

# Read frame and run detection
frame = cv2.imread("test.jpg")
annotated_frame, detections = detector.detect_and_draw(frame)

# Save result
cv2.imwrite("output.jpg", annotated_frame)

πŸ”’ Lock Mode

Lock Mode filters all candidate detections across all pipelines, rendering ONLY bounding boxes that match your specified target string. All non-matching detections are suppressed before drawing.