Object-Intelligence-Backend / PROJECT_DOCUMENTATION.md
muhammadpriv001's picture
docs: comprehensive project documentation for portfolio
fdced5d
|
Raw History Blame Contribute Delete
35.6 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade

Object Intelligence Platform

Open-World Multi-Pipeline Computer Vision System Real-time object detection, open-vocabulary discovery, specific object recognition, and target lock filtering β€” all in one unified platform.

Hugging Face Spaces Live Demo Python 3.10+ License: MIT


Table of Contents

  1. Project Overview
  2. Motivation & Intent
  3. Key Features
  4. System Architecture
  5. ML Pipeline Deep Dive
  6. Backend Architecture
  7. Frontend Architecture
  8. Database Design
  9. API Reference
  10. Lock Mode β€” Target Filtering
  11. Deployment Architecture
  12. Getting Started
  13. Python SDK
  14. OpenCV Examples
  15. Tech Stack
  16. Project Structure
  17. Challenges & Solutions
  18. Future Work
  19. License

1. Project Overview

Object Intelligence Platform is a full-stack computer vision system that combines three distinct detection pipelines into a single unified framework:

Pipeline Model Purpose Capability
Known Objects RT-DETR-L Detect standard objects 80 COCO classes with high accuracy
Open Vocabulary YOLO-World v2 Discover objects by text description Zero-shot detection via natural language prompts
Specific Objects ResNet-18 Embeddings Recognize YOUR personal objects Teach and identify "My Cup" vs any generic cup

The system fuses results from all three pipelines, resolves conflicts using IoU-based NMS with priority weighting, and supports a unique Lock Mode that suppresses all non-target detections in real-time.

Live Links:


2. Motivation & Intent

The Problem

Traditional object detection systems are rigid β€” they detect what they were trained on. A YOLO model trained on COCO can tell you "there is a cup" but cannot tell you "there is YOUR cup." Open-vocabulary models like YOLO-World solve part of this by accepting text prompts, but they still treat all instances of a class as identical.

Our Vision

We wanted to build a system that understands objects at three levels of identity:

  1. Category level β€” "This is a laptop" (RT-DETR, known classes)
  2. Concept level β€” "This matches the description 'wireless mouse'" (YOLO-World, text prompts)
  3. Instance level β€” "This is specifically MY laptop, not just any laptop" (ResNet-18 embeddings)

By combining all three, the platform achieves a richer understanding of visual scenes that approaches how humans perceive and track objects β€” recognizing both what things are and whose they are.

Portfolio Intent

This project demonstrates:

  • ML Engineering: Multi-model pipeline orchestration, embedding-based retrieval, real-time inference
  • Full-Stack Development: Next.js 16 + React 19 frontend, FastAPI + Gradio backend, SQLite persistence
  • System Design: Microservice-like architecture with clean separation of concerns
  • Deployment: GPU-accelerated cloud inference (HuggingFace ZeroGPU) + edge frontend (Vercel)
  • Product Thinking: Lock Mode as a UX innovation for focused object tracking

3. Key Features

Multi-Pipeline Detection

Run any combination of RT-DETR, YOLO-World, and embedding recognition simultaneously. Results are fused using IoU-based Non-Maximum Suppression with priority weighting (specific > known > open vocabulary).

Open Vocabulary Discovery

Type any text description β€” "red cup," "wireless mouse," "screwdriver" β€” and YOLO-World finds matching objects without retraining. Default vocabulary includes 20 common objects; fully customizable at runtime.

Specific Object Recognition

Teach the system your personal objects by uploading 3-10 reference photos. A ResNet-18 backbone extracts 512-dimensional L2-normalized feature vectors. During detection, each bounding box crop is compared against stored embeddings using cosine similarity. If the match exceeds the threshold (0.70), the detection is labeled with your object's identity.

Lock Mode Target Filtering

Enable Lock Mode and specify a target (e.g., "My Cup"). The system suppresses ALL non-matching detections β€” only bounding boxes matching your target are displayed. Uses fuzzy matching with synonym dictionaries and word overlap logic.

Real-Time AR Camera Mode

WebRTC webcam stream with canvas overlay rendering. Bounding boxes are drawn in real-time with color-coded corner brackets, fill overlays, and detection labels. FPS counter and detection count displayed live.

Teaching Studio

Upload reference photos via drag-and-drop or live camera snapshots. Objects are stored in a portable SQLite database with full CRUD operations and ZIP export/import.

Python SDK

Clean object_intelligence package with ObjectDetector class for programmatic access. Supports all modes, lock mode, teaching, and OpenCV integration.


4. System Architecture

                           USER INPUT
                     (Image / Webcam Frame)
                              |
               +--------------+--------------+
               |              |              |
               v              v              v
        +-----------+  +-----------+  +------------+
        |  RT-DETR  |  | YOLO-World|  |  ResNet-18 |
        |  (Known)  |  | (Open Voc)|  | (Specific) |
        +-----------+  +-----------+  +------------+
               |              |              |
               v              v              v
        Known Object    Open Vocabulary  Specific Object
        Detections      Detections       Matches
               |              |              |
               +--------------+--------------+
                              |
                              v
                   +-------------------+
                   |  RESULT FUSION    |
                   |  IoU NMS + Priority|
                   |  Deduplication     |
                   +-------------------+
                              |
                   +----------+----------+
                   |                     |
                   v                     v
              Normal Mode          Lock Mode
              (All Objects)     (Target Only)
                   |                     |
                   v                     v
              +--------------------------------+
              |    VISUAL ANNOTATION           |
              |  Color-coded bounding boxes    |
              |  Lock status banners           |
              +--------------------------------+
                              |
                              v
                   +-------------------+
                   |  API Response     |
                   |  Annotated Image  |
               |   Detection Metadata   |
                   +-------------------+

Pipeline Modes

Mode RT-DETR YOLO-World Embeddings Use Case
combined Yes Yes Yes Full analysis (default)
known Yes No No Fast COCO-only detection
open_vocabulary No Yes No Text-prompted discovery
specific No No Yes Teach/recognize personal objects

5. ML Pipeline Deep Dive

5.1 RT-DETR β€” Known Object Detection

Model: RT-DETR-L (Real-Time Detection Transformer, Large variant) Weights: rtdetr-l.pt (~63 MB) Framework: Ultralytics Classes: 80 COCO standard classes (person, car, chair, bottle, laptop, etc.)

Input Image β†’ RT-DETR-L β†’ Bounding Boxes + Labels + Confidence

RT-DETR is a transformer-based detector that achieves real-time performance with accuracy comparable to slower two-stage detectors. The "L" variant uses a larger backbone for better feature extraction.

Characteristics:

  • High confidence threshold (0.60) for reliable detections
  • Strong on common objects (furniture, electronics, people, vehicles)
  • Fallback to YOLOv8s if RT-DETR fails to load
  • Detection output: type="known", source="rtdetr"

5.2 YOLO-World β€” Open Vocabulary Discovery

Model: YOLO-World v2 (S variant) Weights: yolov8s-worldv2.pt (~25 MB) Framework: Ultralytics Default Prompts: 20 text descriptions loaded from open_vocabulary_classes.txt

Input Image + Text Prompts β†’ YOLO-World β†’ Bounding Boxes + Labels + Confidence

YOLO-World performs zero-shot detection β€” it finds objects matching any text description without task-specific training. The vocabulary is set dynamically via model.set_classes(prompts).

Default Vocabulary (20 classes):

person, cup, laptop, smartphone, water bottle, backpack, screwdriver,
coffee mug, wireless mouse, keyboard, chair, headphone, book, keychain,
glass, glasses, jacket, umbrella, helmet, desk lamp

Characteristics:

  • Lower confidence threshold (0.35) to catch more possibilities
  • Vocabulary is fully customizable at runtime via API or UI
  • Fallback to YOLOv8s if YOLO-World fails to load
  • Detection output: type="open_vocabulary", source="yolo_world"

5.3 ResNet-18 β€” Specific Object Recognition

Model: ResNet-18 (pretrained on ImageNet) Backbone: torchvision.models.resnet18 with final FC layer removed Output: 512-dimensional L2-normalized feature vector Similarity: Cosine similarity (dot product of normalized vectors)

Reference Photos β†’ ResNet-18 β†’ L2-Normalized Embeddings β†’ SQLite Storage
                                                                     |
Detection Crop β†’ ResNet-18 β†’ Query Embedding β†’ Cosine Search β†’ Match?

Teaching Flow:

  1. User provides object name + 3-10 reference photos
  2. Each photo is resized to 224x224, normalized with ImageNet stats
  3. Forward pass through ResNet-18 (minus FC layer) produces 512-dim vector
  4. Vector is L2-normalized and stored in object_embeddings table as JSON

Recognition Flow:

  1. After RT-DETR/YOLO-World produce bounding boxes
  2. Each bounding box is cropped from the original image
  3. Crop is passed through ResNet-18 to get a query embedding
  4. Brute-force cosine similarity search over ALL stored embeddings
  5. If best similarity >= threshold (0.70), detection label is overridden with the specific object name

Characteristics:

  • Threshold: 0.70 similarity for a match
  • Works best with 3-10 diverse reference photos (different angles, lighting)
  • Embeddings are viewpoint-invariant to a degree (ResNet features capture shape, texture, color patterns)
  • Detection output: type="specific", source="embedding", specific_identity="Object Name"

5.4 Result Fusion Engine

The ResultFusionEngine merges detection lists from multiple pipelines:

  1. Priority Sorting: Detections are sorted by type priority (specific=3 > known=2 > open_vocabulary=1), then by confidence
  2. IoU Deduplication: For each detection pair, if IoU >= 0.50, the lower-priority detection is removed
  3. Smart Label Override: If a specific detection overlaps a known/open_vocabulary detection, the existing detection's label is replaced with the specific identity

This ensures that when the system recognizes "My Cup" in a bounding box also detected as "cup" by RT-DETR, the specific identity takes precedence.

5.5 Lock Mode Controller

The LockModeController filters detections based on a target string:

  1. Fuzzy Matching:

    • Strips common prefixes ("my ", "the ", "a ", "an ")
    • Normalizes labels, categories, and specific_identity fields
    • Synonym dictionary: "mobile" = "cell phone" = "phone" = "smartphone"
    • Word overlap matching: "My Cup" matches "coffee cup" (cup is common)
  2. Filtering: Only detections matching the target are kept; all others are suppressed

  3. Visual Feedback:

    • Lock active + match found: Purple lock banner
    • Lock active + NOT FOUND: Red "NOT FOUND" banner
    • Color scheme shifts to pink/magenta when locked

6. Backend Architecture

Entry Point: app.py

The backend is a single app.py file that combines:

  1. Gradio Blocks UI β€” The interactive web interface with 5 tabs (Live Testing, Object Studio, Downloads, Guides, API Reference)
  2. FastAPI Routes β€” Registered via ASGI middleware that intercepts /api/* requests before Gradio's catch-all routes
  3. ZeroGPU Integration β€” @spaces.GPU decorators on key functions for GPU allocation on HuggingFace Spaces

ASGI Middleware

A custom CustomAPIMiddleware (Starlette BaseHTTPMiddleware) intercepts API routes:

Request β†’ CustomAPIMiddleware β†’ (matches /api/*?) β†’ Handle directly
                                            ↓ (no match)
                                   Gradio's route handler

This architecture solves the route conflict between FastAPI custom routes and Gradio's auto-generated /api/{api_name} catch-all routes. The middleware runs at the ASGI level, before FastAPI route matching.

Key Backend Components

File Class Purpose
backend/config.py β€” Central configuration: paths, thresholds, device detection
backend/ml/orchestrator.py ObjectIntelligenceOrchestrator 5-step pipeline orchestrator (lazy singleton)
backend/ml/rtdetr_detector.py RTDETRDetector RT-DETR known object detection
backend/ml/yolo_world_detector.py YOLOWorldDetector YOLO-World open vocabulary detection
backend/ml/embedding_recognizer.py EmbeddingRecognizer ResNet-18 feature extraction + cosine matching
backend/ml/fusion.py ResultFusionEngine IoU-based NMS + priority deduplication
backend/ml/lock_mode.py LockModeController Target filtering + visual annotation
backend/database/storage.py DatabaseManager SQLite CRUD, cosine search, ZIP export
backend/database/models.py β€” Dataclass definitions

ZeroGPU Compatibility

HuggingFace ZeroGPU requires CUDA operations to happen inside @spaces.GPU decorated functions. The orchestrator uses lazy loading β€” models are only instantiated when a GPU function is first called, ensuring all CUDA operations happen within ZeroGPU's allocation context.

# Models load HERE, inside @spaces.GPU context:
@spaces.GPU
def run_gradio_detection(...):
    annotated_bgr, detections, metadata = get_orchestrator().process_frame(...)
    # orchestrator is created on first call, models loaded with GPU available

7. Frontend Architecture

Tech Stack

  • Framework: Next.js 16.3.3 (App Router)
  • React: 19.2.8
  • TypeScript: 5.x
  • 3D Graphics: Three.js 0.185 (wireframe parallax scene)
  • Animations: GSAP 3.15
  • Icons: Lucide React

Pages

Route Page Description
/ Home Hero section, pipeline cards, 6-step visualization, Lock Mode preview
/live Live Testing Upload image or real-time AR camera with canvas overlay
/objects Object Studio Teach objects via file upload or camera snapshots, manage library
/downloads Downloads Model weights, OpenCV examples, database export
/guides Developer Guides 5 inline markdown guides with sidebar navigation

Design System

The frontend uses a custom neumorphism design system:

  • Colors: Primary blue (#1E90FF), emerald (#10b981), pink (#ec4899), dark background (#0a0e17)
  • Components: .neu-card, .neu-btn, .neu-input with dual-direction shadows
  • Animations: fadeInUp, slideInRight, scaleIn with stagger delays
  • Responsive: Desktop navbar hidden on mobile, full-screen animated mobile menu

3D Parallax Scene

The home page features a Three.js WebGL scene with:

  • 7 large background wireframe shapes (Icosahedron, Octahedron, Tetrahedron, Box, Dodecahedron)
  • 14 mid-layer shapes, 18 small front-layer shapes, 20 tiny scattered shapes
  • 6 bounding box outlines (CV-themed), 4 crosshair markers, 4 diamond detection markers
  • 300 particles, 120 twinkling stars
  • Mouse parallax tracking + click scatter physics

Camera AR Mode

Real-time webcam processing with canvas overlay:

  • WebRTC stream β†’ Canvas capture β†’ Backend detection β†’ Canvas overlay rendering
  • Color-coded corner brackets: specific=purple, open_vocab=blue, known=green, lock=pink
  • FPS counter and detection count displayed live
  • Continuous frame loop with requestAnimationFrame

API Communication

// lib/api.ts β€” constructs backend URLs
export function getApiUrl(path: string): string {
  const envUrl = process.env.NEXT_PUBLIC_API_URL || "";
  if (envUrl) {
    // Production: direct cross-origin to HuggingFace Space
    return `${envUrl}/api${path}`;
  }
  // Development: proxied via Next.js rewrites
  return `/backend-api${path}`;
}

8. Database Design

Schema (SQLite)

-- Core entity: a taught object
CREATE TABLE objects (
    id TEXT PRIMARY KEY,              -- "obj_<uuid_hex_10>"
    user_id TEXT,
    name TEXT UNIQUE NOT NULL,
    category TEXT NOT NULL,
    description TEXT,
    status TEXT DEFAULT 'active',
    created_at TEXT NOT NULL,
    updated_at TEXT NOT NULL
);

-- Reference images for each object
CREATE TABLE object_images (
    id TEXT PRIMARY KEY,              -- "img_<uuid_hex_10>"
    object_id TEXT NOT NULL,          -- FK β†’ objects.id
    file_path TEXT NOT NULL,
    image_hash TEXT,
    created_at TEXT NOT NULL
);

-- 512-dimensional embedding vectors (stored as JSON)
CREATE TABLE object_embeddings (
    id TEXT PRIMARY KEY,              -- "emb_<uuid_hex_10>"
    object_id TEXT NOT NULL,          -- FK β†’ objects.id
    model_name TEXT NOT NULL,         -- "resnet18"
    model_version TEXT NOT NULL,      -- "1.0"
    embedding TEXT NOT NULL,          -- JSON array of 512 floats
    created_at TEXT NOT NULL
);

-- Detection event logging
CREATE TABLE detection_logs (
    id TEXT PRIMARY KEY,              -- "log_<uuid_hex_10>"
    user_id TEXT,
    object_name TEXT NOT NULL,
    detection_type TEXT NOT NULL,     -- "known", "open_vocabulary", "specific"
    confidence REAL NOT NULL,
    timestamp TEXT NOT NULL
);

Key Operations

Operation SQL Description
Find best match Brute-force cosine similarity over all embeddings Returns object with similarity >= threshold
Export database ZIP(database.db + manifest.json) Portable library export
Cascade delete Delete embeddings β†’ images β†’ object Clean object removal

9. API Reference

All endpoints are intercepted by CustomAPIMiddleware before reaching Gradio's routes.

GET /api/health

Returns system health status and model load state.

{
  "status": "healthy",
  "models": {
    "rtdetr": true,
    "yolo_world": true,
    "embedding_recognizer": true
  },
  "database": true
}

POST /api/detect

Run detection pipeline on an uploaded image.

Request: multipart/form-data

Field Type Default Description
file File required Image file (JPEG, PNG, etc.)
mode string "combined" Pipeline mode: combined, known, open_vocabulary, specific
lock_mode string "false" Enable lock mode filtering
lock_target string "" Target object name for lock mode
open_vocab_prompts string "" Comma-separated YOLO-World prompts
confidence string "" Detection confidence threshold

Response:

{
  "status": "success",
  "image_base64": "<base64-encoded annotated PNG>",
  "metadata": {
    "mode": "combined",
    "lock_mode": false,
    "lock_target": null,
    "detection_count": 3,
    "detections": [...]
  },
  "detections": [
    {
      "bbox": [120, 80, 340, 290],
      "label": "My Cup",
      "category": "cup",
      "confidence": 0.94,
      "type": "specific",
      "source": "embedding",
      "specific_identity": "My Cup"
    }
  ]
}

GET /api/objects

List all taught objects with image and embedding counts.

POST /api/objects

Teach a new object with reference photos.

Request: multipart/form-data with name, category, description, files (multiple)

DELETE /api/objects/{object_id}

Delete a taught object and all associated data.

GET /api/downloads/export-db

Download the object database as a ZIP archive.


10. Lock Mode β€” Target Filtering

Concept

Lock Mode is a unique feature that restricts the detection output to only objects matching a user-specified target. When enabled:

  1. All detections are evaluated against the target string
  2. Non-matching detections are completely suppressed (not drawn, not returned)
  3. The visual overlay shifts to a pink/magenta color scheme
  4. A lock status banner appears (purple = match found, red = NOT FOUND)

Matching Algorithm

def is_match(detection, target):
    # 1. Strip common prefixes
    target = target.strip().lower()
    for prefix in ["my ", "the ", "a ", "an "]:
        if target.startswith(prefix):
            target = target[len(prefix):]
    
    # 2. Check against label, category, specific_identity
    for field in [label, category, specific_identity]:
        if field and target in field.lower():
            return True
    
    # 3. Synonym dictionary lookup
    if target in synonym_dict and any(s in field for s in synonym_dict[target]):
        return True
    
    # 4. Word overlap matching
    target_words = set(target.split())
    field_words = set(field.lower().split())
    if target_words.issubset(field_words) or field_words.issubset(target_words):
        return True
    
    return False

Synonym Dictionary

synonyms = {
    "mobile": ["cell phone", "phone", "smartphone"],
    "laptop": ["computer", "notebook"],
    "tv": ["television", "monitor", "screen"],
    "sofa": ["couch"],
    "cup": ["mug", "glass"],
    "backpack": ["bag", "knapsack"],
}

Use Cases

  • Warehouse picking: Lock onto a specific item SKU to guide workers
  • Security: Track a specific person or object across frames
  • Accessibility: Help visually impaired users locate specific items
  • Quality control: Focus on a specific defect type in manufacturing

11. Deployment Architecture

                    +---------------------------+
                    |      VERCEL (CDN)         |
                    |   Next.js 16 Frontend     |
                    |  object-intelligence.     |
                    |     vercel.app            |
                    +-------------+-------------+
                                  |
                          HTTPS API calls
                                  |
                    +-------------v-------------+
                    |  HUGGINGFACE SPACES (GPU) |
                    |   ZeroGPU T4 Instance     |
                    |   Gradio 4.44.1 + FastAPI |
                    |   RT-DETR + YOLO-World +  |
                    |   ResNet-18 + SQLite       |
                    +---------------------------+

HuggingFace Spaces (Backend)

  • SDK: Gradio 4.44.1
  • Hardware: GPU (T4 via ZeroGPU allocation)
  • Entry Point: app.py
  • Features:
    • @spaces.GPU decorators for GPU allocation
    • Lazy model loading to satisfy ZeroGPU initialization
    • Startup probe (_zerogpu_startup_probe) for GPU context
    • CORS enabled for cross-origin Vercel frontend
    • Auto-downloads model weights on first inference

Vercel (Frontend)

  • Framework: Next.js 16.3.3 with App Router
  • React: 19.2.8
  • Build: Static export with client-side API calls
  • Environment Variable: NEXT_PUBLIC_API_URL pointing to HuggingFace Space

ZeroGPU Integration Details

ZeroGPU dynamically allocates GPU time to functions decorated with @spaces.GPU. The key challenge was ensuring model loading happens inside GPU context:

# Problem: Models loaded at import time (no GPU)
orchestrator = ObjectIntelligenceOrchestrator()  # ← CUDA fails

# Solution: Lazy loading inside @spaces.GPU function
_orchestrator_instance = None
def get_orchestrator():
    global _orchestrator_instance
    if _orchestrator_instance is None:
        _orchestrator_instance = ObjectIntelligenceOrchestrator()  # ← GPU available
    return _orchestrator_instance

12. Getting Started

Prerequisites

  • Python 3.10+
  • Node.js 18+ (for frontend)
  • CUDA-capable GPU (optional, falls back to CPU)

Installation

# Clone repository
git clone https://github.com/muhammadpriv001/Object-Intelligence.git
cd Object-Intelligence

# Install Python dependencies
pip install -r requirements.txt

# Install frontend dependencies
cd frontend
npm install
cd ..

Running Locally

# Start backend (port 7860)
python app.py

# In a separate terminal, start frontend (port 3000)
cd frontend
npm run dev

Open http://localhost:3000 for the frontend, or http://localhost:7860 for the Gradio UI directly.

Environment Variables

Variable Default Description
NEXT_PUBLIC_API_URL "" (uses proxy) Backend API URL. Set to http://localhost:7860 for local dev

13. Python SDK

Installation

from object_intelligence import ObjectDetector
import cv2

Basic Detection

detector = ObjectDetector()

# Process a frame
frame = cv2.imread("sample.jpg")
annotated_frame, detections = detector.detect_and_draw(frame)

for det in detections:
    print(f"{det['label']}: {det['confidence']:.1%} ({det['type']})")

Lock Mode

detector = ObjectDetector()
detector.set_mode("combined")
detector.lock("My Cup")

frame = cv2.imread("sample.jpg")
annotated_frame, detections = detector.detect_and_draw(frame)
# Only "My Cup" detections are returned

Teaching Objects

detector = ObjectDetector()
detector.add_object(
    name="My Cup",
    images=["cup_front.jpg", "cup_back.jpg", "cup_side.jpg"],
    category="Drinkware",
    description="Personal blue mug with handle"
)

Open Vocabulary

detector = ObjectDetector()
detector.set_mode("open_vocabulary")
detector.set_vocabulary(["red cup", "laptop", "screwdriver", "wireless mouse"])

frame = cv2.imread("sample.jpg")
annotated_frame, detections = detector.detect_and_draw(frame)

14. OpenCV Examples

Five ready-to-run OpenCV webcam scripts in examples/:

Script Description
opencv_known_objects.py RT-DETR detection of standard COCO objects
opencv_open_vocabulary.py YOLO-World text-prompted detection
opencv_specific_object.py Visual embedding recognition
opencv_lock_mode.py Target lock filtering with real-time webcam
opencv_combined.py Full multi-pipeline fusion

15. Tech Stack

Backend

Technology Version Purpose
Python 3.10+ Core language
PyTorch 2.0+ Deep learning framework
Ultralytics 8.1+ RT-DETR and YOLO-World inference
Torchvision 0.15+ ResNet-18 backbone
Gradio 4.44.1 Web UI framework
FastAPI 0.100+ REST API framework
Uvicorn 0.20+ ASGI server
SQLite β€” Embedded database
OpenCV 4.8+ Image processing
NumPy 1.22+ Numerical operations

Frontend

Technology Version Purpose
Next.js 16.3.3 React framework (App Router)
React 19.2.8 UI library
TypeScript 5.x Type-safe JavaScript
Three.js 0.185 3D WebGL graphics
GSAP 3.15 Animation library
Lucide React 1.37 Icon library

Deployment

Service Purpose
HuggingFace Spaces GPU backend (ZeroGPU T4)
Vercel Frontend CDN + static hosting
Git LFS Model weight version control

16. Project Structure

Object-Intelligence/
β”œβ”€β”€ app.py                          # Main entry: Gradio UI + FastAPI + ZeroGPU
β”œβ”€β”€ requirements.txt                # Python dependencies
β”œβ”€β”€ rtdetr-l.pt                     # RT-DETR weights (Git LFS)
β”œβ”€β”€ yolov8s-worldv2.pt              # YOLO-World weights (Git LFS)
β”œβ”€β”€ open_vocabulary_classes.txt     # Default YOLO-World prompts
β”‚
β”œβ”€β”€ backend/
β”‚   β”œβ”€β”€ config.py                   # Central configuration
β”‚   β”œβ”€β”€ database/
β”‚   β”‚   β”œβ”€β”€ models.py               # Dataclass definitions
β”‚   β”‚   └── storage.py              # SQLite manager + cosine search
β”‚   └── ml/
β”‚       β”œβ”€β”€ orchestrator.py         # 5-step pipeline orchestrator
β”‚       β”œβ”€β”€ rtdetr_detector.py      # RT-DETR detector
β”‚       β”œβ”€β”€ yolo_world_detector.py  # YOLO-World detector
β”‚       β”œβ”€β”€ embedding_recognizer.py # ResNet-18 embeddings
β”‚       β”œβ”€β”€ fusion.py               # IoU NMS + priority fusion
β”‚       └── lock_mode.py            # Target filtering + annotations
β”‚
β”œβ”€β”€ object_intelligence/
β”‚   └── detector.py                 # Python SDK client
β”‚
β”œβ”€β”€ examples/                       # OpenCV webcam scripts (5 files)
β”‚
β”œβ”€β”€ docs/                           # Documentation (8 guides)
β”‚
β”œβ”€β”€ database/
β”‚   β”œβ”€β”€ objects.db                  # SQLite database (auto-created)
β”‚   └── object_images/              # Reference image storage
β”‚
└── frontend/                       # Next.js 16 frontend
    β”œβ”€β”€ app/
    β”‚   β”œβ”€β”€ page.tsx                # Home page
    β”‚   β”œβ”€β”€ live/page.tsx           # Live Testing + AR Camera
    β”‚   β”œβ”€β”€ objects/page.tsx        # Object Studio
    β”‚   β”œβ”€β”€ downloads/page.tsx      # Downloads & Resources
    β”‚   └── guides/page.tsx         # Developer Guides
    β”œβ”€β”€ components/
    β”‚   β”œβ”€β”€ Navbar.tsx              # Sticky navigation
    β”‚   β”œβ”€β”€ Footer.tsx              # Footer with tech badges
    β”‚   β”œβ”€β”€ MobileMenu.tsx          # Animated mobile menu
    β”‚   └── ParallaxScene.tsx       # Three.js 3D background
    └── lib/
        └── api.ts                  # Backend API URL construction

17. Challenges & Solutions

Challenge 1: ZeroGPU CUDA Initialization

Problem: Models loaded at module import time triggered CUDA operations outside ZeroGPU's GPU allocation context.

Solution: Implemented lazy loading via get_orchestrator() singleton pattern. Models are only instantiated when a @spaces.GPU decorated function is first called, ensuring all CUDA operations happen within ZeroGPU's allocation context.

Challenge 2: FastAPI Route Conflict with Gradio

Problem: Gradio registers a catch-all route POST /api/{api_name} that intercepts all /api/* requests, preventing custom FastAPI routes from being reached.

Solution: Used Starlette BaseHTTPMiddleware (ASGI middleware) that intercepts requests at the ASGI level, BEFORE FastAPI route matching. This allows custom API handlers to process requests before Gradio's catch-all can intercept them.

Challenge 3: Frontend-Backend CORS

Problem: Cross-origin requests from Vercel frontend to HuggingFace backend require proper CORS handling.

Solution: Added CORS middleware with allow_origins=["*"] to the FastAPI app, enabling cross-origin requests from any domain.

Challenge 4: Model Weight Persistence

Problem: Model weights (~130MB total) need to persist across HuggingFace Space restarts.

Solution: Used Git LFS for version control and Ultralytics auto-download on first inference. The huggingface_hub<1.0 pin ensures Gradio 4.x compatibility.

Challenge 5: Real-Time AR Performance

Problem: Continuous webcam frame processing needs to maintain acceptable FPS while sending frames to the backend.

Solution: Implemented requestAnimationFrame loop with isProcessingRef guard to prevent frame pile-up. JPEG compression at 70% quality reduces upload size. Canvas overlay rendering runs independently of backend calls.


18. Future Work

Short Term

  • Add batch detection for multiple images
  • Implement embedding clustering for automatic object categorization
  • Add WebRTC for lower-latency webcam streaming
  • Support video file upload and frame-by-frame analysis

Medium Term

  • PostgreSQL + pgvector migration for production-scale embedding search
  • User authentication and multi-tenant object libraries
  • Mobile app (React Native) with on-device inference
  • Model fine-tuning pipeline for domain-specific objects

Long Term

  • Edge deployment (TensorRT, ONNX Runtime) for offline use
  • 3D object pose estimation integration
  • Multi-camera tracking and re-identification
  • Natural language scene description generation

19. License

Distributed under the MIT License. See LICENSE for details.

MIT License

Copyright (c) 2026 Muhammad

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

Object Intelligence Platform
Built with PyTorch, Ultralytics, Gradio, FastAPI, Next.js, and Three.js
Live Demo Β· Backend Β· GitHub