Object-Intelligence-Backend / docs /specific-objects.md
muhammadpriv001's picture
Frontend 1.0.0
4346a4c
|
Raw History Blame Contribute Delete
1.41 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade

Specific Object Recognition (Visual Embeddings)

Overview

Pipeline 3 answers the crucial question: "Is this the exact physical object I taught the platform previously?" (e.g., distinguishing "My Cup" from generic cups).

Object Reference Crop ──► Vision Embedding Backbone ──► 512-dim L2 Vector
                                                               β”‚
Query Frame Crop ──────► Vision Embedding Backbone ──► Cosine Similarity (>= 0.65) ──► Match "My Cup"

How Teaching Works

  1. Name & Metadata: User inputs unique object name (e.g. My Coffee Mug).
  2. Reference Photos: User uploads 3-10 reference photos from various angles and lighting.
  3. Feature Extraction: Deep PyTorch vision backbone extracts normalized vector embeddings.
  4. Vector Persistence: Embeddings stored in SQLite object database (objects.db).
  5. Real-time Recognition: Detected bounding box crops are continuously compared against stored embeddings.

Code Example

from object_intelligence import ObjectDetector

detector = ObjectDetector()

# Teach new object
detector.add_object(
    name="My Blue Mug",
    images=["mug_side.jpg", "mug_front.jpg"],
    category="Drinkware",
    description="Personal ceramic mug with handle"
)

# Run specific object recognition
detector.set_mode("specific")
detections = detector.detect(frame)