Object-Intelligence-Backend / docs /specific-objects.md
muhammadpriv001's picture
Frontend 1.0.0
4346a4c
|
Raw History Blame Contribute Delete
1.41 kB
# Specific Object Recognition (Visual Embeddings)
## Overview
Pipeline 3 answers the crucial question: *"Is this the exact physical object I taught the platform previously?"* (e.g., distinguishing **"My Cup"** from generic *cups*).
```
Object Reference Crop ──► Vision Embedding Backbone ──► 512-dim L2 Vector
β”‚
Query Frame Crop ──────► Vision Embedding Backbone ──► Cosine Similarity (>= 0.65) ──► Match "My Cup"
```
## How Teaching Works
1. **Name & Metadata**: User inputs unique object name (e.g. *My Coffee Mug*).
2. **Reference Photos**: User uploads 3-10 reference photos from various angles and lighting.
3. **Feature Extraction**: Deep PyTorch vision backbone extracts normalized vector embeddings.
4. **Vector Persistence**: Embeddings stored in SQLite object database (`objects.db`).
5. **Real-time Recognition**: Detected bounding box crops are continuously compared against stored embeddings.
## Code Example
```python
from object_intelligence import ObjectDetector
detector = ObjectDetector()
# Teach new object
detector.add_object(
name="My Blue Mug",
images=["mug_side.jpg", "mug_front.jpg"],
category="Drinkware",
description="Personal ceramic mug with handle"
)
# Run specific object recognition
detector.set_mode("specific")
detections = detector.detect(frame)
```