File size: 1,636 Bytes
4346a4c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
# Getting Started with Object Intelligence Platform

Welcome to the Open-World Object Intelligence Platform. This platform unifies three distinct computer vision paradigms into a single modular architecture:

1. **Known Object Detection (RT-DETR)**: Ultra-fast, predictable detection for standard trained classes (COCO).
2. **Open-Vocabulary Discovery (YOLO-World)**: Zero-shot visual concept detection driven by text prompts.
3. **Specific Object Recognition (Visual Embeddings)**: Teachable feature embeddings to recognize individual physical objects (e.g. *My Cup*).

---

## ๐Ÿš€ Quick Setup

### 1. Installation
Ensure Python 3.9+ is installed, then run:

```bash
pip install -r requirements.txt
```

### 2. Launching the Gradio Web Application & API
To start the Gradio interface locally or prepare for Hugging Face Spaces deployment:

```bash
python app.py
```

Open your browser at `http://localhost:7860`.

---

## ๐Ÿ“ฆ Python SDK Usage

```python
from object_intelligence import ObjectDetector
import cv2

# Initialize unified detector
detector = ObjectDetector()

# Select mode: 'combined', 'known', 'open_vocabulary', 'specific'
detector.set_mode("combined")

# Enable Lock Mode if desired
detector.lock("My Cup")

# Read frame and run detection
frame = cv2.imread("test.jpg")
annotated_frame, detections = detector.detect_and_draw(frame)

# Save result
cv2.imwrite("output.jpg", annotated_frame)
```

---

## ๐Ÿ”’ Lock Mode
Lock Mode filters all candidate detections across all pipelines, rendering **ONLY** bounding boxes that match your specified target string. All non-matching detections are suppressed before drawing.