Object-Intelligence-Backend / docs /getting-started.md
muhammadpriv001's picture
Frontend 1.0.0
4346a4c
|
Raw History Blame Contribute Delete
1.64 kB
# Getting Started with Object Intelligence Platform
Welcome to the Open-World Object Intelligence Platform. This platform unifies three distinct computer vision paradigms into a single modular architecture:
1. **Known Object Detection (RT-DETR)**: Ultra-fast, predictable detection for standard trained classes (COCO).
2. **Open-Vocabulary Discovery (YOLO-World)**: Zero-shot visual concept detection driven by text prompts.
3. **Specific Object Recognition (Visual Embeddings)**: Teachable feature embeddings to recognize individual physical objects (e.g. *My Cup*).
---
## πŸš€ Quick Setup
### 1. Installation
Ensure Python 3.9+ is installed, then run:
```bash
pip install -r requirements.txt
```
### 2. Launching the Gradio Web Application & API
To start the Gradio interface locally or prepare for Hugging Face Spaces deployment:
```bash
python app.py
```
Open your browser at `http://localhost:7860`.
---
## πŸ“¦ Python SDK Usage
```python
from object_intelligence import ObjectDetector
import cv2
# Initialize unified detector
detector = ObjectDetector()
# Select mode: 'combined', 'known', 'open_vocabulary', 'specific'
detector.set_mode("combined")
# Enable Lock Mode if desired
detector.lock("My Cup")
# Read frame and run detection
frame = cv2.imread("test.jpg")
annotated_frame, detections = detector.detect_and_draw(frame)
# Save result
cv2.imwrite("output.jpg", annotated_frame)
```
---
## πŸ”’ Lock Mode
Lock Mode filters all candidate detections across all pipelines, rendering **ONLY** bounding boxes that match your specified target string. All non-matching detections are suppressed before drawing.