File size: 4,476 Bytes
534cdd5
 
 
 
 
 
 
 
 
 
 
 
 
6d0e7a4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
534cdd5
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
---
title: Object Counter
emoji: πŸ”Ž
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: "5.9.1"
app_file: app.py
pinned: false
license: mit
short_description: Count any object in an image by typing what to look for
---

# Object Counter

Count any object in an image β€” people, cars, sticks, steel rods, or anything else you can name β€” using open-vocabulary object detection. Type what you want counted at runtime; no fixed class list, no training required.

## How it works

This project uses **[YOLO-World](https://github.com/AILab-CVC/YOLO-World)**, an open-vocabulary detector. Instead of being limited to a fixed set of categories (like standard YOLO's 80 COCO classes), you give it a comma-separated list of object names at runtime, and it detects and counts each one in your image:

```
car, person, steel rod, stick
```

The model draws a bounding box around every instance it finds, and the app tallies the boxes per class.

## Repo structure

| File | What it is |
|---|---|
| [`app.py`](./app.py) | Standalone Python script β€” Gradio app with a text box for object classes. Run locally. |
| [`requirements.txt`](./requirements.txt) | Dependencies for `app.py`. |
| [`object_counter_gradio_app.ipynb`](./object_counter_gradio_app.ipynb) | Same Gradio app, packaged as a Colab notebook (free GPU, no local setup). **Recommended starting point.** |
| [`object_counter_yolo_world.ipynb`](./object_counter_yolo_world.ipynb) | Earlier, simpler version with a hardcoded class list β€” kept as a minimal reference example. |

## Getting started

### Option A β€” Colab (no local setup, free GPU)

1. Open [`object_counter_gradio_app.ipynb`](./object_counter_gradio_app.ipynb) in [Google Colab](https://colab.research.google.com/)
2. Enable a GPU runtime: `Runtime β†’ Change runtime type β†’ T4 GPU`
3. Run all cells (Runtime β†’ Run all)
4. A Gradio link/interface appears at the bottom β€” upload an image, type the object(s) you want counted (comma-separated for multiple), and hit Submit

### Option B β€” Run locally

```bash
pip install -r requirements.txt
python app.py
```

Open the local URL Gradio prints (usually `http://127.0.0.1:7860`). Works without a GPU, just slower per image.

## Usage

- Type a single object (`car`) or several (`car, person, bicycle, dog`) β€” no fixed list.
- Be as descriptive as helps: `"steel rod"` detects better than `"metal"`; `"delivery van"` is more specific than `"vehicle"`.
- Adjust the confidence slider if results are off β€” lower catches more objects (with more false positives), higher is stricter.

## Tuning results

- **Missing objects?** Lower the confidence threshold (try 0.05–0.1), or use more descriptive class names.
- **Too many false positives?** Raise the threshold, or make class names more specific.
- **Tightly packed or overlapping objects undercounted** (e.g. a bundle of steel rods)? This is a known limitation of box-based detectors β€” separating touching/overlapping instances into individual boxes is hard. A density-map counting approach (e.g. CountGD), built specifically for packed/repetitive objects, is a planned follow-on for this case.

## Known issue: CUDA/CPU device mismatch

On GPU runtimes, `model.set_classes()` can throw:

```
RuntimeError: Expected all tensors to be on the same device, but got index is on cpu, different from other tensors on cuda:0
```

This is a device-handling bug in how YOLO-World's CLIP text encoder gets loaded, not a config error on your end. Both `app.py` and the Gradio notebook already include the fix: the model is moved to CPU before `set_classes()` and back to the target device (GPU or CPU) afterward. If you still hit this after pulling the latest version here, restart the runtime/kernel and re-run from the model-loading cell.

## Roadmap

- [x] Static image counting (YOLO-World)
- [x] User-defined classes at runtime (no hardcoded list)
- [x] Gradio interface for interactive use
- [x] Standalone Python script version
- [ ] Live webcam / video counting
- [ ] Density-map counting mode for tightly packed objects

## Requirements

Installed automatically via `requirements.txt` / the notebook's install cell:
- `ultralytics`
- `supervision`
- `gradio`
- `opencv-python-headless`
- `pillow`

## License

MIT β€” set in the Spaces config block above. Add a `LICENSE` file with the full MIT text if you also want it to show up as the repo's license on GitHub. Change the `license:` field in that block if you'd prefer something else.