|
Download README.md from bench-labs/objectmodel-v1: direct link, hf CLI and curl.
- Browser
- Download file 18.5 kB
-
https://huggingface.co/bench-labs/objectmodel-v1/resolve/main/README.md
- Command line
-
hf download hf://bench-labs/objectmodel-v1/README.md
-
curl -L -o README.md https://huggingface.co/bench-labs/objectmodel-v1/resolve/main/README.md
18.5 kB
| license: apache-2.0 | |
| tags: | |
| - object-detection | |
| - pytorch | |
| - computer-vision | |
| - research | |
| # ObjectModel-v1 | |
| <table align="center"> | |
| <tr> | |
| <td align="center"> | |
| <a href="https://huggingface.co/spaces/bench-labs/blog?post=ObjectModel-v1.html">📝 Blog Post</a> | |
| </td> | |
| </tr> | |
| </table> | |
| ObjectModel-v1 is a clean-room, compact, end-to-end object detector from Bench Labs. It is a | |
| research implementation, not a benchmark claim. The model tests | |
| whether global semantic reasoning can be compressed into a small fixed latent memory while | |
| precise geometry is recovered by query-conditioned sampling from full-resolution pyramid | |
| features. | |
| The model is NMS-free. It predicts a fixed set of objects and is trained with Hungarian | |
| bipartite matching. | |
| This is v1: a single from-scratch training run with no pretraining, no hyperparameter search | |
| across seeds, and none of the planned ablations run yet. Treat everything below as a first | |
| checkpoint in an ongoing process, not a finished result. v2 is expected to do meaningfully better, | |
| whether that comes from more training compute, an architecture change informed by v1's ablations, | |
| or both. | |
| ## Status | |
| - Architecture, COCO data path, losses, training, evaluation, profiling, and ONNX export are implemented. | |
| - Synthetic forward/loss/backward and data tests are included. | |
| - Full COCO2017 training is complete: 100 epochs on a single RTX 5090, peak val AP 0.358 at epoch 95. | |
| See [Training Progress](#training-progress) below for the full curve. | |
| - No "state of the art" or "beats YOLO/DETR" claim is made. Peak AP (0.358) is below the 40-50 | |
| range originally set as a competitiveness bar against 20-40M-parameter real-time detectors. | |
| This is a real, single-seed, un-pretrained result, not a benchmark claim. See the | |
| [Minimum Validation Protocol](#minimum-validation-protocol) for what a competitive claim actually | |
| requires: controlled baselines, three seeds, and ablations, none of which have been run yet. | |
| ## Training Progress | |
| <img src="assets/ap_progress.svg" alt="Validation AP by epoch across the full 100-epoch run, rising from near-zero to a peak of 0.358 at epoch 95" width="900"> | |
| | | | | |
| |---|---| | |
| | Parameters | 40.8M | | |
| | Epochs | 100 / 100 (complete) | | |
| | Peak val AP (IoU 0.50:0.95) | 0.358 (epoch 95) | | |
| | Final val AP (epoch 100) | 0.356 | | |
| | Val AP50 / AP75 | 0.544 / 0.381 | | |
| | Val AP small / medium / large | 0.188 / 0.387 / 0.493 | | |
| | Val AR@100 | 0.573 | | |
| | Train loss | 4.11 (from 15.77 at epoch 1) | | |
| | Throughput | ~90-110 img/s, batch 32, RTX 5090 | | |
| | Single-frame inference | 30.7 ms / 32.5 FPS (batch 1, eager, RTX 5090) | | |
| Train loss over the same run, log-shaped as expected for a detector trained from scratch: | |
| <img src="assets/loss_progress.svg" alt="Train loss by epoch, dropping sharply in the first few epochs then declining gradually from about 6 to 4.1 over the rest of training" width="900"> | |
| AP50 (IoU 0.50, a looser localization threshold) against AP75 (IoU 0.75, stricter). Both climb | |
| together early on, then AP75 plateaus lower, meaning coarse localization improved faster than | |
| precise localization did: | |
| <img src="assets/ap50_75_progress.svg" alt="AP50 and AP75 by epoch as two lines, AP50 rising to about 0.54 and AP75 to about 0.38" width="900"> | |
| AP broken out by object size. Large objects are detected far more reliably than small ones | |
| throughout training, a common pattern for this class of detector and one of the things the | |
| required ablations would need to explain: | |
| <img src="assets/ap_by_size_progress.svg" alt="AP by object size (small, medium, large) by epoch as three lines, large objects highest at about 0.49, medium at about 0.39, small lowest at about 0.18" width="900"> | |
| Peak AP came at epoch 95 (0.3578), not the final epoch. That is normal late-training fluctuation, | |
| and `best.pt` correctly holds the epoch-95 weights rather than epoch 100's. Against the 40-50 AP | |
| range set as the competitiveness bar, this run falls short. It is a real result from a genuine | |
| from-scratch run, on the low end of what the project was aiming for. | |
| The gap is exactly what the [Minimum Validation Protocol](#minimum-validation-protocol) exists to | |
| characterize properly: whether more decoder layers, more latents, longer training, or pretraining | |
| would close it is unknown without actually running those ablations. | |
| ## Examples | |
| Six detections from the final EMA checkpoint (epoch 95, peak AP) on COCO val2017 images, | |
| confidence ≥ 0.35. Picked for variety, not cherry-picked for perfection. | |
| | | | | |
| |---|---| | |
| |  |  | | |
| | Dessert tray, bottle, bowl, and a dozen individually boxed items | Skate park, parent and child, skateboard at 0.93 confidence | | |
| |  |  | | |
| | Street market, person at 0.94, market umbrella | Rodeo, crowd of people, one animal still mislabeled | | |
| |  |  | | |
| | Family at a table, cups, chair, bottle, two people | Beach, kite at 0.80, people along the shore, a distant boat | | |
| ## Live Tracking Demo | |
| <img src="assets/tracking_demo.gif" alt="ObjectModel-v1 detections chained through a SORT tracker on real pedestrian footage, boxes and ids tracking people across frames" width="420"> | |
| [Full-length video (20s, MP4)](assets/tracking_demo.mp4) | |
| ObjectModel-v1 itself has no temporal component. Every frame is detected independently. The clip | |
| above chains detections through a from-scratch SORT-style tracker (`src/objectmodel_v1/tracking.py`: | |
| constant-velocity Kalman motion model plus IoU/Hungarian frame-to-frame association) to give boxes | |
| a persistent id and a short motion trail. Source footage is `vtest.avi`, OpenCV's standard | |
| pedestrian test clip (BSD-3, ships with OpenCV), genuine video the model never trained on. | |
| This is the same clip run at three points in training, tracker unchanged throughout. Only the | |
| detector's checkpoint improved: | |
| | Checkpoint | Ids issued over 20s | Longest-lived ids | | |
| |---|---|---| | |
| | Epoch 13 (AP 0.227) | ~89 | none survive past a few seconds | | |
| | Epoch 26 (AP 0.288) | ~84 | 3 ids survive nearly the full clip | | |
| | Epoch 95 (AP 0.358, peak) | 92 | 4 ids survive nearly the full clip | | |
| Track count issued does not fall much, since new people keep entering frame throughout the clip | |
| and each one earns a new id, which is correct behavior. Track persistence for people already in | |
| frame improved consistently instead. A separate test on a genuinely different scene, an eye-level | |
| warehouse clip not shown here, also surfaced a real and distinct limitation: the detector still | |
| occasionally hallucinates objects on plain background surfaces, reading a support pillar as | |
| "refrigerator", even at this peak checkpoint. Tracking quality rides on detection quality, and | |
| detection quality on out-of-domain footage (camera angles, lighting, and compression that COCO's | |
| photos do not really cover) is visibly weaker than on COCO's own validation images. | |
| ## Restricted Zone Detection | |
| <img src="assets/zone_demo.gif" alt="A marked polygon zone over a plaza walkway, with tracked people highlighted and a red ALERT banner firing when someone enters the zone" width="420"> | |
| [Full-length video (20s, MP4)](assets/zone_demo.mp4) | |
| A small layer on top of the tracker: `src/objectmodel_v1/zones.py` defines a polygon in the same | |
| pixel coordinates as the tracked boxes, and fires an event only on the transition into or out of | |
| it, not on every frame a track spends inside. It keys off each box's bottom-center point (roughly | |
| where feet touch the ground) rather than the box centroid, since a zone drawn on a floor plane | |
| should care where someone is standing, not where their torso is. | |
| Correctness is checked on a synthetic sequence, not just eyeballed on video: a track walking | |
| through a rectangle produces exactly one `entered` and one `exited` event, none of the frames | |
| spent inside re-fire, a track that never enters never fires, and a fresh id gets independent | |
| state. That test lives with the module. | |
| On the same plaza clip used above, the zone above the walkway fired 91 enter/exit events over 20 | |
| seconds. Read that number carefully: most of it is not 91 different real crossings. It is the | |
| same tracker id churn already described in the section above, the same person's track resetting | |
| and re-entering the zone as a "new" id, plus a handful of misclassified objects (a backpack or | |
| handbag read as its own tracked object) that shouldn't have triggered an event at all. The event | |
| mechanics are verified correct; the real-world event count is exactly as reliable as the | |
| underlying detector and tracker are, which right now is "usable for a demo, not for anything | |
| where a false alert has a real cost." | |
| ```python | |
| from objectmodel_v1.zones import RestrictedZoneMonitor | |
| zone = [(300, 150), (650, 150), (650, 320), (300, 320)] # pixel-space polygon | |
| monitor = RestrictedZoneMonitor(zone) | |
| for frame_index, frame in enumerate(video_frames): | |
| boxes, labels, scores = detect(frame) | |
| tracks = tracker.update(boxes, labels, scores) | |
| for event in monitor.update(tracks, frame_index): | |
| print(event.track_id, event.kind, event.position) # "entered" or "exited" | |
| ``` | |
| ## Architecture | |
| ```text | |
| image | |
| -> compact convolutional backbone (strides 8/16/32) | |
| -> top-down pyramid fusion | |
| -> pooled multi-scale tokens | |
| -> fixed latent memory (global semantics) | |
| -> learned object queries | |
| -> query self-attention | |
| -> cross-attention to latent memory | |
| -> local sampling around the current query box | |
| -> iterative class and box prediction | |
| -> object set (no anchors, no NMS) | |
| ``` | |
| The local sampling radius scales with each query's current width and height. Early decoder | |
| layers can search broadly; later layers focus naturally as boxes are refined. During training, | |
| an optional dense auxiliary head adds one-to-many spatial supervision. It is discarded for | |
| inference and must be evaluated as an ablation, not assumed to help. | |
| ## Installation | |
| Use Python 3.11 or another PyTorch-supported Python version: | |
| ```bash | |
| python3.11 -m venv .venv | |
| .venv/bin/pip install --upgrade pip | |
| .venv/bin/pip install -e '.[coco,export,dev]' | |
| ``` | |
| For a specific CUDA build, install the matching PyTorch wheel first using the command from | |
| <https://pytorch.org/get-started/locally/>, then install ObjectModel-v1. | |
| ## Data | |
| The default configuration expects COCO 2017: | |
| ```text | |
| /path/to/coco/ | |
| annotations/instances_train2017.json | |
| annotations/instances_val2017.json | |
| train2017/*.jpg | |
| val2017/*.jpg | |
| ``` | |
| Category IDs are mapped to contiguous training labels and converted back during evaluation. | |
| Images without target objects are supported. | |
| ## Commands | |
| Profile the model before allocating training compute: | |
| ```bash | |
| objectmodel-profile --config configs/objectmodel_v1.yaml --device cuda | |
| ``` | |
| Overfit a small dataset first. A full single-GPU command is: | |
| ```bash | |
| objectmodel-train \ | |
| --config configs/objectmodel_v1.yaml \ | |
| --data-root /path/to/coco \ | |
| --output outputs/objectmodel_v1 | |
| ``` | |
| Distributed training: | |
| ```bash | |
| torchrun --standalone --nproc_per_node=8 -m objectmodel_v1.train \ | |
| --config configs/objectmodel_v1.yaml \ | |
| --data-root /path/to/coco \ | |
| --output outputs/objectmodel_v1 | |
| ``` | |
| Resume and override configuration values: | |
| ```bash | |
| objectmodel-train \ | |
| --config outputs/objectmodel_v1/config.yaml \ | |
| --data-root /path/to/coco \ | |
| --output outputs/objectmodel_v1 \ | |
| --resume outputs/objectmodel_v1/last.pt \ | |
| --set train.batch_size=8 | |
| ``` | |
| Evaluate the EMA checkpoint with canonical `pycocotools` metrics: | |
| ```bash | |
| objectmodel-eval \ | |
| --config outputs/objectmodel_v1/config.yaml \ | |
| --checkpoint outputs/objectmodel_v1/best.pt \ | |
| --data-root /path/to/coco | |
| ``` | |
| Export raw logits and normalized `cxcywh` boxes to ONNX: | |
| ```bash | |
| objectmodel-export \ | |
| --config outputs/objectmodel_v1/config.yaml \ | |
| --checkpoint outputs/objectmodel_v1/best.pt \ | |
| --output outputs/objectmodel_v1/objectmodel-v1.onnx | |
| ``` | |
| Track detections across video frames (see [Live Tracking Demo](#live-tracking-demo); this is a | |
| post-processing layer over independent per-frame detections, not a model capability): | |
| ```python | |
| from objectmodel_v1.tracking import SortTracker | |
| tracker = SortTracker(iou_threshold=0.3, max_age=5, min_hits=2) | |
| for frame in video_frames: | |
| boxes, labels, scores = detect(frame) # your decode_predictions() call | |
| for t in tracker.update(boxes, labels, scores): | |
| print(t.id, t.box, t.label, t.score) | |
| ``` | |
| ## Minimum Validation Protocol | |
| Before describing ObjectModel-v1 as competitive, run all models on the same COCO train2017 and | |
| val2017 data, image resolution, augmentation budget, training epochs, and hardware. Report: | |
| - COCO AP, AP50, AP75, APS, APM, and APL. | |
| - Parameters, FLOPs/MACs, FP32/FP16/INT8 artifact sizes. | |
| - End-to-end batch-1 median and p95 latency, including preprocessing and decoding. | |
| - Peak training and inference memory, GPU-hours, epochs, and images seen. | |
| - Three seeds for the principal result, with mean and standard deviation. | |
| - Results both from random initialization and with the same permitted pretraining. | |
| Required ablations: | |
| | Experiment | Question | | |
| |---|---| | |
| | latent memory vs flattened feature attention | Does compression preserve useful global context? | | |
| | local sampler disabled | Does high-resolution geometric evidence improve localization? | | |
| | fixed vs box-scaled offsets | Does coarse-to-fine sampling matter? | | |
| | dense auxiliary head disabled | Does added supervision improve convergence? | | |
| | 1/2/3 latent layers | Where is the accuracy/latency optimum? | | |
| | 32/64/96 latents | How aggressively can global context be compressed? | | |
| | 3/4/6 decoder layers | What is the anytime speed/accuracy curve? | | |
| Suggested external baselines are RT-DETR-R18, D-FINE-N/S, LW-DETR-T/S, and YOLOX-S. Use | |
| their official implementations and report their license and measurement setup separately. | |
| ## Possible Extensions and v2 Directions | |
| Two different questions worth separating: what this pipeline could be extended to do without | |
| touching the detector, and what would actually make the detector itself better in v2. | |
| **Extensions on top of the current detector and tracker.** The zone monitor above is one instance | |
| of a general pattern: detector output plus `SortTracker`'s persistent ids is enough to build most | |
| counting and monitoring logic without retraining anything. | |
| - Unique counting: tracks already carry stable ids, so counting distinct people or objects through | |
| a scene (footfall, event attendance, traffic counts) needs no new detection work, just tallying | |
| ids instead of per-frame boxes. | |
| - Movement heatmaps: the demos already compute per-track position trails for the overlays. The same | |
| data aggregated across a full video gives a map of where traffic actually concentrates. | |
| - Dwell time and loitering: track how long a given id stays inside a zone instead of only its | |
| entry and exit. A per-id timer, a small addition to `zones.py`'s pattern. | |
| - Line crossing: detect a track crossing a line in a given direction instead of entering a polygon. | |
| A doorway in/out counter. Same shape as `RestrictedZoneMonitor`, a line-side test instead of | |
| polygon containment. | |
| - Class-based counting: the detector already classifies 80 COCO categories, so counting vehicles, | |
| bicycles, or any other class is the same unique-counting idea applied to a different label. | |
| Behavior or anomaly detection (wrong-way movement, erratic motion) and multi-camera re-identification | |
| are real extensions but not close additions. Both need work this repository does not do yet: | |
| proper motion modeling beyond a Kalman filter, and cross-camera identity matching. | |
| **What v2 would actually need to change.** Three things came out of this run as specific, not | |
| generic: | |
| - Small objects are detected far worse than large ones (APS 0.188 vs APL 0.493, see | |
| [Training Progress](#training-progress)). That gap is bigger than the usual small-vs-large | |
| spread in this class of detector, and points at the fixed latent memory or the local sampler's | |
| pyramid resolution as the first places to look, not just "train longer." | |
| The latent-layer and latent-count ablations above exist specifically to test that. | |
| - Detection quality drops visibly on footage that does not look like COCO's own photography | |
| (see [Live Tracking Demo](#live-tracking-demo)), which limits every downstream extension listed | |
| above. Broader training data or augmentation is the direct lever; the required ablations don't | |
| cover this gap, which is worth adding to that list rather than assuming architecture changes | |
| alone fix it. | |
| - v1 trained from random initialization only. The Minimum Validation Protocol already calls for a | |
| pretrained comparison, and it is the single change most likely to move AP the most, based on how | |
| much of a difference pretraining typically makes at this parameter count in the baselines cited | |
| under Research Basis. | |
| None of this is scheduled work, it's the concrete list of what the data from this run actually | |
| points at, as opposed to a generic "train bigger, train longer." | |
| ## Research Basis | |
| ObjectModel-v1 builds on published, independently attributable ideas: | |
| - DETR: set prediction and Hungarian matching. | |
| - Conditional and Deformable DETR: spatially conditioned/local sparse attention. | |
| - RT-DETR: efficient separation of multi-scale encoding and query decoding. | |
| - D-FINE: evidence that fine-grained iterative localization is valuable. | |
| - DEIM: evidence that one-to-one matching benefits from denser training supervision. | |
| - LW-DETR: evidence that compact transformer detectors can compete with real-time CNNs. | |
| ObjectModel-v1's specific hypothesis is the combination of a **fixed compressed global memory** and | |
| **box-scaled local pyramid sampling**. Publication novelty requires a broader prior-art search | |
| and empirical ablations; this repository does not claim that the combination is patent-new. | |
| ## License | |
| Apache License 2.0. Dataset images, annotations, pretrained weights, and external baselines | |
| retain their own licenses and are not included. | |