Chris Leo commited on
scorevision: push artifact
Browse files- ANALYSIS.md +120 -0
ANALYSIS.md
ADDED
|
@@ -0,0 +1,120 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Detect-crime β King's Model Analysis
|
| 2 |
+
|
| 3 |
+
Element: `manak0/Detect-crime` (subnet 423, public/open-source track).
|
| 4 |
+
|
| 5 |
+
Current king (per Manako dashboard, 2026-05-04): hotkey `5CSeBYpYMriXUPL5zHNrprFparFdQsyKPk4S8dPxmiCZtv9f`, score **0.576**, lifetime $3,730.83.
|
| 6 |
+
|
| 7 |
+
**Correction (after deeper inspection of the manako API):** the king is **NOT** running the `manak0/Detect-crime` baseline despite no `Detect-crime-winner` repo existing on Manako's HF. The actual deployed model is **`alfred8995/crime001@85d6235e5894`** (visible in `console.scorevision.io/api/v2/elements/...` β `challengeDetails[].miners[]`). Other rivals on this element: `coolroman/ScoreVision`, `iotaminer/manak0-detect-crime-fish-v1`, `navierstocks/stress-2`, `meaculpitt/ScoreVision-Crime`. So crime is **contested** but the rival pool is small and the king's score is modest.
|
| 8 |
+
|
| 9 |
+
The original "uncontested baseline" reading was wrong β the no-`-winner`-repo signal is unreliable for crime because Manako appears to publish winner repos selectively (cf. petrol/Person/Vehicle have them, crime/road-signs/fire don't).
|
| 10 |
+
|
| 11 |
+
## 1. Element & scoring
|
| 12 |
+
|
| 13 |
+
- 6 classes (target order from `class_names.txt`): `balaclava, bat, glove, graffiti, hoodie, spray paint`.
|
| 14 |
+
- Element source: `element_trainer/crime` per the model card frontmatter.
|
| 15 |
+
- Pillar wiring (subnet code, [scorevision/vlm_pipeline/non_vlm_scoring/objects.py](../../scorevision/vlm_pipeline/non_vlm_scoring/objects.py)):
|
| 16 |
+
- `IOU` pillar β **label-agnostic** placement: per-frame AUC-F1 over IoU thresholds `(0.3, 0.5)` via Hungarian matching.
|
| 17 |
+
- `MAP50`, `PRECISION`, `RECALL`, `FALSE_POSITIVE`, `COUNT` β all **label-strict** (class names must match GT labels).
|
| 18 |
+
- The dashboard score (0.576) tracks the synthetic-fixed benchmark's `overall_iou` (0.597) almost exactly. Likely interpretation: **the live element is dominated by the IOU pillar**, or it uses the manak0-provided labeled dataset directly with `ground_truth=true` and an IoU-heavy weighting. (Cannot fully confirm without `.env` and a live manifest read.)
|
| 19 |
+
- The synthetic benchmark distribution gives us the per-class headroom map (next section).
|
| 20 |
+
|
| 21 |
+
## 2. King's actual model (`alfred8995/crime001@85d6235e5894`)
|
| 22 |
+
|
| 23 |
+
```
|
| 24 |
+
weights.onnx 19,409,670 bytes YOLOv11s, 390 ONNX nodes
|
| 25 |
+
input [1, 3, 1280, 1280] (letterboxed, RGB, /255)
|
| 26 |
+
output [1, 300, 6] NMS-baked: [x1,y1,x2,y2,conf,cls_id]
|
| 27 |
+
producer pytorch 2.11.0 ultralytics export
|
| 28 |
+
```
|
| 29 |
+
|
| 30 |
+
- **Architecture**: ultralytics YOLOv11s. ~9M params, 19 MB ONNX. Bigger than the manak0/Detect-crime template (which is YOLOv11-nano @ 640).
|
| 31 |
+
- **NMS baked in**: `[1, 300, 6]` final-detection layout β fast, simple decode path.
|
| 32 |
+
- **Class order is different from the baseline**: `[balaclava, hoodie, glove, bat, spray paint, graffiti]`. Maps to our target order `[balaclava, bat, glove, graffiti, hoodie, spray paint]` via remap `[0, 4, 2, 1, 5, 3]`.
|
| 33 |
+
- **Inference recipe** ([king_models/alfred8995_crime001/miner.py](../../king_models/alfred8995_crime001/miner.py)):
|
| 34 |
+
|
| 35 |
+
| Knob | Value | Comment |
|
| 36 |
+
|---|---|---|
|
| 37 |
+
| input size | **1280** | letterboxed, INTER_CUBIC for upscales |
|
| 38 |
+
| conf threshold | **0.52** | high β tuned for FALSE_POSITIVE pillar |
|
| 39 |
+
| NMS IoU | 0.40 | hard NMS |
|
| 40 |
+
| max_det | 150 | per image |
|
| 41 |
+
| TTA | hflip | 2 forward passes; uses TTA-cluster max-score boost |
|
| 42 |
+
| Box sanity filter | min_side=8, min_area=196, max_aspect=8 | drops tiny / degenerate detections (FP killer) |
|
| 43 |
+
| NMS scope | **per-class hard NMS** | already the recommended pattern |
|
| 44 |
+
| CLAHE / hist-eq | none | no luma-gated preprocessing |
|
| 45 |
+
|
| 46 |
+
This is a thoughtfully-tuned chute. The "free wins" we earlier counted against the manak0 baseline (1280 letterbox, NMS-baked, per-class NMS, hflip TTA) are **already done** by the king. Our remaining levers vs. alfred8995 are:
|
| 47 |
+
|
| 48 |
+
1. **Multi-scale TTA** β alfred only does single-scale (1280) hflip. Adding 1536 should help small-object recall.
|
| 49 |
+
2. **Per-class confidence floors** β alfred uses one global 0.52. We push the 4 catastrophic classes lower (0.05β0.10) and keep hoodie/graffiti at the global floor.
|
| 50 |
+
3. **WBF over hard NMS** β averaged box coords yield tighter localizations on borderline IoUβ₯0.5 cases.
|
| 51 |
+
4. **CLAHE on dark frames** β alfred has no preprocessing. CCTV crime footage is night-heavy.
|
| 52 |
+
5. **Better training data** β per-class data quality is the biggest open lever (alfred trained on whatever they had; we can do better with distillation + Roboflow + HF datasets).
|
| 53 |
+
|
| 54 |
+
### Manako synthetic benchmark vs. live element
|
| 55 |
+
|
| 56 |
+
The synthetic-fixed benchmark on `manak0/Detect-crime` (overall_iou 0.597) is run against the **manak0 baseline**, not alfred8995's deployed model. That's why the per-class numbers (balaclava recall 0.034, glove 0.064, etc.) look so bad β they reflect the template, not the king. **The live element score (0.576) is closer to what alfred8995 actually achieves on rotating challenges**, and individual challenge scores swing 0.16β1.00.
|
| 57 |
+
|
| 58 |
+
Two important corollaries:
|
| 59 |
+
|
| 60 |
+
- The on-subnet metric is almost certainly **label-agnostic placement-dominated** (the IOU pillar from `compare_object_placement` in [scorevision/vlm_pipeline/non_vlm_scoring/objects.py](../../scorevision/vlm_pipeline/non_vlm_scoring/objects.py) which uses `label_strict=False`). Evidence: alfred8995's class order differs from the baseline's, but their score still tracks the baseline IoU benchmark β only label-agnostic placement matching produces this behavior.
|
| 61 |
+
- **Class IDs barely matter** as long as the boxes are well-placed. Our miner can use any consistent ordering. Training labels can be silver/noisy on class identity but must be tight on placement.
|
| 62 |
+
|
| 63 |
+
### Per-class baseline (from `benchmark/synthetic/latest.json`, 50 imgs / 611 GT / 326 preds):
|
| 64 |
+
|
| 65 |
+
| metric | overall | balaclava | bat | glove | graffiti | hoodie | spray paint |
|
| 66 |
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 67 |
+
| IoU | **0.597** | 0.029 | 0.037 | 0.137 | 0.232 | **0.539** | 0.070 |
|
| 68 |
+
| mAP@50 | 0.140 | 0.014 | 0.107 | 0.054 | 0.307 | 0.243 | 0.116 |
|
| 69 |
+
| mAP@50β95 | 0.083 | 0.005 | 0.063 | 0.038 | 0.175 | 0.169 | 0.045 |
|
| 70 |
+
| Precision | 0.365 | 0.182 | 0.182 | 0.189 | 0.325 | 0.500 | 0.227 |
|
| 71 |
+
| Recall | **0.195** | **0.034** | 0.143 | **0.064** | 0.321 | 0.274 | 0.161 |
|
| 72 |
+
|
| 73 |
+
Two structural facts jump out:
|
| 74 |
+
|
| 75 |
+
1. **Recall is 0.195 overall** with `pred_count=326 < gt_count=611` β the king under-predicts severely. Lowering confidence thresholds is almost certainly free score.
|
| 76 |
+
2. **`hoodie` carries the overall IoU number alone**: hoodie IoU 0.54 vs balaclava 0.03, glove 0.14, bat 0.04, spray paint 0.07. If the live metric is IoU-dominated, the king's score is essentially "how often is there a hoodie box that roughly overlaps a hoodie GT" β which is why the dashboard sits at 0.576 and not at 0.14 (the mAP). **Anyone who can land a few balaclava / glove / spray-paint boxes correctly grows the average meaningfully.**
|
| 77 |
+
|
| 78 |
+
## 3. The gap to close
|
| 79 |
+
|
| 80 |
+
King = 0.576. Per-class IoU ceiling is 1.0 each. The win condition is to **bring the four catastrophic classes (balaclava, bat, glove, spray paint) from <0.10 IoU into the 0.30β0.50 range** while preserving hoodie/graffiti.
|
| 81 |
+
|
| 82 |
+
Levers in expected impact order:
|
| 83 |
+
|
| 84 |
+
1. **Bigger input + letterbox**. King runs at 640 with stretch resize. Balaclava boxes on a 1408Γ768 frame are routinely <40 px β a 640 stretch destroys them. Move to YOLOv11s/m at 1280 with proper letterbox; this alone should lift balaclava/glove/spray-paint recall by 2β4Γ.
|
| 85 |
+
2. **Per-class confidence floors**. King's global 0.25 over-suppresses the rare classes. Set `balaclava=0.05, bat=0.10, glove=0.05, graffiti=0.20, hoodie=0.20, spray paint=0.10`. The synthetic benchmark FP/FFPI saturation is high (326 preds across 50 imgs β 6.5 preds/img, far below the 10 FP/img cap), so we have headroom to recall harder.
|
| 86 |
+
3. **Multi-scale TTA + WBF**. `{1280, 1536} Γ {orig, hflip}` merged with class-aware Weighted Box Fusion. Same petrol playbook.
|
| 87 |
+
4. **Class-aware NMS at IoU=0.45**. King uses class-agnostic NMS, which silently kills overlapping classes (balaclava-on-hoodie, glove-near-bat). Switch to class-aware to keep both.
|
| 88 |
+
5. **CLAHE on dark frames only**. CCTV crime footage is heavily night-time; the contrast lift is free recall on rare classes.
|
| 89 |
+
6. **Train on *real* data**. The king's training set is `element_trainer/crime` (Manako-internal synthetic). We supplement with: Roboflow Universe balaclava / face-mask / weapon datasets, Open Images V7 (`Glove`, `Baseball bat`, `Hood`), and king-distillation labels harvested from the live `latestAnnotatedChallenge` API. Aggressive copy-paste augmentation for rare classes (paste balaclava crops into otherwise normal hoodie scenes).
|
| 90 |
+
|
| 91 |
+
The training data lever is the biggest because the king's per-class numbers expose a tiny / heavily-imbalanced training set. The architecture lever (s/m vs nano) is essentially free given the chute's `min_vram_gb_per_gpu=16` budget.
|
| 92 |
+
|
| 93 |
+
## 4. Constraints from the runtime
|
| 94 |
+
|
| 95 |
+
- Chute sandbox: only stdlib + pip-installed packages from `chute_config.yml`. **No extra `.py` imports from the HF repo other than `miner.py`** β everything must live in `miner.py`.
|
| 96 |
+
- The king's existing `chute_config.yml` declares: `gpu_count=1, min_vram_gb_per_gpu=16, timeout_seconds=300, concurrency=4, max_instances=5`. We can keep this β YOLOv11s/m at 1280 with TTA fits comfortably under 16 GB.
|
| 97 |
+
- No external egress inside the chute β all assets must be in HF.
|
| 98 |
+
- `predict_batch(batch_images, offset, n_keypoints) -> list[TVFrameResult]` is the required entry point. Frames arrive as BGR `np.uint8` HWC arrays.
|
| 99 |
+
- `n_keypoints` is 0 for OBJECT_DETECTION elements; we still return the right number of `(0, 0)` placeholders.
|
| 100 |
+
|
| 101 |
+
## 5. The plan in this directory
|
| 102 |
+
|
| 103 |
+
```
|
| 104 |
+
scratch/crime_miner/
|
| 105 |
+
βββ ANALYSIS.md β this file
|
| 106 |
+
βββ README.md β recipe + how to run
|
| 107 |
+
βββ miner.py β deployable inference, ports the petrol pattern to 6-class crime
|
| 108 |
+
βββ chute_config.yml β matches king's resource spec
|
| 109 |
+
βββ class_names.txt β target class order β DO NOT REORDER
|
| 110 |
+
βοΏ½οΏ½β training/
|
| 111 |
+
βββ DATASET.md β data sources + pipeline (start here)
|
| 112 |
+
βββ build_dataset.py β public datasets + manako frames + king-distillation
|
| 113 |
+
βββ poll_manako.py β background poller for in-domain frames + king's preds
|
| 114 |
+
βββ train.py β two-stage YOLOv11 training (silver β clean)
|
| 115 |
+
βββ export_onnx.py β export with NMS baked in β [1, 300, 6]
|
| 116 |
+
βββ verify_dataset.py β quick QA over the assembled YOLO dirs
|
| 117 |
+
βββ requirements.txt
|
| 118 |
+
```
|
| 119 |
+
|
| 120 |
+
What's **not** done: the actual training runs. Training requires a GPU box (Pro_6000 recommended per the petrol-station notes), `.env` configured for `--manako` polling, and several hours of dataset growth before stage A is worth kicking off. The pipeline is set up so each step is one command and idempotent.
|