File size: 5,118 Bytes
e8b8483 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 | ---
license: other
license_name: fair-research-license
license_link: https://huggingface.co/facebook/EUPE-ViT-B/blob/main/LICENSE
base_model: facebook/EUPE-ViT-B
tags:
- image-classification
- binary-classification
- minimal-models
- interpretability
- vision-transformer
- circuit-synthesis
library_name: pytorch
datasets:
- detection-datasets/coco
pipeline_tag: image-classification
---
# Zero-Parameter Classifier
Image-level person classification on EUPE-ViT-B features. A 768 pixel image
gives 2304 patch tokens at the final layer; layernorm across the 768 channels
and max-pool across patches gives one 768-D vector. The decision compares two
sums of that vector. The boundary is zero.
```python
patches = backbone(image)["x_norm_patchtokens"] # (2304, 768)
pooled = layernorm(patches, 768).max(dim=0) # (768,)
present = pooled[pos_dims].sum() > pooled[neg_dims].sum()
```
At two dimensions:
```
person present ⟺ feat[48] > feat[637]
```
```python
from infer import PersonDetector
det = PersonDetector.load('d6')
present = det.predict('image.jpg')
```
## Rules
Dimensions are selected on COCO train2017, 118,287 images, and scored on
val2017, 5000 images. The splits are disjoint.
| rule | dims | F1 | precision | recall | slices | LUT4 | CCU2C | ns |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| `d2` | 2 | 0.8681 | 0.8685 | 0.8678 | 4 | 7 | 0 | 0.40 |
| `d4` | 4 | 0.8817 | 0.9069 | 0.8578 | 10 | 8 | 10 | 0.90 |
| `d6` | 6 | 0.8977 | 0.9461 | 0.8541 | 20 | 9 | 20 | 1.40 |
| `d8` | 8 | 0.9039 | 0.9475 | 0.8641 | 30 | 9 | 30 | 1.90 |
| `d12` | 12 | 0.9068 | 0.9500 | 0.8674 | 60 | 10 | 60 | 2.90 |
| `d16` | 16 | 0.9126 | 0.9568 | 0.8723 | 84 | 10 | 84 | 3.90 |
| `d20` | 20 | 0.9217 | 0.9587 | 0.8875 | 108 | 11 | 108 | 4.90 |
| `d40` | 40 | 0.9307 | 0.9698 | 0.8945 | 266 | 12 | 266 | 9.90 |
Train and validation F1 differ by 0.0006 at 40 dimensions and by at most 0.0078
across the set. `d6` is the default.
## Dimensions
```
d2 48 > 637
d4 48 + 71 > 637 + 90
d6 48 + 71 + 292 > 637 + 90 + 82
```
Selection is greedy over the 192 dimensions with the largest class-mean
separation, alternating sides and adding whichever remaining dimension most
improves F1 at a zero boundary. Rules nest; `tests/test_rules.py` checks the
nesting.
Dimension 48 responds to people and to person-associated objects and is
suppressed on non-human animals and on non-anthropogenic structures.
## Offset
A decision of the form `sum(pos) - sum(neg) > t` requires `t` because the two
sums carry a relative offset. Sets selected under a zero boundary carry none. A
40-dimension set selected at `t = 25.28` scores F1 0.7410 on these images when
`t` is set to zero.
Dimension indices and signs are fixed structure. Each rule has no free
parameters and 2 to 40 fixed ones.
## Circuit
`rtl_gen.py` emits one Verilog module per rule. `synth.py` synthesizes them with
[nosis](https://github.com/CharlesCNorton/nosis) for a Lattice ECP5 LFE5U-25F.
Counts are LUT4s, carry cells and slices on that device. Inputs are the selected
channels as signed INT8, post-layernorm and post-max-pool. Output is one bit,
combinational, with no multipliers, no memory and no constants.
`d2` contains no adder and is LUT-bound at 0.40 ns. Wider rules are carry-bound,
with area and delay linear in dimension count.
`tests/test_rtl.py` simulates each module against a Python reference under
Icarus Verilog, on uniform inputs and on inputs at the decision boundary.
## Layout
```
common/ pooled features, the comparison rule, metrics, named pools
cache.py pooled feature cache for a COCO split
choose.py dimension selection on train2017, writes rules.json
verify.py scoring on val2017, writes eval.json
rtl_gen.py Verilog generation from rules.json
synth.py nosis synthesis, writes synth.json
infer.py loader for every rule
rtl/ one module per rule, all generated
tests/ consistency suite, no backbone or dataset required
```
Each measured JSON opens with a provenance block naming its generating script
and the pool it read. `tests/test_artifacts.py` enforces the pairing and that
selection and scoring name different splits.
## Running
```
pip install -e .
python cache.py --split train2017
python cache.py --split val2017
python choose.py
python verify.py
make synth
make test
```
`COCO_ROOT` is the dataset root. `BACKBONE` is the backbone repo id or a local
path. `BACKBONE_SRC` supplies `argus.py` from a local directory; otherwise it is
fetched from the backbone repo. Caching the two splits is a backbone forward
over 123,287 images; every later step reads the cache.
bfloat16 kernels select reduction orders by batch size, so cached values depend
on `--batch`. A cache must be built at one batch size throughout.
## Source backbone
EUPE-ViT-B from Meta FAIR ([arXiv:2603.22387](https://arxiv.org/abs/2603.22387),
Zhu et al., March 2026), distilled from PEcore-G + PElang-G + DINOv3-H+ via a
1.9B proxy teacher. License: FAIR Research License, non-commercial. This
classifier is an artifact derived from that backbone's feature geometry.
|