nanolab

A harness for deciding whether a small vision model is worth shipping.

The models are not the point. This harness has killed more candidates than it has shipped, which is the behaviour you want from a harness. It exists because the most common failure in small-model work is not a bad model β€” it is a good-looking number that a single threshold would have matched.

from nanolab import scalar_baseline, verdict, report
from nanolab.answerability import answerable, zoom_for, camera_spec

1. The bar: cheap baselines that are actually hard to beat

cheap_stats computes ten statistics per image β€” six global (mean, std, Laplacian variance, high-frequency ratio, gradient, entropy) and four spatially aware (centre-vs-edge, radial slope, row and column FFT peaks).

The spatial four exist because of a measured mistake. With global statistics only, a vignetting task showed a baseline of 0.707 against a model's 0.967 β€” an apparent +0.261 margin. Adding a radial-slope statistic took the baseline to 0.907 and the margin to +0.060. A baseline that cannot see where anything is loses to a CNN by default, and flatters it.

For multi-class problems a single threshold is too weak a bar (the implementation uses two cuts, so three bins, and cannot express six classes). Use transferred_baseline plus a small linear model over the same statistics β€” a ~66-parameter logistic. On one task that raised the bar from 0.438 to 0.543 and cut the model's apparent margin from +0.183 to +0.077.

2. The gates: four, in order of what they rule out

verdict() is the single place the logic lives, because re-implementing gates is how gates get silently dropped.

  1. Utility β€” lift over the majority class β‰₯ 0.15. Statistically real is not the same as useful.
  2. Ordinal β€” for ordered targets, mean error below a bin threshold.
  3. Scientific β€” held-out model vs the in-sample scalar. Deliberately optimistic. Asks: is there structure beyond a low-order statistic?
  4. Engineering β€” transferred model vs transferred scalar. Asks: what should actually ship?

Gates 3 and 4 disagree, and that is the point. On block-grid detection the in-sample scalar read 0.954 while the same threshold transferred read 0.565.

3. Answerability: is the task possible at all?

Based on Johnson's criteria (John Johnson, 1958): ~2 px across a target to detect it, 8 to recognise its type, 12.8 to identify a specific one. Measuring CNN attribute classification against per-attribute AP reproduces the recognition figure β€” accuracy degrades sharply below ~8 px on the diagnostic feature, and the effect survives controlling for training-set size (+0.074 / +0.112 / +0.144 within support bands).

answerable(feature_mm=6, reference_mm=1700, reference_px=0.25*720)
# (False, 0.6, 8.0)   -- a nasal cannula from a bedside 720p camera. Not a model problem.

zoom_for(feature_mm=5, input_px=224)
# 140.0   -- crop to 140 mm to put 8 px on a 5 mm feature. Zoom, do not upscale.

This is how we established that IV lines and nasal cannulae are hopeless from a room camera at any practical resolution β€” eight pixels on a 6 mm cannula at bedside framing needs a 9,067-pixel-tall sensor. That conclusion cost one line of arithmetic instead of a labelling campaign.

4. Diagnostics: does resolution buy anything?

resolution_sweep sweeps input size at constant parameter count (valid only for global-average-pooled architectures, else it confounds pixels with capacity).

A result that held in three independent domains β€” garment attributes, clinical attire, clinical scenes: cheap baselines are flat with resolution while learned models are not.

domain model 64 β†’ 256 baseline 64 β†’ 256 margin
clinical attire (mask head) 0.401 β†’ 0.587 0.280 β†’ 0.282 +0.121 β†’ +0.305
clinical scene (hard task) 0.640 β†’ 0.699 0.612 β†’ 0.620 +0.028 β†’ +0.079

Global aggregates wash out detail however many pixels they are given. Resolution is precisely where a model earns its keep over a threshold β€” and if your margin does not grow with resolution, the model may not be doing anything a statistic could not.

5. Splitting: by group, never by sample

Adjacent samples are near-duplicates. A random split puts them on both sides and reports a number that means nothing β€” one quality-gate result went from a meaningless 1.000 to an honest +0.238 when split by group.

Two traps we hit and you will too:

  • Group splits must require both classes on both sides. A size-only split once produced a test set that was 93% one class, and the model scored 0.070 β€” far below chance.
  • Weight by test-set size when groups differ in size. Unweighted averaging over unequal splits reported +0.215 where the sample-weighted figure was +0.129. Constructed units (one image, one patch) are uniform and safe; natural units (a run, a session, a patient) are not.

Installation

pip install -e .

Why publish a harness rather than a benchmark

A benchmark tells you who won. This tells you whether the contest was worth entering. Of the models built with it, several were refused for reasons worth naming: occupancy lost to a threshold on thermal image variance (0.85 vs 0.77); shot-scale lost to a threshold that was itself barely above chance; a studio-vs-street classifier scored 0.994 where one entropy threshold got 0.957.

All three looked like results until the baseline was computed.

6. Placement: where do the cameras go?

nanolab.placement turns answerability into a layout question. A room is a rectangle (metres) plus targets, each with the mm that decide its label and the Johnson level it needs; candidates are wall mounts with a lens, a sensor width and a yaw. A camera covers a target when the target is inside its horizontal FOV and the pinhole model puts enough pixels on the critical dimension (feature_mm / (2Β·dΒ·tan(fov/2)) Β· sensor_px, with slant distance so mount height counts). Fewest cameras covering everything is minimum set cover; the greedy solver is within ln(n)+1 of optimal.

from nanolab.placement import target, wall_candidates, place, format_table, demo
demo()   # 6 m x 4 m two-bed bay, candidates every 1 m on all walls at 2.4 m, 78 and 110 deg lenses
camera                 covers
(0,0) 78deg            bed A (901px), bed B (475px), face A (77px), face B (37px), door (156px), IV pump display (25px)

uncoverable            best px   need   closest camera
cannula A                  3.9   12.8   (1,0) 78deg
cannula B                  3.9   12.8   (5,0) 78deg

1 camera(s) chosen from 40 candidates; greedy is within ln(n)+1 = 2.79x of optimal.

The chosen camera is almost incidental. The useful output is the second block: no wall mount in the room gets a 6 mm cannula past 3.9 px, a factor of 3.3 short of the 12.8 px identify line, even from the nearest wall position with the narrower lens. That is the same conclusion Β§3 reached with one line of arithmetic, now reached with every candidate tried β€” cannulas need a different sensor, not a better spot on the wall.

6. Effective sample size: many samples of one event is still n = 1

Datasets are grouped β€” frames by scene, face track or movie; clips by speaker or upload β€” and the number of independent groups, not rows, governs generalisation. 1,697 frames of one room once gave transfer numbers with a Β±0.255 spread across seeds. effective_n turns that into a number before anything is trained:

from nanolab import effective_n, effective_n_report
effective_n_report("bed-state", X8x8, scene_ids)
# bed-state: 1210 rows in 4 groups, ICC 0.02, n_eff = 202 (16.7% of N)

icc(values, groups) is the one-way random-effects ICC(1) (ANOVA estimator, unbalanced groups via k0 = (N βˆ’ Ξ£n_gΒ²/N)/(Gβˆ’1)); design_effect is Kish's 1 + (mΜ„ βˆ’ 1)Β·ICC; n_eff = N / deff, which lies between G and N. Measure it on a cheap summary of the model input (an 8Γ—8 downsample, a per-clip mean spectrum), not on the label; for a 2-D summary the ICC is taken per column and the median used.

Run on the datasets in this repo (code/effective_n_report.py):

dataset (input summary) group N G ICC (median) n_eff n_eff/N
bed-state, Places365 (8x8) scene 1210 4 0.02 202 0.17
bed-state, Places365 (mean,std,lapvar) scene 1210 4 0.03 137 0.11
mouth-gate test (8x8) movie 6105 4 0.20 20 0.003
mouth-gate test (mean,std,lapvar) movie 6105 4 0.36 11 0.002
mouth-gate test (8x8) face track 6105 1645 0.81 1911 0.31
audio_v2 train (mean spectrum) src (one id per clip) 8807 8807 n/a 8807 1.00
audio_event v1 clips (mean spectrum) src 3234 1804 0.81 1974 0.61
β€” LibriSpeech rows only speaker 994 40 0.33 112 0.11
β€” ESC-50 rows only upload 2000 1524 0.82 1592 0.80

What the numbers say:

  • audio_v2 (transferred) is the sanity case: every clip its own source, so ICC is not estimable and n_eff = N by construction. That is bookkeeping, not proof of independence β€” src does not record the FSD uploader.
  • audio_event v1 (speech/alarm did not transfer): the speech class is 994 rows from 40 LibriSpeech speakers, n_eff β‰ˆ 112. The model saw forty voices. ESC-50 clips from one upload are near-replicates (ICC 0.82) but there are 1,524 uploads, so that side is fine.
  • bed-state (did not transfer): four scene categories; even at ICC 0.02, G = 4 gives a design effect of 6 and n_eff β‰ˆ 200. The ICC is low because a Places scene category is hundreds of different rooms β€” the number to worry about is G = 4 scene types, and the actual failure was the Places-to-room-camera domain shift, which no within-dataset statistic can see.
  • mouth-gate test: 6,105 stacks but 4 movies; n_eff β‰ˆ 20 at movie level. Any test number from it is a 4-movie result. At face-track level the frames are near-replicates (ICC 0.81, 3.7 frames per track), n_eff β‰ˆ 1,900.

The rule this encodes: n_eff is bounded by G whatever the ICC, so report G alongside N and be suspicious of any dataset where n_eff < 0.1 N (the helper warns). A low ICC does not rescue a small G; it only says the groups are internally diverse, which is a different question from how many of them there are.

7. hmm_filter: an unsupervised filter for one noisy sensor (0.3.0)

fit_filter(obs) fits a 2-state HMM by Baum-Welch to a single discrete observer stream with no labels and returns causal P(on | past). Measured on a radar occupancy bit that flickers with a median run of 2 ticks: the learned emission matrix matched the empirical confusion matrix to max |dE| = 0.07, and causal filtering scored 0.442 on state-transition ticks vs 0.104 for a 12-tick majority filter and 0.571 for the raw bit (offline Viterbi 0.675). It replaces the finite-state grammar gate that refused a quarter of the sensor's own ground truth.

Boundary, also measured: with two observers Baum-Welch latched onto the most persistent latent variable (the lighting schedule, not the person). Feed it ONE observer, and read E before trusting the filter.

8. sidecar: per-frame metadata inside H.264 (0.3.0)

inject_mp4(src, dst, json_payload(fn)) puts a distinct SEI user_data_unregistered NAL before the first slice of every frame, patching the mp4's samples in place so B-frame timing survives; extract(path) reads them back from .h264, .mp4 or .mkv. Measured on 300 frames of 720p with a 300-byte JSON payload: 300/300 intact and pts-aligned through remux and Annex-B copy, 0/300 through a default transcode, 300/300 through a transcode with -udu_sei 1. Overhead 325 B/frame, which is 9% of a 0.9 Mbit/s stream and <1% above ~8 Mbit/s. ffmpeg's own h264_metadata=sei_user_data stamps keyframes only (2/300), which is why this exists.

9. sidecar1: what the bytes mean (0.4.0)

sidecar.py moves bytes; sidecar1.py is the Sidecar/1 envelope spec (docs/SIDECAR1.md, 2026-09-05) with a validator and builders. One JSON object per frame; every measurement row carries the answerability record that produced it β€” which model, how many pixels were on the target against the 2 / 8 / 12.8 px Johnson line, whether the gate passed, what the scalar baseline scored on the card β€” and refusals are rows, not absences. p alone is never a validity signal (confidence is worse than random under shift); gate.passed is, and a row without a gate is accepted with a warning. Coordinates default to the original sensor frame, because the record was computed before downscaling.

from nanolab import make, row, gate, validate, fingerprint, rebind
env = make(f=17, t=ts, fp=fingerprint(luma), src={"cam": "bay2", "res": [1920, 1080], "tx_res": [640, 360]},
           m=[row("faces", None, by={"model": "face-nano", "ver": "b77", "kind": "nano"},
                  gate=gate(px_on_target=3.9, task="recognise"), refused=True, why="3.9 px, recognise needs 8")])
validate(env)   # [] -- errors and warnings as strings, empty is valid

fp is a 64-bit pHash of the frame luma (32x32 block mean, 8x8 DCT, sign vs median; specified step by step so a JS receiver can reproduce it). It is how a receiver recovers the frame binding after a transcode strips the SEI: rebind(rows, frame_fps) aligns the two fingerprint sequences monotonically (banded Needleman-Wunsch). Measured on 600 frames after 2x downscale + CRF 28: the hash survives at median 0 / p99 4 bits, but consecutive frames are only ~4 bits apart, so nearest neighbour rebinding is right for 0.86 of frames and alignment for 1.00 (also with 2% drops). On a static scene neither works and ambiguous() says so (1.0); fall back to timestamps.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including resoajoe/nanolab

Article mentioning resoajoe/nanolab