MadStickArt v0.8.0

MadStickArt is a trained, 61,296-parameter neural raster decoder for stick-figure storyboard scenes. It reconstructs 956 ร— 400 grayscale frames from four explicit layout masks: actors, setting, props, and motion. Its output contract is 2.39:1.

Try MadStickArt online โ†’ โ€” the browser trial is deployed separately. Its deployed model/version and renderer are recorded in the Space's model-provenance.json. The renderer included with this release supports 186 settings and 320 props; the deployed trial has its own model and renderer pins and is updated through a separate reviewed stage. Browser suggestions use transparent rules, while Strands decisions run in the downloadable Python pipeline.

The Python pipeline handles structured scene decisions and procedural composition. Strands Decider fills missing controls; explicit user controls and placement remain authoritative. The procedural renderer lays out geometry, and the trained neural decoder rasterizes that layout. This model does not generate story text or choose scene composition from pixels.

Pin revision v0.8.0 from danger-room/MadStickArt for reproducible downloads. The release workflow requires publication to both danger-room/MadStickArt and danger-room/MadStickArt-v0.8.0, with anonymous checksum, inference and Models-list verification for the separate entries. This prepared card does not establish completion of those steps. The manifest retains the stable source repository; publication receipts identify each destination. The separate initial repository danger-room/MadStickArt-v0.1.0 and its v0.1.0 tag remain preserved.

Stable release navigation

Use main for the latest stable release. Pin an immutable version tag for reproducible runs. This card describes v0.8.0.

The original standalone repository remains available at danger-room/MadStickArt-v0.1.0. The stable repository's v0.1.0 preserves that original model on the current release line.

Representative procedural teacher targets

Held-out target and decoder comparison

Component versions

Component Version or revision
Model weights MadStickArt 0.8.0 (v0.8.0)
Dataset/catalog 640 pairs; 1 genres; 4 source titles
Procedural asset library 0.19.0
Scene planner StrandsAgents/strands-decider-2B-hobson-v19 at bb282d786bc251fd4e3068de3ada9ddbb38127cd
Python runtime package madstickartmodel 0.1.0
Catalog SHA-256 d463fcddf21489b144204adb028e4b3357544b16fefec6dfc1759e8c60af6503
Dataset manifest SHA-256 610dd371c033491cb6678597f954484b20e87a4af96139d1bd8104634385e385

This release includes original generalized scene requests. The primary utility catalog has 4 original titles; retained sources preserve their separate catalog and title provenance in the source reports. Targets are saved procedural vector teacher renders; retained rasters are reused byte-for-byte. They are not movie frames or scripts. Source catalogs, code and small evidence reports are included; conditioning tensors, target datasets, planner downloads, decision caches and training logs are excluded.

Held-out decoder results

The checkpoint was selected by validation loss. On the title-isolated test split, it reached foreground IoU 0.6760 at threshold 0.35 on 160 scenes; pixel MAE was 0.0163. Validation foreground IoU was 0.6963 on 13,564 scenes. The bilinear union baseline scored 0.3050 test IoU.

These metrics measure agreement with the synthetic procedural teacher, not human judgments of story quality, composition, or legibility. The primary utility source is labeled drama in the original catalog. Pooled collections have uneven source and genre counts, reported with each comparison.

v0.8 retained-balance1 trains fresh weights for 16 fixed epochs on the unchanged 56,051-row collection: 40,535 unique training rows, 13,564 validation rows and 1,952 test rows. Each epoch presents all training rows once and all 3,840 retained-loop training rows once more, for 44,375 draws and 5,547 optimizer steps. This is 9.47% more optimization exposure than rotation-repair1. Exact per-source multiplicities and all sixteen planned and consumed order hashes are retained with the sampling policy. Validation and test rows never enter those draws.

The utility-v2 primary retains its separate rotation-repair1 data provenance: all 512 original rows and raster bytes plus 128 independently reviewed rendered signed rotations with original title/split ancestry. All other source populations, saved raster bytes, loss terms and validation weights remain unchanged. Validation loss alone selects the checkpoint. The failed rotation-repair1 candidate and its blocked readiness receipt remain under reports/predecessor/; they are not release acceptance evidence.

All nine unchanged gates at threshold 0.35 must pass against the original v0.7 checkpoint 8a9f1838e97121b99dd284e3ab74dae12c1bbbc836cf16528ee4308a530aafd9 on the actual 160-scene primary, full 1,952-scene interval, 5,632 pre-loop tests, 3,840 retained-loop tests and all 11,424 reviewed-pack tests. Canonical legacy and unseen-layout challenge comparisons remain required. The original 128/1,920 scopes remain separately labeled diagnostics with their absent dramatic-rotation failure. These metrics measure procedural-teacher raster agreement; fine strokes can be omitted and dense linework can overlap. Reusing these held-out populations limits the independence of repeated release experiments.

Campers remain a source-level regression on all 128 held-out scenes: foreground IoU 0.699724 to 0.686481 (-0.013243), precision 0.756608 to 0.761494 (+0.004886), and recall 0.902977 to 0.874510 (-0.028467). Native crops preserve the main tent, rope axes and actor poses but still merge crowded rope/hand/head strokes. This material recall loss is accepted with disclosure; aggregate passing scores do not establish uniform improvement.

The 128-scene transit source improves IoU 0.731608 to 0.735160 (+0.003552) and precision 0.792103 to 0.808906 (+0.016803), while recall declines 0.905477 to 0.889671 (-0.015806). Fine route curves and crossings remain softened in native inspection.

The 128-scene plumber source trades detail for slightly higher precision: IoU 0.697656 to 0.687492 (-0.010164), precision 0.770588 to 0.771613 (+0.001025), and recall 0.880544 to 0.863128 (-0.017416). Main plumbing geometry is recognizable in native crops, while small valves and thin cabinet strokes remain imperfect.

On 3,840 retained-loop test scenes, foreground IoU improves from 0.686586 to 0.690124 and precision from 0.743680 to 0.755601, while recall falls from 0.899429 to 0.888442. The retained aggregate passes all nine unchanged gates; that does not mean every source or metric improves.

Fine forms, buttons, valve circles, route marks and dense actor/rope crossings remain softened, darkened or merged. Recognizable object placement does not imply legible small details or exact teacher reproduction.

Rotation and floor placement are supplied by frozen scene geometry and procedural conditioning. Both v0.7 and this candidate follow those axes. Improved rotation-fixture scores reflect neural stroke reconstruction, not newly learned scene or angle selection.

The same title-isolated fixed test populations and native fixtures were exposed during review of the held rotation-repair1 candidate. They are regression evidence, not a newly blind generalization study. Validation loss alone selected epoch 11 after all 16 epochs.

Procedural targets contain some crowded overlaps and faint construction lines. The original action source retains its historical two within-title/split repeated pairs and four condition-collision groups with different targets; its disclosed target limitations are unchanged.

Foreground IoU, precision and recall measure agreement with saved procedural targets at threshold 0.35. Native contact sheets and crops complement these metrics; they are not human narrative-quality scores or browser/ONNX parity evidence.

Measured throughput

The benchmark ran at 956 ร— 400, batch size 1, after 10 neural warmup images plus 3 configured and 1 fresh full-scene warmups, on Apple M5 Max, 40 GPU cores and 128 GB unified memory. The full pipeline measured 5.65 scenes/s across 100 scenes and 100 uncached Strands calls / 502 typed choice questions across 20 genres. Its scope includes layout, neural rendering, annotations, clean and annotated PNGs, and JSON; it met the five-scenes-per-second goal under those conditions. Model download and startup are excluded.

Prepared-scene to PNG measured 55.54 scenes/s and neural forward only measured 1228.20 images/s. Their narrower timing scopes are recorded in reports/benchmark.json. These Apple M-series measurements do not establish browser or base-M5 throughput.

Run the pinned model

hf download danger-room/MadStickArt --revision v0.8.0 --local-dir ./MadStickArt-v0.8.0
cd MadStickArt-v0.8.0
UV_CACHE_DIR="$PWD/.cache/uv" uv sync --locked --extra planner
uv run --no-sync madstick storyboard examples/multigenre-storyboard.jsonl \
  --checkpoint best.pt --device mps --planner strands --planner-device mlx \
  --output generated-storyboard

Use --device cpu when Apple MPS is unavailable. The optional Strands planner downloads its own model weights separately; those weights are not part of MadStickArt. model.safetensors is the canonical safe checkpoint and best.pt preserves training-checkpoint compatibility.

Versioned evidence is included in reports/: training and held-out metrics, benchmark, dataset summary and validation, catalog audit, asset catalog, and component versions. The full training dataset is not distributed.

Downloads last month
160
Safetensors
Model size
61.3k params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using danger-room/MadStickArt 1

Evaluation results

  • Foreground IoU (ink threshold 0.35) on MadStickArt synthetic 1-genre title-held-out test split
    self-reported
    0.676
  • Pixel mean absolute error on MadStickArt synthetic 1-genre title-held-out test split
    self-reported
    0.016