brain-zero / README.md
rfick's picture
Squash history: one commit holding the current files
6bc44d6
|
Raw History Blame Contribute Delete
8.56 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade
metadata
title: 'Brain replay: MASiVar'
emoji: 🧠
colorFrom: indigo
colorTo: blue
sdk: gradio
sdk_version: 6.28.0
app_file: app.py
pinned: false
startup_duration_timeout: 1h
license: mit
short_description: A real brain replayed from packs, to an 84-region connectome

The brain Space on ZeroGPU

The same application as rfick/disco-zero (dmrai-lab/disco-space), serving space/brain.toml (DISCO_CONFIG): a real brain as a compositional replay phantom. The input is the multi-tissue CSD of one scan, MASiVar sub-cIs1 (OpenNeuro ds003416, Cai et al. MRM 2021, CC0): its white-matter FOD field, its WM / GM / CSF fractions and an 84-region parcellation, made once offline and published as the dataset SubstrateCommons/masivar-brain. Each run composes every brain voxel from two replay packs (CACTUS axons for WM, packed spheres for GM; free water in closed form for CSF) at the chosen acquisition, tissue and scanner, adds Rician noise, estimates the tissue responses from the replayed data, fits CSD, tracks from the white matter and scores the 84 x 84 connectome (and its 14 lobar groups) against the connectome of the input FOD tracked with the same tracker, settings, seeds and key. Source and pins: https://github.com/dmrai-lab/disco-space (requirements-zero.txt, app.py, space/sources/brain.py).

The scanner. The menu holds the ideal scanner and two catalogued machines, the Siemens Prisma 3 T and Terra 7 T, each played at every voxel as it delivers there with the head centre at isocentre (voxels binned into encoding classes at 1 % of b). The Hyperfine Swoop 64 mT is not on this menu, and the page says why under it: its own gradient is exact in dmipy-sim's closed form, but it bins this head into 9,752 classes at 2-3 minutes each (docs/scanner.md). The DiSCo Spaces play it exactly.

Where the work runs. The packs' pose responses (the only physics that depends on the tissue and the field) are computed on the CPU and kept in the page's process for the container's lifetime per pack, protocol, tissue and field: for every preset (with its ladder) and the scan's protocol at every field preset when the container starts, for any other run inside its GPU call (the page only looks its cache up before the call, so the call follows the request at once; ZeroGPU's proxy token expires when the call comes late), after which the page keeps them too. Inside the call, torch contracts them with every voxel's FOD, fractions and proton densities (voxels x measurements x 45), then the noise, the reconstruction (dmipy-fit), the tracking and the truth's tracking (dmipy-tract), the scoring. Deterministic algorithms are on and TF32 off inside the call.

GPU quota per visitor

Every run reserves the GPU for the seconds the page shows under the run button. ZeroGPU charges that reservation against the visitor's daily quota, not the Space's: 2 minutes logged out, 5 with a free Hugging Face account, 40 with PRO; a single request above the visitor's quota is refused before it starts, and a run that outlives its reservation is killed. Only the device part of a run holds the GPU, with the packs' responses the page has not cached (part of the reservation shown, which falls once a run has cached them); the page's files and figures after it do not.

Not re-measured at the current pins (dmipy-sim 1b79d2e; the tables below name their own). Measured on an L40S with the BATMAN development fixture (96 x 96 x 60 at 2.5 mm, 90,205 brain voxels, 36,605 in the WM stop mask; dmipy-sim 7f6f1fa, dmipy-fit 0c7dde8 with the single-tissue reconstruction, dmipy-tract 7da22c3; tools/measure_brain.py, 2026-09-30), steady state, seconds:

stage 113 measurements 495 measurements (five shells x 96)
the packs' responses, every tier, on the host CPU (not charged) 4.9-6.1 (8 threads) 7.5-7.9 (8 threads)
the replay (the contraction on the device) 0.04 0.17
noise 0.35 1.4
reconstruction (single-tissue: tournier07 response + CSD) 3.7 8.0
the round trip against the input FOD 0.85 0.85
tracking, density 1 / 2 / 4 (36,605 / 292,840 / 2,342,720 seeds) 0.7 / 1.9 / 13.6 0.7 / 1.9 / 13.6-16.4
the truth's tracking, density 1 / 2 / 4 0.8 / 2.1 / 14.8 0.7 / 2.1 / 23.2
the ladder (three rungs of contraction) 0.07 0.5

The reservation (space/brain.toml, [budget]) is each device stage x 1.5 for the pool, 12 s for the worker's start and the handoff, times 1.3: at the default density 2 and 495 measurements, A alone 49 s, A with the ladder 50 s, A + B + ladder 79 s, so every default configuration with one B fits a logged-out visitor's 2 minutes; density 4 reserves 131 s for A alone and is for logged-in visitors.

Measured on the pool (tools/live.py rfick/brain-zero --config brain.toml, 2026-10-01, the MASiVar asset at revision 2 unless the row says revision 3: 71,052 brain voxels, the scan's own protocol of 485 measurements, density 2, 3 T along the bore, every tier, SNR 30 at M0 = 1, the multi-tissue reconstruction; the packs' responses cached on the page):

run reserved device held handoff to the page wall through the API connectome Pearson log(1 + count) vs the input's
A + ladder 50 s 13.6 s 4.5 s 103 s 0.903 (lobar 0.980), 228,760 streamlines
A + B (field → 7 T) + ladder 78 s 22.1 s 6.6 s 149 s A 0.903, B 0.889, A vs B 0.962
A + ladder, the windowed WM pack (single_bundle_1s_c3_seg125ms, window 0 of 8; 2026-10-01 13:10, logged out) 50 s 13.9 s 4.7 s 96 s 0.893 (lobar 0.975); the container served after 22 s, its warm-up in a process beside the page
A + B (every shell's pulse timing → long-TE δ 30 / Δ 120 ms, TE 160 ms: two windows of both packs) + ladder (2026-10-01 14:55, logged out, the responses cached by the warm-up) 78 s 22.8 s 6.8 s 362 s (the pool's queue included) A 0.893, B 0.839 (lobar 0.949; the median voxel's b = 0 SNR falls from 7.0 to 3.7 at TE 160 ms), A vs B 0.943
A + ladder on revision 3 of the asset (the responses from the eroded mask, Tournier 2013; 2026-10-01 16:20, with an account token: a logged-out address has two runs a day) 50 s 13.4 s 4.5 s 81 s 0.932 (lobar 0.981); the FOD round trip against the new truth: principal peak 46.5° median, AFD r 0.826 (the truth now holds 5,498 two-peak WM voxels against 2,272)
A + ladder on revision 2 of the WM pack (built from the kept walks, each window's own contact channel, no contact envelope; dmipy-sim#528; 2026-10-02 00:03, account token, the page's cache warm) 50 s 22.9 s of compute — 96 s 0.930 (lobar 0.983), the same to the third digit as on the resegmented cut

The warm-up's timeline (Space commit 44fa8da7, 2026-10-01 18:06 UTC: 32 entries, one per response; the container has 192 CPUs visible and a cgroup quota of 16, 104 GB; RSS 4-15 GB): the default run's responses 42 s after the process started, every preset with its ladder by 2.5 min, the one-window knobs by 7 min, the first two-window class at 9 min (112 s: its band's coupling tables are built once per container), the second at 9.5 min, then the field presets in ascending field. Until an entry is cached, a run asking for it computes it inside its own reservation at the cold price; a logged-out visitor's 120 s then holds the default run from the first minute, a two-window class once the warm-up reaches it, and another field once that field is cached. Before dmipy-sim#532 the first two-window entry took 1292 s on this pool (1217 s of it page faults in the field factor's numpy route) and the 11.7 T band 936 s.

Device stages of A at 485 measurements on the pool: replay 0.3 s, noise 1.7, the three-tissue responses 1.2, MT-CSD 3.1, the round trip 1.2, tracking 2.5, the truth's tracking 2.3 (B reuses A's). The page then spends 9-17 s writing files and 4-6 s on states and figures, outside the reservation.

The payload that crosses from the GPU worker back to the page is one Result per run (the DWI before and after the noise in float32 on the brain's bounding box, the FOD field, the streamlines): 1.5 GB per run at 495 measurements and density 2 (0.27 GB of it streamlines; 3.4 GB at density 4), plus the ladder's rungs; the pool hands it off in 4.5 s for A alone and 6.6 s for A + B.

The truth connectome is reproducible bit for bit at one key; between two keys at the same density its Pearson of log(1 + count) over the 3,486 region pairs is 0.834 / 0.945 / 0.982 at density 1 / 2 / 4, the floor of the score.