MuonShower models
Flow-matching generators and muon-count heads for simulating cosmic-ray muon showers
(CORSIKA I3MCTree_preMuonProp, primaries with zenith in [55, 65] degrees).
Code lives in the MuonShower repo. Clone it,
install environment.yml, then either run the end-to-end script or load models directly --
don't load .safetensors/.pth files by hand, hf_pipeline.py reconstructs the right
model class from the config.json next to each weight file and loads the state dict for you.
Quick start: end-to-end reproduction
CUDA_VISIBLE_DEVICES=0 ./venv/bin/python reproduce.py \
--generator block_diffusion_bs200 --counter block_count_200 \
--data_dir /path/to/your/corsika/test/split
Downloads every checkpoint in this repo (~3.1 GB, cached after the first run), predicts
each event's muon count, samples the shower, decodes to physical units, and writes
events.csv / particles.csv to --out_dir (default ./reproduce_out). --data_dir must
point at your own CORSIKA HDF5 split -- that dataset is not hosted here, only the weights
are. For a quick check instead of the full split: add --max_files 1 --n_events 5.
Loading a model directly
from hf_pipeline import load_generator, load_counter
gen = load_generator("block_diffusion_bs200") # {"model", "matcher", "config"}
cnt = load_counter("block_count_200") # {"nblocks", "count", "lastblock", "config"}
Generators (generators/)
| name | block layout | objective |
|---|---|---|
original_2000 |
1 block x 2000 slots | plain flow matching, no auxiliary terms |
block_diffusion_bs200 |
10 blocks x 200 slots | two-stream block diffusion (BD3-LM style), flow-matching loss on real slots only |
block_diffusion_bs500 |
4 blocks x 500 slots | same recipe, different block size |
All are MuonDiffusionTransformer (dim=1024, depth=16), trained with DELTA_ANGLES=True
(muon zenith/azimuth encoded relative to the primary) -- decoding requires that setting in
muon_dataset.py/flow_matching.py to match.
Counters (counters/)
| name | what it predicts | input |
|---|---|---|
poisson_2000 |
total muon count N, direct Poisson regression | primary only |
block_count_10 |
n_blocks (Poisson) + per-block count, block_size=10 | primary [+ muons for the count head] |
block_count_200 |
n_blocks + per-block count + a dedicated last-block head, block_size=200 | primary [+ muons] |
block_count_500 |
n_blocks + per-block count, block_size=500 | primary [+ muons] |
block_count_200/lastblock_classifier.pth is generator-specific. It was trained on
muon prefixes generated by generators/block_diffusion_bs200 (checkpoint_044), not on
real data, so that training matches what it sees at inference. Using it with a different
generator checkpoint is untested and likely to underperform its measured numbers:
| stage-3 head | full-N MAE (400-event stratified test) | predictions landing on a multiple of 200 |
|---|---|---|
| naively trained (target-alignment bug) | 107.2 | 0.000 |
| trained on true prefixes (ceiling, not deployable) | 76.8 | 0.177 |
| trained on generated prefixes (shipped) | 65.5 | 0.000 |
poisson_2000 baseline (no generator involved) |
57.2 | 0.003 |
No count read off a generated grid can beat the primary-only baseline: the generated
muons carry no information about the true count beyond what the primary already gives
(x_gen = f(primary, noise), noise independent of the true shower). The value of the
two-stage pipeline is removing degenerate behavior (the 200/400/600/800 quantization above),
not beating the baseline on accuracy.
classifiers_nblocks (the nblocks head in block_count_200) was not trained with
tail-balancing; its exact-match accuracy degrades from ~0.99 (single-block events, 96.9% of
the data) to 0.14-0.55 on events needing 2-6 blocks.
Caveats carried over from the source repo
block_diffusion_*checkpoints should not be re-selected byval/flow_lossalone -- that metric improved monotonically through training while across-event generation diversity (what any downstream count read-off depends on) got worse. These checkpoints are the ones actually evaluated above, not the training-loss optimum.original_2000and theblock_diffusion_*runs are trained on different data splits / epoch counts and are not loss-comparable to each other; compare downstream task metrics, not raw training loss, across families.