HybridDiffusion 2B - variable block-size curriculum (run-2)
Continuation of run-1 from step 6750 -> 7250 (500 optimizer steps), with exactly
one deliberate change from hybrid_diffusion_2b.toml: block size is sampled per
optimizer step instead of being fixed at 3.
| base | run-1 checkpoint step-6750 (model + optimizer + lr_scheduler + train_state) |
| ladder | b in {3, 4, 8, 16} at weights 0.35 / 0.15 / 0.15 / 0.35 |
| sampler | per optimizer step, broadcast from rank 0, seed 1234; pure function of (seed, step) |
| doc_alignment | 192 (= 2^6*3, so b=3 stays usable; 3 does not divide 64) |
| geometry sampler | OFF (mask_geometry_bernoulli = 1.0) - one variable only |
| lr | 1e-5 flat, no warmup (optimizer state carried, so no optimizer shock) |
| seq_len / global batch | 4096 / 256 |
| parallelism | HSDP, shard_degree=4 x replicate_degree=2, 8x RTX PRO 6000 Blackwell |
| format | torch.distributed.checkpoint (DCP), full state: model + optimizer + lr_scheduler + train_state |
Contents
checkpoint-7250-variable-block-size/ is a complete DCP directory (.metadata + __*_0.distcp shards),
resumable with torchtitan - not model-only.
Load
from huggingface_hub import snapshot_download
p = snapshot_download("Arushhh/hybrid-diffusion-2b-variable-block-size-curriculum", allow_patterns="checkpoint-7250-variable-block-size/*")
Then point checkpoint.initial_load_path at that directory.
Caveats
- 500 steps at 0.35 mass on b=16 is ~180 optimizer steps of b=16 exposure, not 500.
max_norm=1.0clips hard and unevenly by rung: b=3 retains ~51% of gradient magnitude, b=16 ~3.6%. Relevant when interpreting any null result.- Trained on shards held out from run-1's consumed region (Long-SFT 00365-00397,
Math 00128-00137, IF 00004-00007). Long-SFT
data-00398is the audit set and is excluded from training.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support