HybridDiffusion 2B - variable block-size curriculum (run-2)

Continuation of run-1 from step 6750 -> 7250 (500 optimizer steps), with exactly one deliberate change from hybrid_diffusion_2b.toml: block size is sampled per optimizer step instead of being fixed at 3.

base run-1 checkpoint step-6750 (model + optimizer + lr_scheduler + train_state)
ladder b in {3, 4, 8, 16} at weights 0.35 / 0.15 / 0.15 / 0.35
sampler per optimizer step, broadcast from rank 0, seed 1234; pure function of (seed, step)
doc_alignment 192 (= 2^6*3, so b=3 stays usable; 3 does not divide 64)
geometry sampler OFF (mask_geometry_bernoulli = 1.0) - one variable only
lr 1e-5 flat, no warmup (optimizer state carried, so no optimizer shock)
seq_len / global batch 4096 / 256
parallelism HSDP, shard_degree=4 x replicate_degree=2, 8x RTX PRO 6000 Blackwell
format torch.distributed.checkpoint (DCP), full state: model + optimizer + lr_scheduler + train_state

Contents

checkpoint-7250-variable-block-size/ is a complete DCP directory (.metadata + __*_0.distcp shards), resumable with torchtitan - not model-only.

Load

from huggingface_hub import snapshot_download
p = snapshot_download("Arushhh/hybrid-diffusion-2b-variable-block-size-curriculum", allow_patterns="checkpoint-7250-variable-block-size/*")

Then point checkpoint.initial_load_path at that directory.

Caveats

  • 500 steps at 0.35 mass on b=16 is ~180 optimizer steps of b=16 exposure, not 500.
  • max_norm=1.0 clips hard and unevenly by rung: b=3 retains ~51% of gradient magnitude, b=16 ~3.6%. Relevant when interpreting any null result.
  • Trained on shards held out from run-1's consumed region (Long-SFT 00365-00397, Math 00128-00137, IF 00004-00007). Long-SFT data-00398 is the audit set and is excluded from training.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support