YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

MoS-27B-Checkpoints

Epoch checkpoints of the Qwen3.8-27B DFlash2 MoS line (project MoS, branch ryan/mos-improve, experiments/dflash2/qwen3.8-27b-100k/). Started 2026-09-22, when the general archive repo ryan-0608/MoS-Aurora-Experiment-Archive reached the Hugging Face 20,000-file limit.

Layout

dflash2_27b_traj_20260915/main-800k/<arm>/epoch<N>/
    model.safetensors            drafter weights (speculators DFlash2MoSDraftModel)
    optimizer_state_dict.pt      full AdamW state -> exact resume
    scheduler_state_dict.pt, training_state.json, config.json, config.py, train_command.txt
    val_metrics.json             trainer held-out 10% validation at the end of the epoch
    serving/                     fixed-500 serving result for this checkpoint
        acceptance_summary.json  pooled AL = sum completion tokens / sum verify calls
        acceptance_trace.jsonl   per-request counts
        client.log, server.log

Arms

arm (path) Ryan's name recipe
dense4 27B dense (4-epoch baseline) dense drafter warm-started from z-lab/Qwen3.8-27B-DFlash2 @50307d4c; base LR 1e-4; cosine over 4 epochs, warmup 0.005; same corpus, 3 verifier + 5 trainer layout, global batch and 73,670 steps per epoch as x4rand; training (started 2026-09-25)
x4rand 27B 4expert rand LR6 K=4 full-width expert MLPs per draft layer, random init (gate/up N(0,0.02), down N(0,1e-3)); shared MLP, attention and the rest warm-started from z-lab/Qwen3.8-27B-DFlash2 @50307d4c; base LR 1e-4, expert LR 6e-4, router LR 5e-4; top-2, tau 0.9, balance 0.01; cosine over 4 epochs, warmup 0.005; corpus Current-800K-Self-27B-Traj (719,455 train rows); 73,670 steps per epoch

Results (target Qwen/Qwen3.8-27B @1d4bf0f2)

Serving protocol: sglang main f5866545, TP2, DFLASH 16 draft tokens, fixed-500 prompts, max 128 new tokens, greedy, concurrency 1, 500/500 completed. Same protocol as the dense and K4-jitter references.

arm epoch global_step train-side val AL serving pooled AL vs dense same epoch
dense4 1 73,670 5.0262 5.0571 โ€”
dense4 2 147,335 5.0261 5.0510 โ€”
x4rand 1 73,670 5.1485 5.2314 +3.45% (dense4 5.0571)
x4rand 2 147,335 5.2056 5.2922 +4.78% (dense4 5.0510)
x4rand 3 221,001 5.2202 5.3261 +5.46% vs dense E2 5.0502 (dense stopped after E2)
x4rand 4 294,682 5.2201 5.3113 +5.17% vs dense E2 5.0502; best epoch = 3

Reference arms (model-only, no optimizer) are in ryan-0608/MoS-Aurora-Experiment-Archive under dflash2_27b_traj_20260915/main-800k/: dense/epoch1, dense/epoch2, mos/epoch1 (K4 shared+1% jitter, expert LR 1e-4; serving 5.1501, +1.83%).

Use

Resume training: pass the epoch directory as the checkpoint dir to speculators train.py (same train_command.txt, --epochs 4). Serve: export_d2_for_sglang.py <epoch dir> <out> then sglang --speculative-algorithm DFLASH --speculative-draft-model-path <out>; the MoS architecture needs apply_serving_dflash2_mos.py on sglang main.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support