next JEV stage 1, last-block CoT decoder
Stage 1 (CoT only) with a new rationale decoder (aux_decoder: last_block): a new cross-attention block over the 128 workspace tokens, then an exact copy of Qwen3.5-0.8B's last block (layer 24, full causal attention) and final norm, with tied embeddings in and out. One pass through the backbone, no loops (loops: 1). The decoder's learning rate is 2x the rest (aux_learning_rate_scale: 2.0).
Training was stopped on purpose at step 2693 of 6347 to start stage 2; checkpoint-00002600/ is the last stage 1 checkpoint and the stage 2 init. Not validated on the full validation set.
- Checkpoints:
checkpoint-00002600/(stage 2 init) andcheckpoint-00002300/.checkpoint.pt= FP32 weights + Adam state + training state; load withnext_jev.train.load_checkpoint - Code: github.com/JonnesLin/next_jev, branch
feat/stage1-last-block-decoder(commit56cc80c), configconfigs/stage1_last_block_decoder.json - Data: stage 1 multi-response CoT from
JonesLin/multi-model-cot-2730, 406,167 train prompts - Workspace: 128 tokens, 1 pass,
middle_workspaceblocks 1-24 - Training: global batch 64 on 2x H100 NVL, lr 8e-05 (decoder 1.6e-04), cosine to 10%, no warmup, BF16 autocast with FP32 master weights
metrics.jsonl: per-step training metrics up to step 2600eval.jsonl: every 200 steps, the CoT loss and greedy decodes of 10 held-out validation rows (5 text, 5 image). CoT loss on these rows: 3.74 at step 200, 2.18 at step 2200, 2.10 at step 2600.- W&B: https://wandb.ai/fast-ssl/next-jev/runs/59ca1d12a252417d9bc05aaba997cc71
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support