next JEV stage 1, last-block CoT decoder

Stage 1 (CoT only) with a new rationale decoder (aux_decoder: last_block): a new cross-attention block over the 128 workspace tokens, then an exact copy of Qwen3.5-0.8B's last block (layer 24, full causal attention) and final norm, with tied embeddings in and out. One pass through the backbone, no loops (loops: 1). The decoder's learning rate is 2x the rest (aux_learning_rate_scale: 2.0). Training was stopped on purpose at step 2693 of 6347 to start stage 2; checkpoint-00002600/ is the last stage 1 checkpoint and the stage 2 init. Not validated on the full validation set.

  • Checkpoints: checkpoint-00002600/ (stage 2 init) and checkpoint-00002300/. checkpoint.pt = FP32 weights + Adam state + training state; load with next_jev.train.load_checkpoint
  • Code: github.com/JonnesLin/next_jev, branch feat/stage1-last-block-decoder (commit 56cc80c), config configs/stage1_last_block_decoder.json
  • Data: stage 1 multi-response CoT from JonesLin/multi-model-cot-2730, 406,167 train prompts
  • Workspace: 128 tokens, 1 pass, middle_workspace blocks 1-24
  • Training: global batch 64 on 2x H100 NVL, lr 8e-05 (decoder 1.6e-04), cosine to 10%, no warmup, BF16 autocast with FP32 master weights
  • metrics.jsonl: per-step training metrics up to step 2600
  • eval.jsonl: every 200 steps, the CoT loss and greedy decodes of 10 held-out validation rows (5 text, 5 image). CoT loss on these rows: 3.74 at step 200, 2.18 at step 2200, 2.10 at step 2600.
  • W&B: https://wandb.ai/fast-ssl/next-jev/runs/59ca1d12a252417d9bc05aaba997cc71
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JonesLin/next-jev-stage1-last-block

Finetuned
(468)
this model