CAT-YOKO-12B

YOCO-style causal encoder-decoder MoE, upcycled from openbmb/MiniCPM5-2B-Base (Apache-2.0, Llama GQA).

This card is AvrovaDonz/CAT-YOKO. Training code: AvrovaDonz2026/CAT-YOKO. Operator implementations and evidence: AvrovaDonz/CAT-YOKO-KERNEL.

library_name: transformers only means tokenizer / shard convention follows the Hugging Face ecosystem. The graph is the in-repo cat_yoko PyTorch implementation, not an AutoModelForCausalLM architecture on Hub.

Spec

Name CAT-YOKO-12B
Architecture YOCO-style causal encoder-decoder MoE
Base MiniCPM5-2B (Llama GQA, untied)
(d) 2048
(V) 130560
(L) 42 (16 encoder + 26 decoder)
Attention 16 Q / 2 KV, head_dim=128
FFN SwiGLU intermediate 6144
MoE 1 shared + 20 routed
top-(k) encoder 7 / decoder 10
Stored params 12,250,381,312 (โ‰ˆ12.25B)
Encoder active โ‰ˆ2.03B / input token
Decoder active โ‰ˆ4.33B / output token

Curriculum C1

Stage tokens Encoder Trainable
B0 8B frozen new modules only
B1 27B frozen decoder + lm_head + final RMSNorm
B2 15B trainable all

Wall-clock (published ledger)

Theoretical hours on a 50B-token envelope, not a measured run. On B200 / SM100, allowed linear GEMMs use TeNvfp4Linear (TE NVFP4BlockScaling; B0 frozen encoder is FPROP, WGRAD is B1/B2). Without TE or on sm_120, Nvfp4Linear E2M1/16 emulation. Attn softmax / SDPA stay fp32.

Recipe H100-h Role
C1+NVFP4 571 published wall-clock. B0 student bf16; B1/B2 allowed GEMMs NVFP4
C1+FP8 729 Hopper / Ada fallback
joint bf16 1325 100% baseline

Data and tokenizer

Published recipe Ultra-FineWeb en 55% / zh 30% + UltraData-Math 10% + StarCoder 5%; planned mixture
Current real-text pilot Ultra-FineWeb en 60% / zh 30% + UltraData-Math L2 10%; no code slice
Prepared pilot 79,998,976 training tokens / 999,424 validation tokens, packed at 4096
Tokenizer MiniCPM5-2B-Base tokenizer files; identities recorded in the release manifest

Current B0 snapshot

The latest completed B0 recovery snapshot is step 81864. Adam / absolute cursor: 47062; next corpus row: 8000. B0 remains in its original budget.

Use trainable.pt for optimizer/RNG/cursor recovery or weights-only.pt for the identical overlay without Adam. Both require MiniCPM5-2B-Base upcycling; no full 12B graph is uploaded. Snapshot instructions and verified manifest record base, data, runtime and archived deployed code. Older downloader pins select historical snapshots.

Real packed inputs: 192,765,952, including repeated rows; unique training tokens: 79,998,976.

Historical weights

Path Source Notes
checkpoints/b0-rocm-realtext/step-53307/trainable.pt Historical B0 recovery snapshot Step 53307; overlay plus CPU Adam, RNG and cursor; identities above
checkpoints/b0-rocm-realtext/step-53307/weights-only.pt Same snapshot, smaller overlay Same 132 weights byte-for-byte; no Adam
checkpoints/b0-rocm-realtext/step-52616/ Historical ROCm real-text release Step 52616, cursor/Adam 17814; preceding packed-attention/native CPU Adam backend; both original files and hashes retained
checkpoints/b0-full/trainable.pt Historical Vast B200 DummyStream B0 Step 26940, 130,041,856 cumulative phase tokens; sha256 7eebc9a4da78d79be71bbe52881f2a0eaffd899f58ada3a3325f410eca181955; weights only
checkpoints/b0-3090-bf16/trainable.pt Historical RTX 3090 BF16 DummyStream B0 Step 33800; sha256 2dc31406c240ee8631eb49c22908e41734c6558325b4c19270dd7ab95679e690; weights only
checkpoints/b0/trainable.pt Historical 6000D --try, 32 steps sha256 9012e5ac55c2f59ef7cacc34d5769444413d070116dbff0696c7b258b9aa0636; smoke snapshot
checkpoints/b0-nvfp4-try/trainable.pt Historical 6000D NVFP4 --try, 2 steps sha256 461b4ffc05fd46e2668448393789764ccf9dd673644040fe4527259b176a510e; smoke snapshot
checkpoints/b1/, checkpoints/b2/ Future stages No B1/B2 trained snapshot is included in this release

The historical paths are preserved. The complete 23 GiB graph is not included. Historical logs: B200, 6000D, 3090. ROCm reports: real-text training, operator checks.

Validation observation

At completed step 81864, fixed held-out rows 0โ€“31 have NLL 7.29112438273621 (32 batches, 130,877.0 valid tokens). Rows 32โ€“63 have NLL 7.387959246989646 (32 batches, 130,878.0 valid tokens). These are two slices of the same pilot validation corpus, not external benchmarks.

License

Artifact License
This repo's code and derived weights Apache-2.0
MiniCPM5-2B base Apache-2.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for AvrovaDonz/CAT-YOKO

Finetuned
(5)
this model