SImi-2B public weights
Public weights-only mirror. No architecture source, no trainer, no private recovery tree.
SImi-2B is a custom Llama-style causal LM trained in JAX/Flax on an AMD MI300X. This repo is not a Transformers checkpoint. AutoModel.from_pretrained will not load it.
Created by Dakuwon Moody.
Status
| Item | Value |
|---|---|
| Public role | Weights only. No source. |
| Student | SImiModel / SImi-2B |
| Parameters | ~2.32B (2,321,856,000 config estimate) |
| Live graph | nn.scan over 30 layers |
| Tokenizer | gpt2 (50,257) |
| Training mix | Full Hugging Face agent-grade streams (no toy slices) |
| Hardware | AMD Instinct MI300X, bf16, remat off, micro-batch 3, seq 1024 |
| Optimizer | Adafactor on the live run (step 99000 on disk is the older Adam snapshot) |
| Save cadence | every 5000 steps under distill/step_<n>/ |
| Latest public tree | distill/step_180000/ |
| Hosted inference | Not supported |
Checkpoints
| Path | What |
|---|---|
distill/step_99000/ |
Canonical unrolled restore. 30 layer_* blocks, Adam/MultiSteps Orbax tree. This is the known-good init. |
distill/step_100000/ .. later |
Scanned Adafactor Orbax trees (scanned_layers). Written by the live MI300X loop. |
Local copies are deleted only after the Orbax blobs exist here.
Restore notes (this ROCm/Orbax stack):
- Step 99000: restore as host numpy, ignore the old Adam opt state, stack
layer_0..layer_29intonn.scan. - Later scanned steps: Orbax
StandardRestorecan fail (Layout/ TensorStore). A metadata-shaped numpy restore of the params works. Fresh Adafactor opt is fine; numpy drops optax NamedTuples. - Do not treat a
metadata.jsonstep number as proof a folder exists.
Architecture
Llama-style decoder-only:
- RMSNorm pre-norm
- RoPE
- GQA
- SwiGLU
- Untied token embedding and LM head
- Causal LM, packed next-token CE
| Property | Value |
|---|---|
| Vocabulary | 50,257 (GPT-2 BPE) |
| Hidden width | 2,560 |
| Depth | 30 layers |
| Query heads | 20 |
| KV heads | 4 |
| Head dimension | 128 |
| SwiGLU intermediate | 6,912 |
| Max sequence | 1,024 |
| RoPE theta | 10,000 |
| RMSNorm eps | 1e-5 |
| Compute / param dtype | bfloat16 |
| Estimated parameters | 2,321,856,000 |
Tokenizer
| Property | Value |
|---|---|
| Tokenizer | gpt2 |
| Vocab size | 50,257 |
| BOS / EOS / PAD | 50256 |
| Packing | 1,024 tokens |
Training mix
Same full HF streams as Uni and Vegeta:
- instruction: Tulu-3, OpenHermes-2.5, UltraChat
- agent / tool: Orca AgentInstruct, Agent-FLAN, xLAM, ToolACE, Glaive, Hermes function calling
- Downloads last month
- 17,504
Model tree for YNSScarSaiyan/simi-weights
Base model
YNSScarSaiyan/simi