SImi-2B public weights

Public weights-only mirror. No architecture source, no trainer, no private recovery tree.

SImi-2B is a custom Llama-style causal LM trained in JAX/Flax on an AMD MI300X. This repo is not a Transformers checkpoint. AutoModel.from_pretrained will not load it.

Created by Dakuwon Moody.

Status

Item Value
Public role Weights only. No source.
Student SImiModel / SImi-2B
Parameters ~2.32B (2,321,856,000 config estimate)
Live graph nn.scan over 30 layers
Tokenizer gpt2 (50,257)
Training mix Full Hugging Face agent-grade streams (no toy slices)
Hardware AMD Instinct MI300X, bf16, remat off, micro-batch 3, seq 1024
Optimizer Adafactor on the live run (step 99000 on disk is the older Adam snapshot)
Save cadence every 5000 steps under distill/step_<n>/
Latest public tree distill/step_180000/
Hosted inference Not supported

Checkpoints

Path What
distill/step_99000/ Canonical unrolled restore. 30 layer_* blocks, Adam/MultiSteps Orbax tree. This is the known-good init.
distill/step_100000/ .. later Scanned Adafactor Orbax trees (scanned_layers). Written by the live MI300X loop.

Local copies are deleted only after the Orbax blobs exist here.

Restore notes (this ROCm/Orbax stack):

  • Step 99000: restore as host numpy, ignore the old Adam opt state, stack layer_0..layer_29 into nn.scan.
  • Later scanned steps: Orbax StandardRestore can fail (Layout / TensorStore). A metadata-shaped numpy restore of the params works. Fresh Adafactor opt is fine; numpy drops optax NamedTuples.
  • Do not treat a metadata.json step number as proof a folder exists.

Architecture

Llama-style decoder-only:

  • RMSNorm pre-norm
  • RoPE
  • GQA
  • SwiGLU
  • Untied token embedding and LM head
  • Causal LM, packed next-token CE
Property Value
Vocabulary 50,257 (GPT-2 BPE)
Hidden width 2,560
Depth 30 layers
Query heads 20
KV heads 4
Head dimension 128
SwiGLU intermediate 6,912
Max sequence 1,024
RoPE theta 10,000
RMSNorm eps 1e-5
Compute / param dtype bfloat16
Estimated parameters 2,321,856,000

Tokenizer

Property Value
Tokenizer gpt2
Vocab size 50,257
BOS / EOS / PAD 50256
Packing 1,024 tokens

Training mix

Same full HF streams as Uni and Vegeta:

  • instruction: Tulu-3, OpenHermes-2.5, UltraChat
  • agent / tool: Orca AgentInstruct, Agent-FLAN, xLAM, ToolACE, Glaive, Hermes function calling
Downloads last month
17,504
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for YNSScarSaiyan/simi-weights

Finetuned
(1)
this model

Space using YNSScarSaiyan/simi-weights 1