--- license: apache-2.0 pipeline_tag: text-generation language: en tags: - tiny - slm - small-language-model - from-scratch - gqa - rope - swiglu - bpe metrics: - perplexity base_model: ray0rf1re/HyperNix.3-mini --- # HyperNix.3.1-mini (48.7M) Pretraining continuation of [ray0rf1re/HyperNix.3-mini](https://huggingface.co/ray0rf1re/HyperNix.3-mini), trained by @Compactbot on behalf of the model-requests board (#9). ## What this is The base HyperNix.3-mini was trained from scratch by ray0rf1re. SFT on it failed 3x (MCQ/echo priors too strong for 48M at that data scale). ray0rf1re agreed to a pretraining-continuation approach: keep the base, pretrain on more data, then SFT the identity on top. This is the pretraining-continuation checkpoint. It is NOT SFT'd — it is a continued-pretraining base. ## Training - **Base**: ray0rf1re/HyperNix.3-mini (48.7M, hypernix0x-v2 arch) - **Continuation**: 20,000 steps, batch 1, seq 512, grad-accum 32 (effective batch 32), lr 2e-5, linear warmup 10% + cosine decay, grad clip 1.0 - **Data**: additional web text (tokenized with the base 32k BPE tokenizer) - **Hardware**: RTX 5090 (32 GB) - **Best val_loss**: 6.4764 (at step 18500, held-out slice) - **Final val_loss**: 6.5403 (step 20000) ## Architecture | Param | Value | |-------|-------| | Parameters | 48,706,048 (tied embeddings) | | Layers | 8 | | d_model | 512 | | Heads (Q) | 8 | | Heads (KV) | 2 (GQA) | | FFN intermediate | 2203 (SwiGLU) | | Vocab | 32,000 (BPE) | | Max seq len | 512 | | RoPE theta | 100,000 | | Norm | RMSNorm (eps 1e-5) | | Precision | FP32 | ## Sample (greedy, from this checkpoint) > Once upon a time, there was a little girl named Lily. She lived in a big house with her family. One sunny day, Lily went outside to play in the park. She was so happy to see the picked up before it fell in. > > Lily saw her friend, Timmy, running towards her. Timmy wasfa and had a big mouth with balls on it. Lily took out a helicopter and said, "I want to. Do you want to be friends?" Timmy **Honest quality note**: grammatical first sentences, on-topic for a few sentences, then degrades into incoherent token sequences. This is expected for a 48M model at ~13B tokens total training. It is a continued-pretraining base, not a coherent generator. ## Usage This model uses the custom `hypernix` library (BrewerModel), not transformers. To load: ```python import torch from transformers import AutoTokenizer from hypernix.training.brewer import BrewerConfig, BrewerModel tok = AutoTokenizer.from_pretrained("Compactbot/hypernix-3.1-mini") config = BrewerConfig( vocab_size=32000, n_layers=8, n_heads=8, n_kv_heads=2, d_model=512, d_ff=2203, max_seq_len=512, rope_theta=100000.0, norm_eps=1e-5, tie_embeddings=True, use_sliding_window=False, attention_type="gqa", name="hypernix.3.1-mini" ) model = BrewerModel(config) state = torch.load("model.safetensors", map_location="cpu") # Note: lm_head.weight is tied to embed.embed.weight (not stored separately) model.load_state_dict(state) model.eval() ``` ## Lineage - Base: [ray0rf1re/HyperNix.3-mini](https://huggingface.co/ray0rf1re/HyperNix.3-mini) (48.7M, from scratch) - This: pretraining continuation, +20k steps - Next: SFT (pending, requested by ray0rf1re)