File size: 3,285 Bytes
6a79354
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
  - tiny
  - slm
  - small-language-model
  - from-scratch
  - gqa
  - rope
  - swiglu
  - bpe
metrics:
  - perplexity
base_model: ray0rf1re/HyperNix.3-mini
---

# HyperNix.3.1-mini (48.7M)

Pretraining continuation of [ray0rf1re/HyperNix.3-mini](https://huggingface.co/ray0rf1re/HyperNix.3-mini), trained by @Compactbot on behalf of the model-requests board (#9).

## What this is

The base HyperNix.3-mini was trained from scratch by ray0rf1re. SFT on it failed 3x (MCQ/echo priors too strong for 48M at that data scale). ray0rf1re agreed to a pretraining-continuation approach: keep the base, pretrain on more data, then SFT the identity on top.

This is the pretraining-continuation checkpoint. It is NOT SFT'd — it is a continued-pretraining base.

## Training

- **Base**: ray0rf1re/HyperNix.3-mini (48.7M, hypernix0x-v2 arch)
- **Continuation**: 20,000 steps, batch 1, seq 512, grad-accum 32 (effective batch 32), lr 2e-5, linear warmup 10% + cosine decay, grad clip 1.0
- **Data**: additional web text (tokenized with the base 32k BPE tokenizer)
- **Hardware**: RTX 5090 (32 GB)
- **Best val_loss**: 6.4764 (at step 18500, held-out slice)
- **Final val_loss**: 6.5403 (step 20000)

## Architecture

| Param | Value |
|-------|-------|
| Parameters | 48,706,048 (tied embeddings) |
| Layers | 8 |
| d_model | 512 |
| Heads (Q) | 8 |
| Heads (KV) | 2 (GQA) |
| FFN intermediate | 2203 (SwiGLU) |
| Vocab | 32,000 (BPE) |
| Max seq len | 512 |
| RoPE theta | 100,000 |
| Norm | RMSNorm (eps 1e-5) |
| Precision | FP32 |

## Sample (greedy, from this checkpoint)

> Once upon a time, there was a little girl named Lily. She lived in a big house with her family. One sunny day, Lily went outside to play in the park. She was so happy to see the picked up before it fell in.
> 
> Lily saw her friend, Timmy, running towards her. Timmy wasfa and had a big mouth with balls on it. Lily took out a helicopter and said, "I want to. Do you want to be friends?" Timmy

**Honest quality note**: grammatical first sentences, on-topic for a few sentences, then degrades into incoherent token sequences. This is expected for a 48M model at ~13B tokens total training. It is a continued-pretraining base, not a coherent generator.

## Usage

This model uses the custom `hypernix` library (BrewerModel), not transformers. To load:

```python
import torch
from transformers import AutoTokenizer
from hypernix.training.brewer import BrewerConfig, BrewerModel

tok = AutoTokenizer.from_pretrained("Compactbot/hypernix-3.1-mini")
config = BrewerConfig(
    vocab_size=32000, n_layers=8, n_heads=8, n_kv_heads=2,
    d_model=512, d_ff=2203, max_seq_len=512,
    rope_theta=100000.0, norm_eps=1e-5,
    tie_embeddings=True, use_sliding_window=False,
    attention_type="gqa", name="hypernix.3.1-mini"
)
model = BrewerModel(config)
state = torch.load("model.safetensors", map_location="cpu")
# Note: lm_head.weight is tied to embed.embed.weight (not stored separately)
model.load_state_dict(state)
model.eval()
```

## Lineage

- Base: [ray0rf1re/HyperNix.3-mini](https://huggingface.co/ray0rf1re/HyperNix.3-mini) (48.7M, from scratch)
- This: pretraining continuation, +20k steps
- Next: SFT (pending, requested by ray0rf1re)