Compactbot commited on
Commit
6a79354
·
verified ·
1 Parent(s): d343f42

Ship HyperNix.3.1-mini: pretraining continuation of HyperNix.3-mini, 20k steps, best val_loss 6.4764

Browse files
Files changed (2) hide show
  1. README.md +90 -0
  2. config.json +17 -0
README.md ADDED
@@ -0,0 +1,90 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ pipeline_tag: text-generation
4
+ language: en
5
+ tags:
6
+ - tiny
7
+ - slm
8
+ - small-language-model
9
+ - from-scratch
10
+ - gqa
11
+ - rope
12
+ - swiglu
13
+ - bpe
14
+ metrics:
15
+ - perplexity
16
+ base_model: ray0rf1re/HyperNix.3-mini
17
+ ---
18
+
19
+ # HyperNix.3.1-mini (48.7M)
20
+
21
+ Pretraining continuation of [ray0rf1re/HyperNix.3-mini](https://huggingface.co/ray0rf1re/HyperNix.3-mini), trained by @Compactbot on behalf of the model-requests board (#9).
22
+
23
+ ## What this is
24
+
25
+ The base HyperNix.3-mini was trained from scratch by ray0rf1re. SFT on it failed 3x (MCQ/echo priors too strong for 48M at that data scale). ray0rf1re agreed to a pretraining-continuation approach: keep the base, pretrain on more data, then SFT the identity on top.
26
+
27
+ This is the pretraining-continuation checkpoint. It is NOT SFT'd — it is a continued-pretraining base.
28
+
29
+ ## Training
30
+
31
+ - **Base**: ray0rf1re/HyperNix.3-mini (48.7M, hypernix0x-v2 arch)
32
+ - **Continuation**: 20,000 steps, batch 1, seq 512, grad-accum 32 (effective batch 32), lr 2e-5, linear warmup 10% + cosine decay, grad clip 1.0
33
+ - **Data**: additional web text (tokenized with the base 32k BPE tokenizer)
34
+ - **Hardware**: RTX 5090 (32 GB)
35
+ - **Best val_loss**: 6.4764 (at step 18500, held-out slice)
36
+ - **Final val_loss**: 6.5403 (step 20000)
37
+
38
+ ## Architecture
39
+
40
+ | Param | Value |
41
+ |-------|-------|
42
+ | Parameters | 48,706,048 (tied embeddings) |
43
+ | Layers | 8 |
44
+ | d_model | 512 |
45
+ | Heads (Q) | 8 |
46
+ | Heads (KV) | 2 (GQA) |
47
+ | FFN intermediate | 2203 (SwiGLU) |
48
+ | Vocab | 32,000 (BPE) |
49
+ | Max seq len | 512 |
50
+ | RoPE theta | 100,000 |
51
+ | Norm | RMSNorm (eps 1e-5) |
52
+ | Precision | FP32 |
53
+
54
+ ## Sample (greedy, from this checkpoint)
55
+
56
+ > Once upon a time, there was a little girl named Lily. She lived in a big house with her family. One sunny day, Lily went outside to play in the park. She was so happy to see the picked up before it fell in.
57
+ >
58
+ > Lily saw her friend, Timmy, running towards her. Timmy wasfa and had a big mouth with balls on it. Lily took out a helicopter and said, "I want to. Do you want to be friends?" Timmy
59
+
60
+ **Honest quality note**: grammatical first sentences, on-topic for a few sentences, then degrades into incoherent token sequences. This is expected for a 48M model at ~13B tokens total training. It is a continued-pretraining base, not a coherent generator.
61
+
62
+ ## Usage
63
+
64
+ This model uses the custom `hypernix` library (BrewerModel), not transformers. To load:
65
+
66
+ ```python
67
+ import torch
68
+ from transformers import AutoTokenizer
69
+ from hypernix.training.brewer import BrewerConfig, BrewerModel
70
+
71
+ tok = AutoTokenizer.from_pretrained("Compactbot/hypernix-3.1-mini")
72
+ config = BrewerConfig(
73
+ vocab_size=32000, n_layers=8, n_heads=8, n_kv_heads=2,
74
+ d_model=512, d_ff=2203, max_seq_len=512,
75
+ rope_theta=100000.0, norm_eps=1e-5,
76
+ tie_embeddings=True, use_sliding_window=False,
77
+ attention_type="gqa", name="hypernix.3.1-mini"
78
+ )
79
+ model = BrewerModel(config)
80
+ state = torch.load("model.safetensors", map_location="cpu")
81
+ # Note: lm_head.weight is tied to embed.embed.weight (not stored separately)
82
+ model.load_state_dict(state)
83
+ model.eval()
84
+ ```
85
+
86
+ ## Lineage
87
+
88
+ - Base: [ray0rf1re/HyperNix.3-mini](https://huggingface.co/ray0rf1re/HyperNix.3-mini) (48.7M, from scratch)
89
+ - This: pretraining continuation, +20k steps
90
+ - Next: SFT (pending, requested by ray0rf1re)
config.json ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "vocab_size": 32000,
3
+ "n_layers": 8,
4
+ "n_heads": 8,
5
+ "n_kv_heads": 2,
6
+ "d_model": 512,
7
+ "d_ff": 2203,
8
+ "max_seq_len": 512,
9
+ "rope_theta": 100000.0,
10
+ "norm_eps": 1e-05,
11
+ "dropout": 0.0,
12
+ "tie_embeddings": true,
13
+ "use_sliding_window": false,
14
+ "sliding_window_size": 512,
15
+ "attention_type": "gqa",
16
+ "name": "hypernix.3.1-mini"
17
+ }