kacperwikiel commited on
Commit
3f431df
·
verified ·
1 Parent(s): fba7868

Release GoLLeM 149M after 20B continuation tokens with full GLINT evaluation

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ loss-curve.png filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright [yyyy] [name of copyright owner]
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
NOTICE ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ GoLLeM-149M-20B
2
+
3
+ Weights and architecture derive from SlayerLab/gollem-v5-ckpts
4
+ revision 963440ce6ab4ada7e95da4c1faa28ebb00c082d4, published under Apache-2.0.
5
+ modeling_gollem.py extracts the inference classes unchanged from the pinned
6
+ GoLLeM r6 training implementation. The continuation, export and release scripts
7
+ were prepared for this 149M run.
8
+
9
+ The likelihood routines in glint_metrics.py are adapted from
10
+ Glint-Research/Glint-1.3/benchmark.py, revision
11
+ c3f99e246aa2c64382f9668dc864533617482d0a, whose repository declares MIT.
12
+ Copyright remains with the original contributors. MIT license text:
13
+
14
+ Permission is hereby granted, free of charge, to any person obtaining a copy
15
+ of this software and associated documentation files (the "Software"), to deal
16
+ in the Software without restriction, including without limitation the rights
17
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
18
+ copies of the Software, and to permit persons to whom the Software is
19
+ furnished to do so, subject to the following conditions:
20
+
21
+ The above copyright notice and this permission notice shall be included in
22
+ all copies or substantial portions of the Software.
23
+
24
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
25
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
26
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
27
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
28
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
29
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
30
+ THE SOFTWARE.
README.md ADDED
@@ -0,0 +1,133 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ library_name: pytorch
6
+ pipeline_tag: text-generation
7
+ base_model: SlayerLab/gollem-v5-ckpts
8
+ tags:
9
+ - gollem
10
+ - tiny-lm
11
+ - muon
12
+ - value-residual
13
+ - continued-pretraining
14
+ - glint-tiny-ml-leaderboard
15
+ ---
16
+
17
+ # GoLLeM 149M — 20B-token continuation
18
+
19
+ A **148,910,738-parameter English causal language model**, grown from the GoLLeM-v5
20
+ 122.8M final checkpoint and continued for **20,000,014,336 additional tokens**.
21
+ This is a base completion model, not an instruction-tuned assistant.
22
+
23
+ ## Evaluation
24
+
25
+ Full final-checkpoint evaluation, with no best-checkpoint selection:
26
+
27
+ | Metric | Result | Coverage |
28
+ |---|---:|---|
29
+ | BLiMP accuracy | **79.47%** | 67,000 pairs / 67 tasks |
30
+ | ARC-Easy accuracy | **54.97%** | 2,376 test examples |
31
+ | WikiText-2 byte perplexity | **2.212573** | Full test text |
32
+ | WikiText-2 token perplexity | 21.452430 | Same evaluation |
33
+ | GLINT overall score, fixed board snapshot | **77.11235** | Calculated |
34
+ | GLINT efficiency score, fixed board snapshot | **77.13585** | Calculated |
35
+
36
+ The likelihood protocol follows `Glint-Research/Glint-1.3/benchmark.py` at
37
+ `c3f99e246aa2c64382f9668dc864533617482d0a`: first-256-token truncation for BLiMP
38
+ and ARC, summed unnormalized log probabilities, zero-shot ARC scoring, and
39
+ non-overlapping 256-token WikiText contexts. Inference uses BF16 autocast with
40
+ FP32 log probabilities. Byte perplexity is converted from token perplexity using
41
+ the measured token/UTF-8-byte ratio. Dataset revisions are pinned in
42
+ [evaluation.json](evaluation.json). Loader order and complete task counts are checked.
43
+
44
+ Against the [GLINT board](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard)
45
+ at revision `2c5ea9682babfe2c410ebbbf78113cc98f08a2f9`, this would place **4th in
46
+ overall score** and **8th in size-adjusted efficiency** among the existing entries
47
+ plus this model. This is a calculated, self-reported placement, not an accepted
48
+ leaderboard entry. Competitor scores were not independently re-evaluated here;
49
+ some competitors use lm-eval-harness and different context/scoring conventions.
50
+ The ranking and normalized score can change as the board changes. This release
51
+ does not establish a leaderboard win.
52
+
53
+ ## Use
54
+
55
+ This release uses a small custom PyTorch architecture. Use the included loader;
56
+ it is not packaged for `transformers.AutoModelForCausalLM`.
57
+
58
+ ```bash
59
+ pip install huggingface-hub
60
+ hf download SlayerLab/gollem-149m-20b --local-dir gollem-149m-20b
61
+ cd gollem-149m-20b
62
+ pip install -r requirements.txt
63
+ python generate.py --prompt "The scientific method is" --max-new-tokens 64
64
+ ```
65
+
66
+ A matching
67
+ CUDA-enabled PyTorch installation is needed for GPU inference/evaluation.
68
+
69
+ ```python
70
+ from load_model import load_model
71
+ model, tokenizer = load_model('.', device='cuda')
72
+ # model(input_ids) returns (logits, optional_loss)
73
+ ```
74
+
75
+ Weights remain FP32 in `model.safetensors`; tied embeddings are restored by
76
+ `safetensors.torch.load_model`. Export verification checks every model tensor
77
+ and compares logits against the original checkpoint at lengths 16, 256 and 1024.
78
+ The generation CLI uses BF16 autocast on CUDA, FP32 on CPU, and greedy decoding
79
+ by default. It recomputes context rather than using a KV cache.
80
+
81
+ ## Architecture and training
82
+
83
+ - 19 transformer layers; hidden size 768; 12 attention heads; FFN size 2160.
84
+ - Vocabulary 12,288; tied input/output embeddings; context length 1024.
85
+ - RMSNorm, RoPE (theta 100,000), QK normalization, SwiGLU and value residuals.
86
+ - Source: `SlayerLab/gollem-v5-ckpts`, revision
87
+ `963440ce6ab4ada7e95da4c1faa28ebb00c082d4`, file
88
+ `final_128m_16x768/ckpt_760k.pt` (24,903,680,000 source training tokens).
89
+ - FFNs were widened with zero new down-projection columns; three residual blocks
90
+ were appended with zero output projections. Initial FP32 logits exactly matched
91
+ the donor on the tested contexts. Optimizer state was reset for continuation.
92
+ - Continuation: 20,000,014,336 tokens, 39,147 actual optimizer updates, seed 1337.
93
+ Inherited token lineage totals 44,903,694,336; newly added parameters only saw
94
+ the 20B-token continuation.
95
+ - Cleaned ARC-MIX pool: 9,391,706,576 tokens, reconstructed from the published
96
+ SlayerLab corpus/removal list. SHA256:
97
+ `ecfd0a4040a0ece6f07f728f092a1b373b6d253b4adfbe259c5107c1fc16681d`.
98
+ Repeated shuffled passes; training corpus is not redistributed here. This work
99
+ did not independently audit all training/evaluation overlap.
100
+ - Muon for hidden matrices and AdamW for auxiliary parameters; cosine learning
101
+ rate 2e-4 to 2e-5, 250 reference-unit warmup; Muon LR ratio 33.3333.
102
+ - Eight H100 PCIe GPUs. Effective batch started at 256 sequences, then became
103
+ 512 after 524,288,000 continuation tokens. Learning rate remained indexed by
104
+ token position. Grouped Muon, BF16 gradient communication and later CUDA graphs
105
+ improved throughput. Numerical ordering changed; this is not a bitwise replay
106
+ of the original trainer.
107
+ - Optimized sustained full-segment throughput: 1.045M tokens/sec including
108
+ startup, validation and saves; later ordinary windows reached about 1.14M.
109
+ These are measurements on this host, not expected performance on every H100 node.
110
+
111
+ `training_manifest.json` records configuration and hashes. Its `reference_step`
112
+ uses 262,144 tokens/unit; it is not the actual optimizer-update count after the
113
+ batch change. `training_metrics.jsonl` records the training history.
114
+
115
+ ## Reproduce evaluation
116
+
117
+ ```bash
118
+ python evaluate_glint.py --model-dir . --out reproduced-results.json
119
+ ```
120
+
121
+ This evaluates all three datasets on CUDA using pinned dataset revisions. It can
122
+ take tens of minutes. `board_snapshot.json` pins the normalization constants.
123
+ See `release_verification.json` for export and official-function parity checks.
124
+
125
+ ## Limitations and license
126
+
127
+ English base model intended for research. It can generate incorrect, repetitive,
128
+ biased or inappropriate text and is not optimized for instruction following.
129
+ Small changes in precision, context length, tokenizer handling or benchmark
130
+ harness can change scores. Overall score and size-adjusted efficiency are distinct.
131
+
132
+ Weights and GoLLeM implementation: Apache-2.0, following the source model's
133
+ published license. See `LICENSE` and `NOTICE` for source and evaluation attribution.
board_snapshot.json ADDED
@@ -0,0 +1,1263 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "paramLogMin": 2.9822712330395684,
3
+ "paramLogMax": 8.176091259055681,
4
+ "wikiMinLog": 0.6205764877251099,
5
+ "wikiMaxLog": 6.214608098422191,
6
+ "ranking": [
7
+ {
8
+ "name": "JugnuLM-110M-R2+",
9
+ "org": "altslate",
10
+ "params": "110M",
11
+ "blimp": 82.52,
12
+ "arc": 55.13,
13
+ "wiki": 1.8735,
14
+ "tokens": "25B",
15
+ "releaseDate": "2026-09-19",
16
+ "links": {
17
+ "card": "https://huggingface.co/altslate/JugnuLM-110M-R2plus"
18
+ },
19
+ "score": 79.1735740039279,
20
+ "efficiency": 80.20023332743875
21
+ },
22
+ {
23
+ "name": "GPT-X2-125M",
24
+ "org": "axiomiclabs",
25
+ "params": "125M",
26
+ "blimp": 81.28,
27
+ "arc": 57.07,
28
+ "aci": 54.53,
29
+ "wiki": 1.86,
30
+ "tokens": "75B",
31
+ "releaseDate": "2026-04-22",
32
+ "links": {
33
+ "card": "https://huggingface.co/AxiomicLabs/GPT-X2-125M"
34
+ },
35
+ "score": 79.45,
36
+ "efficiency": 80.05561878992458
37
+ },
38
+ {
39
+ "name": "Haidass-143M-v1",
40
+ "org": "DALab",
41
+ "params": "143M",
42
+ "blimp": 79.05,
43
+ "arc": 60.23,
44
+ "wiki": 1.8892,
45
+ "tokens": "100B",
46
+ "releaseDate": "2026-08-11",
47
+ "links": {
48
+ "card": "https://huggingface.co/DALabCommunity/Haidass-143M-v1"
49
+ },
50
+ "score": 79.66718100767177,
51
+ "efficiency": 79.8263615325101
52
+ },
53
+ {
54
+ "name": "JugnuLM-53M",
55
+ "org": "altslate",
56
+ "params": "53M",
57
+ "blimp": 78.14,
58
+ "arc": 51.43,
59
+ "wiki": 2.04,
60
+ "tokens": "12B",
61
+ "releaseDate": "2026-09-05",
62
+ "links": {
63
+ "card": "https://huggingface.co/altslate/JugnuLM-53M"
64
+ },
65
+ "score": 75.97290550501255,
66
+ "efficiency": 79.27738349200743
67
+ },
68
+ {
69
+ "name": "Glint-2",
70
+ "org": "glintresearch",
71
+ "params": "1.06M",
72
+ "blimp": 66.36,
73
+ "arc": 36.8,
74
+ "aci": 48.02,
75
+ "wiki": 3.09,
76
+ "tokens": "~300B",
77
+ "releaseDate": "2026-07-19",
78
+ "links": {
79
+ "card": "https://huggingface.co/Glint-Research/Glint-2"
80
+ },
81
+ "score": 64.6953799614223,
82
+ "efficiency": 78.09071110206555
83
+ },
84
+ {
85
+ "name": "Supra-50M-Instruct",
86
+ "org": "supralabs",
87
+ "params": "51.8M",
88
+ "blimp": 76.3,
89
+ "arc": 52.2,
90
+ "aci": 53.28,
91
+ "wiki": 2.56,
92
+ "tokens": "20B",
93
+ "releaseDate": "2026-05-21",
94
+ "links": {
95
+ "card": "https://huggingface.co/SupraLabs/Supra-50M-Instruct",
96
+ "base": "https://huggingface.co/SupraLabs/Supra-50M-Base"
97
+ },
98
+ "score": 74.26326441586103,
99
+ "efficiency": 77.5644874221056
100
+ },
101
+ {
102
+ "name": "Glint-1.3 (merged)",
103
+ "org": "glintresearch",
104
+ "params": "982K",
105
+ "blimp": 68.7,
106
+ "arc": 32.5,
107
+ "aci": 52.73,
108
+ "wiki": 3.08,
109
+ "tokens": "100B",
110
+ "releaseDate": "2026-05-13",
111
+ "links": {
112
+ "card": "https://huggingface.co/Glint-Research/Glint-1.3"
113
+ },
114
+ "score": 64.06136182059981,
115
+ "efficiency": 77.5301302449298
116
+ },
117
+ {
118
+ "name": "Ivme-Conversate-v3-Base",
119
+ "org": "ivmelabs",
120
+ "params": "24.79M",
121
+ "blimp": 78.49,
122
+ "arc": 39.02,
123
+ "wiki": 2.1362,
124
+ "tokens": "~15B",
125
+ "releaseDate": "2026-09-13",
126
+ "links": {
127
+ "card": "https://huggingface.co/IvmeLabs/Ivme-Conversate-v3-Base"
128
+ },
129
+ "score": 71.67833464673554,
130
+ "efficiency": 77.07312862611622
131
+ },
132
+ {
133
+ "name": "GoLLeM-v5 128M",
134
+ "org": "Fabryka AI",
135
+ "params": "122.8M",
136
+ "blimp": 79.09,
137
+ "arc": 53.24,
138
+ "wiki": 2.2538,
139
+ "tokens": "24.9B",
140
+ "releaseDate": "2026-09-28",
141
+ "links": {
142
+ "card": "https://huggingface.co/SlayerLab/gollem-v5-ckpts"
143
+ },
144
+ "score": 76.29901139538448,
145
+ "efficiency": 76.93725470597208
146
+ },
147
+ {
148
+ "name": "GPT-S-5M",
149
+ "org": "axiomiclabs",
150
+ "params": "5.16M",
151
+ "blimp": 72.27,
152
+ "arc": 35.69,
153
+ "aci": 53.07,
154
+ "wiki": 2.57,
155
+ "tokens": "25B",
156
+ "releaseDate": "2026-05-19",
157
+ "links": {
158
+ "card": "https://huggingface.co/AxiomicLabs/GPT-S-5M"
159
+ },
160
+ "score": 67.3933667970715,
161
+ "efficiency": 76.88794431148578
162
+ },
163
+ {
164
+ "name": "Supra-50M-Base",
165
+ "org": "supralabs",
166
+ "params": "51.8M",
167
+ "blimp": 76.3,
168
+ "arc": 46,
169
+ "aci": 53.53,
170
+ "wiki": 2.04,
171
+ "tokens": "20B",
172
+ "releaseDate": "2026-05-27",
173
+ "links": {
174
+ "card": "https://huggingface.co/SupraLabs/Supra-50M-Base"
175
+ },
176
+ "score": 73.5495721716792,
177
+ "efficiency": 76.81906943472619
178
+ },
179
+ {
180
+ "name": "MicroSupra-1k",
181
+ "org": "supralabs",
182
+ "params": "1K",
183
+ "blimp": 58.61,
184
+ "arc": 26.39,
185
+ "aci": 0.57,
186
+ "wiki": 11.27,
187
+ "tokens": "—",
188
+ "releaseDate": "2026-05-13",
189
+ "links": {
190
+ "card": "https://huggingface.co/SupraLabs/MicroSupra-1k"
191
+ },
192
+ "score": 50.93160731709393,
193
+ "efficiency": 76.31048511062035
194
+ },
195
+ {
196
+ "name": "Hydrion-v1-Base",
197
+ "org": "opengcm",
198
+ "params": "114.1M",
199
+ "blimp": 80.08,
200
+ "arc": 47.26,
201
+ "wiki": 2.04,
202
+ "tokens": "~2.5B",
203
+ "links": {},
204
+ "score": 75.22957217167921,
205
+ "efficiency": 76.0899885430591
206
+ },
207
+ {
208
+ "name": "GoLLeM-v5 64M",
209
+ "org": "Fabryka AI",
210
+ "params": "62.9M",
211
+ "blimp": 76.16,
212
+ "arc": 48.19,
213
+ "wiki": 2.3474,
214
+ "tokens": "24.9B",
215
+ "releaseDate": "2026-09-27",
216
+ "links": {
217
+ "card": "https://huggingface.co/SlayerLab/gollem-v5-ckpts"
218
+ },
219
+ "score": 73.39654671767353,
220
+ "efficiency": 76.06345060452743
221
+ },
222
+ {
223
+ "name": "Ivme-Conversate-v2-Base",
224
+ "org": "ivmelabs",
225
+ "params": "23.85M",
226
+ "blimp": 75.09,
227
+ "arc": 39.98,
228
+ "aci": 53.72,
229
+ "wiki": 2.225,
230
+ "tokens": "~12.85B",
231
+ "releaseDate": "2026-07-07",
232
+ "links": {
233
+ "card": "https://huggingface.co/IvmeLabs/Ivme-Conversate-v2-Base",
234
+ "demo": "https://huggingface.co/spaces/IvmeLabs/Ivme-Conversate-Demo"
235
+ },
236
+ "score": 70.62231190929478,
237
+ "efficiency": 76.05176278508509
238
+ },
239
+ {
240
+ "name": "GoLLeM-v5 64M (v1 recipe)",
241
+ "org": "Fabryka AI",
242
+ "params": "62.9M",
243
+ "blimp": 75.83,
244
+ "arc": 47.94,
245
+ "wiki": 2.3718,
246
+ "tokens": "13.1B",
247
+ "releaseDate": "2026-09-24",
248
+ "links": {
249
+ "card": "https://huggingface.co/SlayerLab/gollem-v5-ckpts"
250
+ },
251
+ "score": 73.14159516582735,
252
+ "efficiency": 75.79923524784323
253
+ },
254
+ {
255
+ "name": "Pollock 1.5 Mini LM 128M",
256
+ "org": "Fabryka AI",
257
+ "params": "127.43M",
258
+ "blimp": 78.497,
259
+ "arc": 47.7273,
260
+ "wiki": 1.9435,
261
+ "tokens": "21.50B",
262
+ "releaseDate": "2026-09-20",
263
+ "links": {
264
+ "card": "https://huggingface.co/SlayerLab/pollock-mini-lm-125m"
265
+ },
266
+ "score": 75.14642835516423,
267
+ "efficiency": 75.65875244066669
268
+ },
269
+ {
270
+ "name": "Hummingbird-V2",
271
+ "org": "juinron",
272
+ "params": "9.59M",
273
+ "blimp": 72.9,
274
+ "arc": 38.85,
275
+ "wiki": 2.9551,
276
+ "tokens": "10B",
277
+ "releaseDate": "2026-09-23",
278
+ "links": {
279
+ "card": "https://huggingface.co/juinron/Hummingbird-V2"
280
+ },
281
+ "score": 67.82470273244657,
282
+ "efficiency": 75.62254586041912
283
+ },
284
+ {
285
+ "name": "Michel-Nano-v2",
286
+ "org": "finnianx",
287
+ "params": "9.94M",
288
+ "blimp": 72.52,
289
+ "arc": 35.9,
290
+ "aci": 54.25,
291
+ "wiki": 2.46,
292
+ "tokens": "6.5B",
293
+ "releaseDate": "2026-06-13",
294
+ "links": {
295
+ "card": "https://huggingface.co/finnianx/michel-nano-v2"
296
+ },
297
+ "score": 67.80736215979108,
298
+ "efficiency": 75.50158990691395
299
+ },
300
+ {
301
+ "name": "min-spark",
302
+ "org": "minimalabs",
303
+ "params": "5.76M",
304
+ "blimp": 69.19,
305
+ "arc": 37.08,
306
+ "wiki": 2.7747,
307
+ "tokens": "10B",
308
+ "releaseDate": "2026-08-06",
309
+ "links": {
310
+ "card": "https://huggingface.co/MinimaLabs/min-spark"
311
+ },
312
+ "score": 66.37337572758646,
313
+ "efficiency": 75.41900255718414
314
+ },
315
+ {
316
+ "name": "GoLLeM-v5 32M",
317
+ "org": "Fabryka AI",
318
+ "params": "31.6M",
319
+ "blimp": 73.48,
320
+ "arc": 44.44,
321
+ "wiki": 2.5386,
322
+ "tokens": "16B",
323
+ "releaseDate": "2026-09-23",
324
+ "links": {
325
+ "card": "https://huggingface.co/SlayerLab/gollem-v5-ckpts"
326
+ },
327
+ "score": 70.78661838489879,
328
+ "efficiency": 75.39597759176215
329
+ },
330
+ {
331
+ "name": "TextModel-v1",
332
+ "org": "benchlabs",
333
+ "params": "122.7M",
334
+ "blimp": 80.31,
335
+ "arc": 49.71,
336
+ "aci": 55.044,
337
+ "wiki": 2.628,
338
+ "tokens": "1.64B",
339
+ "releaseDate": "2026-07-29",
340
+ "links": {
341
+ "card": "https://huggingface.co/TobiasLogic/TextModel-v1"
342
+ },
343
+ "score": 74.61371791371184,
344
+ "efficiency": 75.2404050492078
345
+ },
346
+ {
347
+ "name": "Gros-Michel-90m-Base-v2",
348
+ "org": "finnianx",
349
+ "params": "95M",
350
+ "blimp": 80.2,
351
+ "arc": 43.18,
352
+ "aci": 53.9,
353
+ "wiki": 2.08,
354
+ "tokens": "9B",
355
+ "releaseDate": "2026-07-06",
356
+ "links": {
357
+ "card": "https://huggingface.co/finnianx/Gros-Michel-90m-Base-v2"
358
+ },
359
+ "score": 73.79386500847113,
360
+ "efficiency": 75.20307015909175
361
+ },
362
+ {
363
+ "name": "ObsidianSmall-Base",
364
+ "org": "dreamw",
365
+ "params": "8.7M",
366
+ "blimp": 71.891,
367
+ "arc": 33.291,
368
+ "wiki": 2.467,
369
+ "tokens": "4B",
370
+ "releaseDate": "2026-07-31",
371
+ "links": {
372
+ "card": "https://huggingface.co/Dream-W/ObsidianSmall-Base"
373
+ },
374
+ "score": 66.71109716428224,
375
+ "efficiency": 74.65256171819846
376
+ },
377
+ {
378
+ "name": "GoLLeM-v5 16M",
379
+ "org": "Fabryka AI",
380
+ "params": "17.4M",
381
+ "blimp": 70.08,
382
+ "arc": 40.91,
383
+ "wiki": 2.6746,
384
+ "tokens": "16B",
385
+ "releaseDate": "2026-09-23",
386
+ "links": {
387
+ "card": "https://huggingface.co/SlayerLab/gollem-v5-ckpts"
388
+ },
389
+ "score": 68.16564952787132,
390
+ "efficiency": 74.30485232133755
391
+ },
392
+ {
393
+ "name": "Gros-Michel-90m-Base",
394
+ "org": "finnianx",
395
+ "params": "91.1M",
396
+ "blimp": 78.35,
397
+ "arc": 41.5,
398
+ "aci": 54.51,
399
+ "wiki": 2.07,
400
+ "tokens": "6.5B",
401
+ "releaseDate": "2026-06-28",
402
+ "links": {
403
+ "card": "https://huggingface.co/finnianx/Gros-Michel-90m-Base"
404
+ },
405
+ "score": 72.64591517652724,
406
+ "efficiency": 74.16051667816828
407
+ },
408
+ {
409
+ "name": "Supra-Mini-v6",
410
+ "org": "supralabs",
411
+ "params": "1.41M",
412
+ "blimp": 61.86,
413
+ "arc": 30.26,
414
+ "aci": 51.04,
415
+ "wiki": 3,
416
+ "tokens": "—",
417
+ "releaseDate": "2026-05-30",
418
+ "links": {
419
+ "card": "https://huggingface.co/SupraLabs/Supra-Mini-v6-1M"
420
+ },
421
+ "score": 61.19151293252804,
422
+ "efficiency": 73.13141194113591
423
+ },
424
+ {
425
+ "name": "CreekwardGoat-500K",
426
+ "org": "wonderfulmonkey",
427
+ "params": "500k",
428
+ "blimp": 60.61,
429
+ "arc": 30.35,
430
+ "aci": 49.58,
431
+ "wiki": 4.1228,
432
+ "tokens": "300M",
433
+ "releaseDate": "2026-07-16",
434
+ "links": {
435
+ "card": "https://huggingface.co/wonderfulmonkey/CreekwardGoat-500K"
436
+ },
437
+ "score": 58.91044477014074,
438
+ "efficiency": 72.95870925891221
439
+ },
440
+ {
441
+ "name": "Michel-Micro",
442
+ "org": "finnianx",
443
+ "params": "28.4M",
444
+ "blimp": 69.75,
445
+ "arc": 38.59,
446
+ "aci": 53.5,
447
+ "wiki": 2.3,
448
+ "tokens": "2.6B",
449
+ "releaseDate": "2026-06-09",
450
+ "links": {
451
+ "card": "https://huggingface.co/finnianx/michel-micro"
452
+ },
453
+ "score": 68.18143346822255,
454
+ "efficiency": 72.92550367504299
455
+ },
456
+ {
457
+ "name": "KeyLM-75M-Instruct",
458
+ "org": "minimalabs",
459
+ "params": "75M",
460
+ "blimp": 76.02,
461
+ "arc": 39.1,
462
+ "aci": 52.26,
463
+ "wiki": 2.17,
464
+ "tokens": "~18B",
465
+ "releaseDate": "2026-05-29",
466
+ "links": {
467
+ "card": "https://huggingface.co/MinimaLabs/KeyLM-75M-Instruct",
468
+ "base": "https://huggingface.co/MinimaLabs/KeyLM-75M"
469
+ },
470
+ "score": 70.78812412850577,
471
+ "efficiency": 72.83953798119279
472
+ },
473
+ {
474
+ "name": "Archaea-74M",
475
+ "org": "GODELEV",
476
+ "params": "74M",
477
+ "blimp": 74.91,
478
+ "arc": 39.06,
479
+ "aci": 52.58,
480
+ "wiki": 2.2,
481
+ "tokens": "~1.2B",
482
+ "releaseDate": "2026-06-01",
483
+ "links": {
484
+ "card": "https://huggingface.co/GODELEV/Archaea-74M"
485
+ },
486
+ "score": 70.32297626040025,
487
+ "efficiency": 72.40037555332115
488
+ },
489
+ {
490
+ "name": "Michel-Tiny",
491
+ "org": "finnianx",
492
+ "params": "55.7M",
493
+ "blimp": 74.08,
494
+ "arc": 37.33,
495
+ "aci": 52.9,
496
+ "wiki": 2.29,
497
+ "tokens": "1.3B",
498
+ "releaseDate": "2026-06-05",
499
+ "links": {
500
+ "card": "https://huggingface.co/finnianx/michel-tiny"
501
+ },
502
+ "score": 69.23073081506284,
503
+ "efficiency": 72.09813447714069
504
+ },
505
+ {
506
+ "name": "KeyLM-75M",
507
+ "org": "minimalabs",
508
+ "params": "75M",
509
+ "blimp": 76.1,
510
+ "arc": 35.65,
511
+ "aci": 52.27,
512
+ "wiki": 2.08,
513
+ "tokens": "~18B",
514
+ "releaseDate": "2026-05-29",
515
+ "links": {
516
+ "card": "https://huggingface.co/MinimaLabs/KeyLM-75M"
517
+ },
518
+ "score": 69.91719834180446,
519
+ "efficiency": 71.94337308488805
520
+ },
521
+ {
522
+ "name": "Supra-1.5-Instruct-exp",
523
+ "org": "supralabs",
524
+ "params": "51.8M",
525
+ "blimp": 67.4,
526
+ "arc": 45.9,
527
+ "aci": 53.17,
528
+ "wiki": 2.7,
529
+ "tokens": "23B",
530
+ "releaseDate": "2026-06-12",
531
+ "links": {
532
+ "card": "https://huggingface.co/SupraLabs/Supra-1.5-50M-Instruct-exp",
533
+ "demo": "https://huggingface.co/spaces/SupraLabs/Supra1.5-50M-Instruct-Demo"
534
+ },
535
+ "score": 68.87932797416606,
536
+ "efficiency": 71.94121899055959
537
+ },
538
+ {
539
+ "name": "Glint-1",
540
+ "org": "glintresearch",
541
+ "params": "1M",
542
+ "blimp": 61.2,
543
+ "arc": 32,
544
+ "wiki": 4.45,
545
+ "tokens": "100B",
546
+ "releaseDate": "2026-05-02",
547
+ "links": {
548
+ "card": "https://huggingface.co/Glint-Research/Glint-1"
549
+ },
550
+ "score": 59.20203385107235,
551
+ "efficiency": 71.60417983768008
552
+ },
553
+ {
554
+ "name": "Supra-Mini-v5",
555
+ "org": "supralabs",
556
+ "params": "7.87M",
557
+ "blimp": 63.5,
558
+ "arc": 34.4,
559
+ "aci": 52.56,
560
+ "wiki": 2.73,
561
+ "tokens": "—",
562
+ "releaseDate": "2026-05-16",
563
+ "links": {
564
+ "card": "https://huggingface.co/SupraLabs/Supra-Mini-v5-8M"
565
+ },
566
+ "score": 63.68015163197632,
567
+ "efficiency": 71.52774878766438
568
+ },
569
+ {
570
+ "name": "Michel-Nano",
571
+ "org": "finnianx",
572
+ "params": "5.96M",
573
+ "blimp": 65.23,
574
+ "arc": 33.38,
575
+ "aci": 54.25,
576
+ "wiki": 3.25,
577
+ "tokens": "1.1B",
578
+ "releaseDate": "2026-06-10",
579
+ "links": {
580
+ "card": "https://huggingface.co/finnianx/michel-nano"
581
+ },
582
+ "score": 62.87789324852722,
583
+ "efficiency": 71.35741139542888
584
+ },
585
+ {
586
+ "name": "Supra-Mini-v4",
587
+ "org": "supralabs",
588
+ "params": "2.62M",
589
+ "blimp": 60.7,
590
+ "arc": 31.5,
591
+ "aci": 50.42,
592
+ "wiki": 3.17,
593
+ "tokens": "—",
594
+ "releaseDate": "2026-05-14",
595
+ "links": {
596
+ "card": "https://huggingface.co/SupraLabs/Supra-Mini-v4-2M"
597
+ },
598
+ "score": 60.88973848518265,
599
+ "efficiency": 71.1934620365565
600
+ },
601
+ {
602
+ "name": "Byrne-86M-Base",
603
+ "org": "quazim0t0",
604
+ "params": "86M",
605
+ "blimp": 73.56,
606
+ "arc": 39.31,
607
+ "aci": 52.42,
608
+ "wiki": 2.3753,
609
+ "tokens": "—",
610
+ "releaseDate": "2026-06-18",
611
+ "links": {
612
+ "card": "https://huggingface.co/Quazim0t0/Byrne-86M-Base",
613
+ "base": "https://huggingface.co/Quazim0t0/Byrne-86M"
614
+ },
615
+ "score": 69.4994751776389,
616
+ "efficiency": 71.11587440455948
617
+ },
618
+ {
619
+ "name": "Supra-50M-Reasoning",
620
+ "org": "supralabs",
621
+ "params": "51.8M",
622
+ "blimp": 64.14,
623
+ "arc": 45.16,
624
+ "aci": 53.16,
625
+ "wiki": 2.6,
626
+ "tokens": "20B",
627
+ "releaseDate": "2026-06-04",
628
+ "links": {
629
+ "card": "https://huggingface.co/SupraLabs/Supra-50M-Reasoning",
630
+ "demo": "https://huggingface.co/spaces/SupraLabs/Supra-50M-Reasoning-Demo"
631
+ },
632
+ "score": 67.77087912849917,
633
+ "efficiency": 70.78349629651903
634
+ },
635
+ {
636
+ "name": "Escarda-86M-Base",
637
+ "org": "quazim0t0",
638
+ "params": "85.7M",
639
+ "blimp": 71.44,
640
+ "arc": 38.01,
641
+ "aci": 52.51,
642
+ "wiki": 2.2228,
643
+ "tokens": "~20B",
644
+ "releaseDate": "2026-05-22",
645
+ "links": {
646
+ "card": "https://huggingface.co/Quazim0t0/Escarda-86M-Base",
647
+ "discussion": "https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard/discussions/21"
648
+ },
649
+ "score": 68.75487327030368,
650
+ "efficiency": 70.36399981030752
651
+ },
652
+ {
653
+ "name": "Dumb 1.2",
654
+ "org": "56m",
655
+ "params": "34.6M",
656
+ "blimp": 70.51,
657
+ "arc": 34.97,
658
+ "wiki": 2.7195,
659
+ "tokens": "1.1B",
660
+ "links": {},
661
+ "score": 66.22978068417677,
662
+ "efficiency": 70.29127820867504
663
+ },
664
+ {
665
+ "name": "PotentSulfurLM 500K",
666
+ "org": "mihaipopa",
667
+ "params": "587K",
668
+ "blimp": 59.01,
669
+ "arc": 27.06,
670
+ "aci": 51.4,
671
+ "wiki": 4.52,
672
+ "tokens": "~200M",
673
+ "releaseDate": "2026-05-27",
674
+ "links": {
675
+ "card": "https://huggingface.co/MihaiPopa-1/PotentSulfurLM-500K-Base"
676
+ },
677
+ "score": 56.7323639102486,
678
+ "efficiency": 69.88073138943594
679
+ },
680
+ {
681
+ "name": "Glint-0.4",
682
+ "org": "glintresearch",
683
+ "params": "1M",
684
+ "blimp": 58.5,
685
+ "arc": 31,
686
+ "wiki": 5.01,
687
+ "tokens": "10B",
688
+ "releaseDate": "2026-04-19",
689
+ "links": {
690
+ "card": "https://huggingface.co/Glint-Research/Glint-0.4"
691
+ },
692
+ "score": 57.26240121419818,
693
+ "efficiency": 69.25821644562583
694
+ },
695
+ {
696
+ "name": "kirk-tung",
697
+ "org": "rtc",
698
+ "params": "53.1M",
699
+ "blimp": 72.81,
700
+ "arc": 30.3,
701
+ "aci": 52.77,
702
+ "wiki": 2.38,
703
+ "tokens": "1.1B",
704
+ "releaseDate": "2026-06-09",
705
+ "links": {
706
+ "card": "https://huggingface.co/rtc2022/kirk-tung"
707
+ },
708
+ "score": 66.23436296685166,
709
+ "efficiency": 69.11003843213756
710
+ },
711
+ {
712
+ "name": "Supra-Mini-v3",
713
+ "org": "supralabs",
714
+ "params": "468K",
715
+ "blimp": 55.3,
716
+ "arc": 27.3,
717
+ "aci": 48.96,
718
+ "wiki": 4.49,
719
+ "tokens": "—",
720
+ "releaseDate": "2026-05-14",
721
+ "links": {
722
+ "card": "https://huggingface.co/SupraLabs/Supra-Mini-v3-0.5M"
723
+ },
724
+ "score": 55.61537817827218,
725
+ "efficiency": 69.03166322587406
726
+ },
727
+ {
728
+ "name": "CinnabarLM 1.5M",
729
+ "org": "mihaipopa",
730
+ "params": "1.71M",
731
+ "blimp": 60.51,
732
+ "arc": 26.68,
733
+ "aci": 51.28,
734
+ "wiki": 4.23,
735
+ "tokens": "~50M",
736
+ "releaseDate": "2026-05-19",
737
+ "links": {
738
+ "card": "https://huggingface.co/MihaiPopa-1/CinnabarLM-1.5M-Base"
739
+ },
740
+ "score": 57.50082074544151,
741
+ "efficiency": 68.25683128055626
742
+ },
743
+ {
744
+ "name": "CinnabarLM 1.4M",
745
+ "org": "mihaipopa",
746
+ "params": "1.51M",
747
+ "blimp": 60.7,
748
+ "arc": 24.58,
749
+ "aci": 51.57,
750
+ "wiki": 4.09,
751
+ "tokens": "~30M",
752
+ "releaseDate": "2026-05-19",
753
+ "links": {
754
+ "card": "https://huggingface.co/MihaiPopa-1/CinnabarLM-1.4M-Base"
755
+ },
756
+ "score": 57.06470724773393,
757
+ "efficiency": 68.03589444426096
758
+ },
759
+ {
760
+ "name": "CinnabarLM 4M",
761
+ "org": "mihaipopa",
762
+ "params": "4.23M",
763
+ "blimp": 62.87,
764
+ "arc": 27.36,
765
+ "aci": 50.99,
766
+ "wiki": 3.77,
767
+ "tokens": "~80M",
768
+ "releaseDate": "2026-05-05",
769
+ "links": {
770
+ "card": "https://huggingface.co/MihaiPopa-1/CinnabarLM-4M-Base"
771
+ },
772
+ "score": 59.20016492992388,
773
+ "efficiency": 68.0323450909838
774
+ },
775
+ {
776
+ "name": "Dumb-1.2-RC1",
777
+ "org": "56m",
778
+ "params": "34.6M",
779
+ "blimp": 64.87,
780
+ "arc": 34.13,
781
+ "aci": 54.16,
782
+ "wiki": 2.838,
783
+ "tokens": "210M",
784
+ "releaseDate": "2026-06-28",
785
+ "links": {
786
+ "card": "https://huggingface.co/56m/Dumb-1.2-RC1"
787
+ },
788
+ "score": 63.81563160872634,
789
+ "efficiency": 67.72908303685495
790
+ },
791
+ {
792
+ "name": "Byrne-86M",
793
+ "org": "quazim0t0",
794
+ "params": "86M",
795
+ "blimp": 70.33,
796
+ "arc": 34.68,
797
+ "aci": 52.16,
798
+ "wiki": 2.6839,
799
+ "tokens": "—",
800
+ "releaseDate": "2026-06-18",
801
+ "links": {
802
+ "card": "https://huggingface.co/Quazim0t0/Byrne-86M",
803
+ "base": "https://huggingface.co/Quazim0t0/Byrne-86M-Base"
804
+ },
805
+ "score": 66.15163269721914,
806
+ "efficiency": 67.69016874627584
807
+ },
808
+ {
809
+ "name": "Byrne-100M-Ultra-MC-sft",
810
+ "org": "quazim0t0",
811
+ "params": "113.9M",
812
+ "blimp": 78,
813
+ "arc": 26.5,
814
+ "wiki": 2.383,
815
+ "tokens": "~2B",
816
+ "releaseDate": "2026-08-19",
817
+ "links": {
818
+ "card": "https://huggingface.co/Quazim0t0/Byrne-100M-Ultra-MC"
819
+ },
820
+ "score": 66.69019002372889,
821
+ "efficiency": 67.45783133477885
822
+ },
823
+ {
824
+ "name": "SRLM-1M",
825
+ "org": "martico2432",
826
+ "params": "904k",
827
+ "blimp": 53.1,
828
+ "arc": 28.66,
829
+ "wiki": 4.455,
830
+ "tokens": "~16M",
831
+ "releaseDate": "2026-07-04",
832
+ "links": {
833
+ "card": "https://huggingface.co/Martico2432/srlm-1m"
834
+ },
835
+ "score": 55.382009072148584,
836
+ "efficiency": 67.2175930561982
837
+ },
838
+ {
839
+ "name": "Byrne-100M-Ultra-MC-dpo",
840
+ "org": "quazim0t0",
841
+ "params": "113.9M",
842
+ "blimp": 77.9,
843
+ "arc": 25.5,
844
+ "wiki": 2.385,
845
+ "tokens": "~2B",
846
+ "releaseDate": "2026-08-19",
847
+ "links": {
848
+ "card": "https://huggingface.co/Quazim0t0/Byrne-100M-Ultra-MC"
849
+ },
850
+ "score": 66.31852442080252,
851
+ "efficiency": 67.08188765331347
852
+ },
853
+ {
854
+ "name": "Dumb-1.2-Preview-0625",
855
+ "org": "56m",
856
+ "params": "34.6M",
857
+ "blimp": 64.05,
858
+ "arc": 32.79,
859
+ "aci": 54.44,
860
+ "wiki": 2.875,
861
+ "tokens": "~115-165M",
862
+ "releaseDate": "2026-06-25",
863
+ "links": {
864
+ "card": "https://huggingface.co/56m/Dumb-1.2-Preview-0625"
865
+ },
866
+ "score": 63.01844758825059,
867
+ "efficiency": 66.88301223323896
868
+ },
869
+ {
870
+ "name": "Byrne-100M-Ultra-MC-base",
871
+ "org": "quazim0t0",
872
+ "params": "113.9M",
873
+ "blimp": 81.1,
874
+ "arc": 20.5,
875
+ "wiki": 2.308,
876
+ "tokens": "~2B",
877
+ "releaseDate": "2026-08-19",
878
+ "links": {
879
+ "card": "https://huggingface.co/Quazim0t0/Byrne-100M-Ultra-MC"
880
+ },
881
+ "score": 65.91407674024059,
882
+ "efficiency": 66.67278455420127
883
+ },
884
+ {
885
+ "name": "Echo88-150M-Instruct",
886
+ "org": "exnivo",
887
+ "params": "150M",
888
+ "blimp": 72.67,
889
+ "arc": 31.4,
890
+ "aci": 50.83,
891
+ "wiki": 2.45,
892
+ "tokens": "~1.47B",
893
+ "releaseDate": "2026-05-05",
894
+ "links": {
895
+ "card": "https://huggingface.co/exnivo/Echo88-150M-Instruct"
896
+ },
897
+ "score": 66.38163401278912,
898
+ "efficiency": 66.38163401278912
899
+ },
900
+ {
901
+ "name": "Supra-Mini-v2",
902
+ "org": "supralabs",
903
+ "params": "168K",
904
+ "blimp": 53.5,
905
+ "arc": 26.8,
906
+ "aci": 48.95,
907
+ "wiki": 7.79,
908
+ "tokens": "—",
909
+ "releaseDate": "2026-05-12",
910
+ "links": {
911
+ "card": "https://huggingface.co/SupraLabs/Supra-Mini-v2-0.1M"
912
+ },
913
+ "score": 51.565520922818735,
914
+ "efficiency": 66.21356504337999
915
+ },
916
+ {
917
+ "name": "Ivme-Conversate-v1-Base",
918
+ "org": "ivmelabs",
919
+ "params": "22.03M",
920
+ "blimp": 61.4,
921
+ "arc": 30.85,
922
+ "aci": 53.66,
923
+ "wiki": 3.14,
924
+ "tokens": "~1.57B",
925
+ "releaseDate": "2026-06-05",
926
+ "links": {
927
+ "card": "https://huggingface.co/IvmeLabs/Ivme-Conversate-22M-Base"
928
+ },
929
+ "score": 60.96306546788355,
930
+ "efficiency": 65.8522330672023
931
+ },
932
+ {
933
+ "name": "Glimmer 1",
934
+ "org": "glintresearch",
935
+ "params": "11.9K",
936
+ "blimp": 52.43,
937
+ "arc": 25.46,
938
+ "aci": 50,
939
+ "wiki": 14.73,
940
+ "tokens": "500K",
941
+ "releaseDate": "2026-06-16",
942
+ "links": {
943
+ "card": "https://huggingface.co/Glint-Research/Glimmer-1-Base"
944
+ },
945
+ "score": 46.966205163181236,
946
+ "efficiency": 65.50622039282746
947
+ },
948
+ {
949
+ "name": "Aurora-80K",
950
+ "org": "auroraairesearch",
951
+ "params": "80K",
952
+ "blimp": 52.31,
953
+ "arc": 26.05,
954
+ "wiki": 9.78,
955
+ "tokens": "80M",
956
+ "releaseDate": "2026-08-20",
957
+ "links": {
958
+ "card": "https://huggingface.co/AuroraAI-Research/Aurora-80K"
959
+ },
960
+ "score": 49.563250998988714,
961
+ "efficiency": 65.17994379492052
962
+ },
963
+ {
964
+ "name": "MicroLM2-1M",
965
+ "org": "cromia",
966
+ "params": "1.71M",
967
+ "blimp": 54.2,
968
+ "arc": 27.4,
969
+ "aci": 49.54,
970
+ "wiki": 4.82,
971
+ "tokens": "~4.5B",
972
+ "releaseDate": "2026-05-22",
973
+ "links": {
974
+ "card": "https://huggingface.co/CromIA/MicroLM2-1M"
975
+ },
976
+ "score": 54.85944428751184,
977
+ "efficiency": 65.12136321418728
978
+ },
979
+ {
980
+ "name": "TinyMoE-100m-2x8-retrained",
981
+ "org": "flamef0x",
982
+ "params": "99.8M",
983
+ "blimp": 66.01,
984
+ "arc": 33.88,
985
+ "aci": 51.69,
986
+ "wiki": 2.879,
987
+ "tokens": "—",
988
+ "releaseDate": "2026-07-05",
989
+ "links": {
990
+ "card": "https://huggingface.co/FlameF0X/TinyMoE-100m-2x8-retrained"
991
+ },
992
+ "score": 64.02682960753184,
993
+ "efficiency": 65.11757145717472
994
+ },
995
+ {
996
+ "name": "DistillSupra-0.2M",
997
+ "org": "supralabs",
998
+ "params": "0.2M",
999
+ "blimp": 51.84,
1000
+ "arc": 26.77,
1001
+ "aci": 49.26,
1002
+ "wiki": 8.0696,
1003
+ "tokens": "1.5M",
1004
+ "releaseDate": "2026-05-15",
1005
+ "links": {
1006
+ "card": "https://huggingface.co/SupraLabs/DistillSupra-0.2M"
1007
+ },
1008
+ "score": 50.792064507542136,
1009
+ "efficiency": 64.8501466534245
1010
+ },
1011
+ {
1012
+ "name": "Cosmos-T2-Accelerate-Beta2",
1013
+ "org": "wop",
1014
+ "params": "9.96M",
1015
+ "blimp": 69,
1016
+ "arc": 28,
1017
+ "aci": 50.21,
1018
+ "wiki": 6.72,
1019
+ "tokens": "~10M",
1020
+ "releaseDate": "2026-06-03",
1021
+ "links": {
1022
+ "card": "https://huggingface.co/wop/Cosmos-T2-Accelerate-Beta2"
1023
+ },
1024
+ "score": 58.012606314477104,
1025
+ "efficiency": 64.59052961312662
1026
+ },
1027
+ {
1028
+ "name": "Ivme-Conversate-N-v1-Base",
1029
+ "org": "ivmelabs",
1030
+ "params": "9.55M",
1031
+ "blimp": 59.24,
1032
+ "arc": 26.81,
1033
+ "wiki": 4.95,
1034
+ "tokens": "~836M",
1035
+ "releaseDate": "2026-08-08",
1036
+ "links": {
1037
+ "card": "https://huggingface.co/IvmeLabs/Ivme-Conversate-S-v1-Base"
1038
+ },
1039
+ "score": 56.18419403051056,
1040
+ "efficiency": 62.653539676937676
1041
+ },
1042
+ {
1043
+ "name": "Cosmos-T2-Accelerate-beta",
1044
+ "org": "wop",
1045
+ "params": "5.03M",
1046
+ "blimp": 54.3,
1047
+ "arc": 26.3,
1048
+ "aci": 48.4,
1049
+ "wiki": 6.26,
1050
+ "tokens": "~22M",
1051
+ "releaseDate": "2026-06-02",
1052
+ "links": {
1053
+ "card": "https://huggingface.co/wop/Cosmos-T2-Accelerate-beta"
1054
+ },
1055
+ "score": 52.96846121101516,
1056
+ "efficiency": 60.48732290167559
1057
+ },
1058
+ {
1059
+ "name": "StorySupra-10M",
1060
+ "org": "supralabs",
1061
+ "params": "12.6M",
1062
+ "blimp": 61.47,
1063
+ "arc": 28.45,
1064
+ "aci": 49.86,
1065
+ "wiki": 8.76,
1066
+ "tokens": "—",
1067
+ "releaseDate": "2026-05-15",
1068
+ "links": {
1069
+ "card": "https://huggingface.co/SupraLabs/StorySupra-10M"
1070
+ },
1071
+ "score": 54.0729003655732,
1072
+ "efficiency": 59.672568692026175
1073
+ },
1074
+ {
1075
+ "name": "Cosmos-T2-Accelerate-Preview",
1076
+ "org": "wop",
1077
+ "params": "9.96M",
1078
+ "blimp": 56.95,
1079
+ "arc": 26.8,
1080
+ "aci": 50.71,
1081
+ "wiki": 6.73,
1082
+ "tokens": "462M",
1083
+ "releaseDate": "2026-06-01",
1084
+ "links": {
1085
+ "card": "https://huggingface.co/wop/Cosmos-T2-Accelerate-Preview",
1086
+ "demo": "https://huggingface.co/spaces/wop/Cosmos-T2-Chat"
1087
+ },
1088
+ "score": 53.58707907863565,
1089
+ "efficiency": 59.663201466020375
1090
+ },
1091
+ {
1092
+ "name": "Cosmos-T2A-low",
1093
+ "org": "wop",
1094
+ "params": "9.96M",
1095
+ "blimp": 47.8,
1096
+ "arc": 31,
1097
+ "aci": 49.54,
1098
+ "wiki": 5.31,
1099
+ "tokens": "~46.7M",
1100
+ "releaseDate": "2026-06-04",
1101
+ "links": {
1102
+ "card": "https://huggingface.co/wop/Cosmos-T2A-low"
1103
+ },
1104
+ "score": 53.349199024172016,
1105
+ "efficiency": 59.398348709381324
1106
+ },
1107
+ {
1108
+ "name": "Glint-0.3",
1109
+ "org": "glintresearch",
1110
+ "params": "1M",
1111
+ "blimp": 47.3,
1112
+ "arc": 25.5,
1113
+ "wiki": 7.87,
1114
+ "tokens": "~100M",
1115
+ "releaseDate": "2026-04-04",
1116
+ "links": {
1117
+ "card": "https://huggingface.co/Glint-Research/Glint-0.3"
1118
+ },
1119
+ "score": 49.00463935440382,
1120
+ "efficiency": 59.2705483402886
1121
+ },
1122
+ {
1123
+ "name": "TinyMoE-100m-2x8",
1124
+ "org": "flamef0x",
1125
+ "params": "99.8M",
1126
+ "blimp": 61.13,
1127
+ "arc": 25.88,
1128
+ "aci": 50.66,
1129
+ "wiki": 3.878,
1130
+ "tokens": "~625M",
1131
+ "releaseDate": "2026-06-15",
1132
+ "links": {
1133
+ "card": "https://huggingface.co/FlameF0X/TinyMoE-100m-2x8"
1134
+ },
1135
+ "score": 57.95852986999959,
1136
+ "efficiency": 58.945893986893935
1137
+ },
1138
+ {
1139
+ "name": "Supra-Mini-0.1M",
1140
+ "org": "supralabs",
1141
+ "params": "117K",
1142
+ "blimp": 51.77,
1143
+ "arc": 26.39,
1144
+ "aci": 52.23,
1145
+ "wiki": 25.17,
1146
+ "tokens": "500M",
1147
+ "releaseDate": "2026-05-18",
1148
+ "links": {
1149
+ "card": "https://huggingface.co/SupraLabs/Supra-Mini-0.1M"
1150
+ },
1151
+ "score": 43.863715881997,
1152
+ "efficiency": 56.987416612177725
1153
+ },
1154
+ {
1155
+ "name": "Blink-1-Instruct",
1156
+ "org": "glintresearch",
1157
+ "params": "1.09K",
1158
+ "blimp": 52.46,
1159
+ "arc": 26.8,
1160
+ "wiki": 70.44,
1161
+ "tokens": "100B",
1162
+ "releaseDate": "2026-06-21",
1163
+ "links": {
1164
+ "card": "https://huggingface.co/Glint-Research/Blink-1-Base",
1165
+ "base": "https://huggingface.co/Glint-Research/Blink-1-Base"
1166
+ },
1167
+ "score": 38.09820128775882,
1168
+ "efficiency": 56.94501186635348
1169
+ },
1170
+ {
1171
+ "name": "Blink-1-Base",
1172
+ "org": "glintresearch",
1173
+ "params": "1.09K",
1174
+ "blimp": 52.84,
1175
+ "arc": 26.6,
1176
+ "wiki": 71.35,
1177
+ "tokens": "100B",
1178
+ "releaseDate": "2026-06-21",
1179
+ "links": {
1180
+ "card": "https://huggingface.co/Glint-Research/Blink-1-Base"
1181
+ },
1182
+ "score": 38.0817146487351,
1183
+ "efficiency": 56.920369446944996
1184
+ },
1185
+ {
1186
+ "name": "peacebell-v1-148M",
1187
+ "org": "wayneworkman",
1188
+ "params": "148.55M",
1189
+ "blimp": 56.26,
1190
+ "arc": 27.36,
1191
+ "wiki": 6.8841,
1192
+ "tokens": "~35B",
1193
+ "releaseDate": "2026-09-19",
1194
+ "links": {
1195
+ "card": "https://huggingface.co/wayneworkman2012/peacebell-v1-148M"
1196
+ },
1197
+ "score": 53.40884446359834,
1198
+ "efficiency": 53.43053473259743
1199
+ },
1200
+ {
1201
+ "name": "Cosmos-T2-80M-Test",
1202
+ "org": "wop",
1203
+ "params": "87.60M",
1204
+ "blimp": 57.61,
1205
+ "arc": 25,
1206
+ "aci": 49.5,
1207
+ "wiki": 11.24,
1208
+ "tokens": "~18M",
1209
+ "releaseDate": "2026-05-31",
1210
+ "links": {
1211
+ "card": "https://huggingface.co/wop/Cosmos-T2-80M-Test"
1212
+ },
1213
+ "score": 50.15082355189342,
1214
+ "efficiency": 51.278566526214455
1215
+ },
1216
+ {
1217
+ "name": "Cosmos-T-80M",
1218
+ "org": "wop",
1219
+ "params": "79.7M",
1220
+ "blimp": 50.47,
1221
+ "arc": 27.82,
1222
+ "aci": 49.77,
1223
+ "wiki": 12.42,
1224
+ "tokens": "~21M",
1225
+ "releaseDate": "2026-05-30",
1226
+ "links": {
1227
+ "card": "https://huggingface.co/wop/Cosmos-T-80M"
1228
+ },
1229
+ "score": 48.11596794513753,
1230
+ "efficiency": 49.38807879600606
1231
+ },
1232
+ {
1233
+ "name": "Glint-0.2",
1234
+ "org": "glintresearch",
1235
+ "params": "1M",
1236
+ "blimp": 49.8,
1237
+ "arc": 27,
1238
+ "wiki": 636.4,
1239
+ "tokens": "~100M",
1240
+ "releaseDate": "2026-03-22",
1241
+ "links": {
1242
+ "card": "https://huggingface.co/Glint-Research/Glint-0.2"
1243
+ },
1244
+ "score": 25.599999999999998,
1245
+ "efficiency": 30.962905910561155
1246
+ },
1247
+ {
1248
+ "name": "Glint-0.1",
1249
+ "org": "glintresearch",
1250
+ "params": "1M",
1251
+ "blimp": 46.7,
1252
+ "arc": 21,
1253
+ "wiki": 4106963.13,
1254
+ "tokens": "~100M",
1255
+ "releaseDate": "2026-03-09",
1256
+ "links": {
1257
+ "card": "https://huggingface.co/Glint-Research/Glint-0.1"
1258
+ },
1259
+ "score": 22.566666666666666,
1260
+ "efficiency": 27.294124090429563
1261
+ }
1262
+ ]
1263
+ }
config.json ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "vocab": 12288,
3
+ "n_layer": 19,
4
+ "n_embd": 768,
5
+ "n_head": 12,
6
+ "block": 1024,
7
+ "ffn_mult": 2.8125,
8
+ "norm": "rmsnorm",
9
+ "norm_eps": 1e-06,
10
+ "pos": "rope",
11
+ "rope_theta": 100000.0,
12
+ "ffn": "swiglu",
13
+ "value_residual": true,
14
+ "qk_norm": true,
15
+ "logit_cap": 0.0,
16
+ "z_loss": 0.0
17
+ }
evaluate_glint.py ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Full pinned GLINT-style evaluation of the released safetensors weights."""
2
+ import argparse,json,math
3
+ from pathlib import Path
4
+ import torch
5
+ import pyarrow.parquet as pq
6
+ from huggingface_hub import hf_hub_download,list_repo_files
7
+ from load_model import load_model
8
+ import glint_metrics as metrics
9
+
10
+ def main():
11
+ p=argparse.ArgumentParser();p.add_argument('--model-dir',default='.');p.add_argument('--out',default='glint-results.json');a=p.parse_args()
12
+ root=Path(a.model_dir);torch.set_num_threads(4)
13
+ model,tok=load_model(root,'cuda')
14
+ reference=json.loads((root/'evaluation.json').read_text())
15
+ revisions=reference['dataset_revisions']
16
+ files={r:list_repo_files(r,repo_type='dataset',revision=rev) for r,rev in revisions.items()}
17
+ def rows(repo,config,split):
18
+ fs=[f for f in files[repo] if f.endswith('.parquet') and f.split('/')[0]==config and (split in Path(f).name or '/'+split+'/' in '/'+f)]
19
+ result=[]
20
+ for f in sorted(fs):result.extend(pq.read_table(hf_hub_download(repo,f,repo_type='dataset',revision=revisions[repo])).to_pylist())
21
+ if not result:raise ValueError(f'No data: {repo}/{config}/{split}')
22
+ return result
23
+ metrics._rows=rows
24
+ @torch.inference_mode()
25
+ def logits(x):
26
+ with torch.autocast('cuda',dtype=torch.bfloat16):return model(x)[0].float()
27
+ text=' '.join(r['text'] for r in rows('Salesforce/wikitext','wikitext-2-raw-v1','test')).strip()
28
+ ppl=metrics.compute_perplexity(logits,tok,text,'cuda')
29
+ byte_ppl=math.exp(math.log(ppl)*len(tok.encode(text).ids)/len(text.encode('utf-8')))
30
+ result={'parameters':sum(p.numel() for p in model.parameters()),'dataset_revisions':revisions,'protocol':reference['protocol'],'wikitext2_token_ppl':ppl,'wikitext2_byte_ppl':byte_ppl}
31
+ result.update(metrics.evaluate_blimp(logits,tok,'cuda'));result.update(metrics.evaluate_arc_easy(logits,tok,'cuda'))
32
+ assert result['blimp_n']==67000 and result['arc_n']==2376
33
+ board=json.loads((root/'board_snapshot.json').read_text())
34
+ wiki=100*max(0,min(1,1-(math.log(min(byte_ppl,500))-board['wikiMinLog'])/(board['wikiMaxLog']-board['wikiMinLog'])))
35
+ result['overall_score_fixed_board_snapshot']=(result['blimp_acc']+result['arc_easy_acc']+wiki)/3
36
+ bonus=1+.5*max(0,min(1,(board['paramLogMax']-math.log10(result['parameters']))/(board['paramLogMax']-board['paramLogMin'])))
37
+ result['efficiency_fixed_board_snapshot']=result['overall_score_fixed_board_snapshot']*bonus
38
+ Path(a.out).write_text(json.dumps(result,indent=2));print(json.dumps(result,indent=2))
39
+ if __name__=='__main__':main()
evaluation.json ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "step": 76294,
3
+ "parameters": 148910738,
4
+ "dataset_revisions": {
5
+ "Salesforce/wikitext": "b08601e04326c79dfdd32d625aee71d232d685c3",
6
+ "nyu-mll/blimp": "877fba0801ffb7cbd8c39c1ff314a46f053f6036",
7
+ "allenai/ai2_arc": "210d026faf9955653af8916fad021475a3f00453"
8
+ },
9
+ "continuation_tokens": 20000014336,
10
+ "provenance": {
11
+ "source_repo": "SlayerLab/gollem-v5-ckpts",
12
+ "source_revision": "963440ce6ab4ada7e95da4c1faa28ebb00c082d4",
13
+ "source_file": "final_128m_16x768/ckpt_760k.pt",
14
+ "source_sha256": "95d8b43fcd9c5fa06566da13dd03856b00300bad6f37fb4dc9442c18be9ac372",
15
+ "source_steps": 760000,
16
+ "source_training_tokens": 24903680000,
17
+ "parameters": 148910738,
18
+ "checks": [
19
+ {
20
+ "length": 16,
21
+ "max_absolute_logit_error": 0.0
22
+ },
23
+ {
24
+ "length": 128,
25
+ "max_absolute_logit_error": 0.0
26
+ },
27
+ {
28
+ "length": 256,
29
+ "max_absolute_logit_error": 0.0
30
+ }
31
+ ],
32
+ "optimizer": "reset for continued pretraining",
33
+ "method": "FFN expansion with zero new down columns; append three identity residual blocks"
34
+ },
35
+ "protocol": "GoLLeM r6 GLINT parity; bf16 forward, fp32 logprob; context 256",
36
+ "wikitext2_token_ppl": 21.452430290724198,
37
+ "bytes_per_token": 3.8604970807412586,
38
+ "wikitext2_byte_ppl": 2.2125733814416444,
39
+ "blimp_acc": 79.47,
40
+ "blimp_n": 67000,
41
+ "arc_easy_acc": 54.97,
42
+ "arc_n": 2376,
43
+ "efficiency_fixed_board_snapshot": 77.1358484476474,
44
+ "leader_reported_efficiency_snapshot": 80.20023332743875,
45
+ "comparison_caveat": "Competitor scores are self-reported, not independently re-evaluated under this protocol"
46
+ }
generate.py ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Simple causal completion CLI (no KV cache)."""
2
+ import argparse
3
+ from contextlib import nullcontext
4
+ import torch
5
+ from load_model import load_model
6
+
7
+ def main():
8
+ p=argparse.ArgumentParser()
9
+ p.add_argument('--model-dir',default='.')
10
+ p.add_argument('--prompt',default='The scientific method is')
11
+ p.add_argument('--max-new-tokens',type=int,default=64)
12
+ p.add_argument('--temperature',type=float,default=0.0)
13
+ p.add_argument('--device',default='cuda' if torch.cuda.is_available() else 'cpu')
14
+ args=p.parse_args();torch.set_num_threads(4)
15
+ model,tok=load_model(args.model_dir,args.device)
16
+ ids=tok.encode(args.prompt).ids
17
+ if not ids:raise ValueError('Prompt must encode to at least one token')
18
+ eos=tok.token_to_id('<|endoftext|>')
19
+ with torch.inference_mode():
20
+ for _ in range(args.max_new_tokens):
21
+ x=torch.tensor([ids[-model.block:]],device=args.device)
22
+ ctx=torch.autocast('cuda',dtype=torch.bfloat16) if args.device.startswith('cuda') else nullcontext()
23
+ with ctx:logits=model(x)[0][0,-1].float()
24
+ nxt=int(logits.argmax()) if args.temperature<=0 else int(torch.multinomial(torch.softmax(logits/args.temperature,dim=-1),1))
25
+ ids.append(nxt)
26
+ if nxt==eos:break
27
+ print(tok.decode(ids,skip_special_tokens=True))
28
+ if __name__=='__main__':main()
glint_metrics.py ADDED
@@ -0,0 +1,118 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """GLINT-1.3 likelihood protocol adapted to a logits callable; see NOTICE."""
2
+ import math
3
+ import torch
4
+ import torch.nn.functional as F
5
+ import numpy as np
6
+
7
+ def _rows(*args):
8
+ raise RuntimeError("Use evaluate_glint.py to install the pinned dataset loader")
9
+
10
+ def tokenize_many(tokenizer, texts, max_length=256):
11
+ all_ids = []
12
+ for text in texts:
13
+ ids = tokenizer.encode(text).ids
14
+ ids = [i for i in ids if i < tokenizer.get_vocab_size()]
15
+ if len(ids) > max_length:
16
+ ids = ids[:max_length]
17
+ all_ids.append(ids)
18
+ return all_ids
19
+
20
+ def batch_log_probs(logits_fn, tokenizer, texts, device, max_length=256, batch_size=128):
21
+ all_ids = tokenize_many(tokenizer, texts, max_length)
22
+ results = [-float("inf")] * len(all_ids)
23
+ with torch.inference_mode():
24
+ for start in range(0, len(all_ids), batch_size):
25
+ end = min(start + batch_size, len(all_ids))
26
+ batch = all_ids[start:end]
27
+ batch_indices = [j for j in range(start, end) if len(batch[j-start]) >= 2]
28
+ batch_seqs = [batch[j-start] for j in range(start, end) if len(batch[j-start]) >= 2]
29
+ if not batch_seqs:
30
+ continue
31
+ max_len = max(len(s) for s in batch_seqs)
32
+ B = len(batch_seqs)
33
+ padded_np = np.zeros((B, max_len - 1), dtype=np.int64)
34
+ targets_np = np.zeros((B, max_len - 1), dtype=np.int64)
35
+ mask_np = np.zeros((B, max_len - 1), dtype=bool)
36
+ for j, ids in enumerate(batch_seqs):
37
+ padded_np[j, :len(ids)-1] = ids[:-1]
38
+ targets_np[j, :len(ids)-1] = ids[1:]
39
+ mask_np[j, :len(ids)-1] = True
40
+ padded = torch.from_numpy(padded_np).to(device)
41
+ targets = torch.from_numpy(targets_np).to(device)
42
+ mask = torch.from_numpy(mask_np).to(device)
43
+ logits = logits_fn(padded)
44
+ log_probs = F.log_softmax(logits, dim=-1)
45
+ log_probs_flat = log_probs.view(-1, logits.size(-1))
46
+ targets_flat = targets.view(-1)
47
+ gathered = log_probs_flat[torch.arange(targets_flat.size(0), device=device), targets_flat]
48
+ gathered = gathered.view(B, -1)
49
+ gathered[~mask] = 0.0
50
+ sums = gathered.sum(dim=-1).tolist()
51
+ for bi, val in zip(batch_indices, sums):
52
+ results[bi] = val
53
+ return results
54
+
55
+ def compute_perplexity(logits_fn, tokenizer, text, device, max_length=256):
56
+ ids = tokenizer.encode(text).ids
57
+ ids = [i for i in ids if i < tokenizer.get_vocab_size()]
58
+ if len(ids) < 2:
59
+ return float("inf")
60
+ nll = 0.0; n_tokens = 0
61
+ for i in range(0, len(ids) - 1, max_length):
62
+ chunk = ids[i:i + max_length + 1]
63
+ if len(chunk) < 2:
64
+ continue
65
+ inputs = torch.tensor([chunk[:-1]], device=device)
66
+ targets = torch.tensor([chunk[1:]], device=device)
67
+ with torch.no_grad():
68
+ logits = logits_fn(inputs)
69
+ loss = F.cross_entropy(logits.view(-1, logits.size(-1)), targets.view(-1), reduction="sum")
70
+ nll += loss.item(); n_tokens += targets.numel()
71
+ return math.exp(nll / n_tokens) if n_tokens > 0 else float("inf")
72
+
73
+ BLIMP_CONFIGS = [
74
+ "adjunct_island","anaphor_gender_agreement","anaphor_number_agreement","animate_subject_passive",
75
+ "animate_subject_trans","causative","complex_NP_island","coordinate_structure_constraint_complex_left_branch",
76
+ "coordinate_structure_constraint_object_extraction","determiner_noun_agreement_1","determiner_noun_agreement_2",
77
+ "determiner_noun_agreement_irregular_1","determiner_noun_agreement_irregular_2","determiner_noun_agreement_with_adj_2",
78
+ "determiner_noun_agreement_with_adj_irregular_1","determiner_noun_agreement_with_adj_irregular_2",
79
+ "determiner_noun_agreement_with_adjective_1","distractor_agreement_relational_noun",
80
+ "distractor_agreement_relative_clause","drop_argument","ellipsis_n_bar_1","ellipsis_n_bar_2",
81
+ "existential_there_object_raising","existential_there_quantifiers_1","existential_there_quantifiers_2",
82
+ "existential_there_subject_raising","expletive_it_object_raising","inchoative","intransitive",
83
+ "irregular_past_participle_adjectives","irregular_past_participle_verbs","irregular_plural_subject_verb_agreement_1",
84
+ "irregular_plural_subject_verb_agreement_2","left_branch_island_echo_question","left_branch_island_simple_question",
85
+ "matrix_question_npi_licensor_present","npi_present_1","npi_present_2","only_npi_licensor_present","only_npi_scope",
86
+ "passive_1","passive_2","principle_A_c_command","principle_A_case_1","principle_A_case_2","principle_A_domain_1",
87
+ "principle_A_domain_2","principle_A_domain_3","principle_A_reconstruction","regular_plural_subject_verb_agreement_1",
88
+ "regular_plural_subject_verb_agreement_2","sentential_negation_npi_licensor_present","sentential_negation_npi_scope",
89
+ "sentential_subject_island","superlative_quantifiers_1","superlative_quantifiers_2","tough_vs_raising_1",
90
+ "tough_vs_raising_2","transitive","wh_island","wh_questions_object_gap","wh_questions_subject_gap",
91
+ "wh_questions_subject_gap_long_distance","wh_vs_that_no_gap","wh_vs_that_no_gap_long_distance",
92
+ "wh_vs_that_with_gap","wh_vs_that_with_gap_long_distance",
93
+ ]
94
+
95
+ def evaluate_blimp(logits_fn, tokenizer, device):
96
+ import os
97
+ ds = []
98
+ for c in BLIMP_CONFIGS:
99
+ ds.extend(_rows("nyu-mll/blimp", c, "train"))
100
+ assert len(ds) == 67000, f"BLiMP: {len(ds)} par, oczekiwano 67000 (67 fenomenow x 1000)"
101
+ good = batch_log_probs(logits_fn, tokenizer, [e["sentence_good"] for e in ds], device)
102
+ bad = batch_log_probs(logits_fn, tokenizer, [e["sentence_bad"] for e in ds], device)
103
+ correct = sum(1 for g, b in zip(good, bad) if g > b)
104
+ return {"blimp_acc": round(correct/len(ds)*100, 2), "blimp_n": len(ds)}
105
+
106
+ def evaluate_arc_easy(logits_fn, tokenizer, device):
107
+ ds = _rows("allenai/ai2_arc", "ARC-Easy", "test")
108
+ correct = 0; total = 0
109
+ for ex in ds:
110
+ q = ex["question"]; ch = ex["choices"]
111
+ full = [q + " " + t for t in ch["text"]]
112
+ lps = batch_log_probs(logits_fn, tokenizer, full, device, batch_size=4)
113
+ lpq = batch_log_probs(logits_fn, tokenizer, [q], device)[0]
114
+ best = max(range(len(lps)), key=lambda j: lps[j] - lpq)
115
+ if ch["label"][best] == ex["answerKey"]:
116
+ correct += 1
117
+ total += 1
118
+ return {"arc_easy_acc": round(correct/total*100, 2), "arc_n": total}
load_model.py ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Load the local safetensors release without Transformers remote code."""
2
+ import json
3
+ from pathlib import Path
4
+ from types import SimpleNamespace
5
+ import torch
6
+ from safetensors.torch import load_model as load_safetensors
7
+ from tokenizers import Tokenizer
8
+ from modeling_gollem import GPT
9
+
10
+ def load_model(directory='.',device='cpu'):
11
+ directory=Path(directory)
12
+ cfg=SimpleNamespace(**json.loads((directory/'config.json').read_text()))
13
+ model=GPT(cfg.vocab,cfg.n_layer,cfg.n_embd,cfg.n_head,cfg.block,cfg)
14
+ load_safetensors(model,str(directory/'model.safetensors'),strict=True)
15
+ return model.eval().to(device),Tokenizer.from_file(str(directory/'tokenizer.json'))
loss-curve.png ADDED

Git LFS Details

  • SHA256: f1874ce1fa78f5d02508a444d0f139c099024abdf3eca157058182090027ea8d
  • Pointer size: 131 Bytes
  • Size of remote file: 263 kB
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:05e2a496967bdd36da3e1bdd4fb8b10277b92a319b1f5f1e61062454eecc7b2e
3
+ size 595664120
modeling_gollem.py ADDED
@@ -0,0 +1,140 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """GoLLeM inference architecture, extracted unchanged from the pinned r6 trainer.
2
+ Source: SlayerLab/gollem-v5-ckpts; Apache-2.0. See README and NOTICE.
3
+ """
4
+ import torch
5
+ from torch import nn
6
+ from torch.nn import functional as F
7
+
8
+ class RMSNorm(nn.Module):
9
+ """Qwen3-style RMSNorm (fp32-compute dla stabilnosci). 1D weight -> AdamW w split-Muon."""
10
+ def __init__(self, d, eps=1e-6):
11
+ super().__init__()
12
+ self.weight = nn.Parameter(torch.ones(d))
13
+ self.eps = eps
14
+
15
+ def forward(self, x):
16
+ return x * torch.rsqrt(x.float().pow(2).mean(-1, keepdim=True) + self.eps).to(x.dtype) * self.weight
17
+
18
+
19
+ def make_norm(d, cfg):
20
+ return RMSNorm(d, cfg.norm_eps) if cfg.norm == "rmsnorm" else nn.LayerNorm(d)
21
+
22
+
23
+ def apply_rope(x, base=100000.0):
24
+ """Parameter-free RoPE na [B,H,T,D] (interleaved-conv, port z qwen_model.py). Train==eval
25
+ MUSZA uzywac tej samej konwencji (self-contained eval -> spojne)."""
26
+ _, _, T, dim = x.shape
27
+ pos = torch.arange(T, device=x.device, dtype=torch.float32)
28
+ freq = 1.0 / (base ** (torch.arange(0, dim, 2, device=x.device, dtype=torch.float32) / dim))
29
+ ang = torch.outer(pos, freq)
30
+ cos, sin = ang.cos().to(x.dtype)[None, None], ang.sin().to(x.dtype)[None, None]
31
+ even, odd = x[..., ::2], x[..., 1::2]
32
+ return torch.stack((even * cos - odd * sin, even * sin + odd * cos), dim=-1).flatten(-2)
33
+
34
+
35
+ class SwiGLU(nn.Module):
36
+ """Qwen3 gated-MLP: down(silu(gate(x))*up(x)). 3x 2D bez-bias -> wszystkie do Muon."""
37
+ def __init__(self, d, hidden):
38
+ super().__init__()
39
+ self.gate = nn.Linear(d, hidden, bias=False)
40
+ self.up = nn.Linear(d, hidden, bias=False)
41
+ self.down = nn.Linear(hidden, d, bias=False)
42
+
43
+ def forward(self, x):
44
+ return self.down(F.silu(self.gate(x)) * self.up(x))
45
+
46
+
47
+ class Block(nn.Module):
48
+ def __init__(self, d, nh, block, cfg, is_first=False):
49
+ super().__init__()
50
+ self.ln1 = make_norm(d, cfg)
51
+ self.ln2 = make_norm(d, cfg)
52
+ self.qkv = nn.Linear(d, 3 * d)
53
+ self.proj = nn.Linear(d, d)
54
+ if cfg.ffn == "swiglu":
55
+ self.mlp = SwiGLU(d, int(round(cfg.ffn_mult * d)))
56
+ else:
57
+ self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
58
+ self.nh = nh
59
+ self.d = d
60
+ self.cfg = cfg
61
+ self.is_first = is_first
62
+ if cfg.value_residual and not is_first:
63
+ self.vr_lambda = nn.Parameter(torch.zeros(1))
64
+ if cfg.qk_norm:
65
+ hd = d // nh
66
+ self.q_norm = RMSNorm(hd, cfg.norm_eps)
67
+ self.k_norm = RMSNorm(hd, cfg.norm_eps)
68
+
69
+ def forward(self, x, v0=None):
70
+ B, T, D = x.size()
71
+ h = self.ln1(x)
72
+ q, k, v = self.qkv(h).split(self.d, dim=2)
73
+ hd = D // self.nh
74
+ q = q.view(B, T, self.nh, hd).transpose(1, 2)
75
+ k = k.view(B, T, self.nh, hd).transpose(1, 2)
76
+ v = v.view(B, T, self.nh, hd).transpose(1, 2)
77
+ if self.cfg.qk_norm:
78
+ q = self.q_norm(q)
79
+ k = self.k_norm(k)
80
+ if self.cfg.pos == "rope":
81
+ q = apply_rope(q, self.cfg.rope_theta)
82
+ k = apply_rope(k, self.cfg.rope_theta)
83
+ if self.cfg.value_residual:
84
+ if self.is_first:
85
+ v0 = v
86
+ else:
87
+ v = v + self.vr_lambda * v0
88
+ y = F.scaled_dot_product_attention(q, k, v, is_causal=True)
89
+ y = y.transpose(1, 2).contiguous().view(B, T, D)
90
+ x = x + self.proj(y)
91
+ x = x + self.mlp(self.ln2(x))
92
+ return x, v0
93
+
94
+
95
+ class GPT(nn.Module):
96
+ def __init__(self, vocab, n_layer, n_embd, n_head, block, cfg):
97
+ super().__init__()
98
+ self.cfg = cfg
99
+ self.tok = nn.Embedding(vocab, n_embd)
100
+ self.use_rope = cfg.pos == "rope"
101
+ if not self.use_rope:
102
+ self.pos = nn.Embedding(block, n_embd)
103
+ self.blocks = nn.ModuleList([Block(n_embd, n_head, block, cfg, is_first=(i == 0)) for i in range(n_layer)])
104
+ self.lnf = make_norm(n_embd, cfg)
105
+ self.head = nn.Linear(n_embd, vocab, bias=False)
106
+ self.head.weight = self.tok.weight # tie
107
+ self.block = block
108
+ self.apply(self._init)
109
+
110
+ def _init(self, m):
111
+ if isinstance(m, nn.Linear):
112
+ nn.init.normal_(m.weight, 0.0, 0.02)
113
+ if m.bias is not None:
114
+ nn.init.zeros_(m.bias)
115
+ elif isinstance(m, nn.Embedding):
116
+ nn.init.normal_(m.weight, 0.0, 0.02)
117
+
118
+ def forward(self, idx, targets=None):
119
+ B, T = idx.size()
120
+ x = self.tok(idx)
121
+ if not self.use_rope:
122
+ pos = torch.arange(T, device=idx.device)
123
+ x = x + self.pos(pos)[None]
124
+ v0 = None
125
+ for b in self.blocks:
126
+ x, v0 = b(x, v0)
127
+ logits = self.head(self.lnf(x))
128
+ cap = getattr(self.cfg, "logit_cap", 0.0)
129
+ if cap and cap > 0:
130
+ logits = cap * torch.tanh(logits / cap)
131
+ loss = None
132
+ if targets is not None:
133
+ flat = logits.view(-1, logits.size(-1))
134
+ loss = F.cross_entropy(flat, targets.view(-1))
135
+ zc = getattr(self.cfg, "z_loss", 0.0)
136
+ if zc and zc > 0:
137
+ lse = torch.logsumexp(flat, dim=-1)
138
+ loss = loss + zc * (lse * lse).mean()
139
+ return logits, loss
140
+
release_verification.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "weights_and_logits": "exact match to final training checkpoint at contexts 16, 256, 1024",
3
+ "glint_likelihood_and_perplexity_functions": "exact match on golden text fixtures including padding, truncation and empty input",
4
+ "official_harness_revision": "c3f99e246aa2c64382f9668dc864533617482d0a",
5
+ "token_12285": "<|endoftext|>",
6
+ "status": "passed"
7
+ }
requirements.txt ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ torch==2.11.0
2
+ tokenizers==0.22.2
3
+ safetensors==0.6.2
4
+ huggingface-hub==0.36.0
5
+ numpy>=1.26,<3
6
+ pyarrow>=20,<25
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
training_manifest.json ADDED
@@ -0,0 +1,72 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "checkpoint_sha256": "c4e8b5fb5730242d790d469806a8735be2424b9351dcdc25f643980ba3028236",
3
+ "weights_sha256": "05e2a496967bdd36da3e1bdd4fb8b10277b92a319b1f5f1e61062454eecc7b2e",
4
+ "parameters": 148910738,
5
+ "reference_step": 76294,
6
+ "optimizer_updates": 39147,
7
+ "continuation_tokens": 20000014336,
8
+ "provenance": {
9
+ "source_repo": "SlayerLab/gollem-v5-ckpts",
10
+ "source_revision": "963440ce6ab4ada7e95da4c1faa28ebb00c082d4",
11
+ "source_file": "final_128m_16x768/ckpt_760k.pt",
12
+ "source_sha256": "95d8b43fcd9c5fa06566da13dd03856b00300bad6f37fb4dc9442c18be9ac372",
13
+ "source_steps": 760000,
14
+ "source_training_tokens": 24903680000,
15
+ "parameters": 148910738,
16
+ "checks": [
17
+ {
18
+ "length": 16,
19
+ "max_absolute_logit_error": 0.0
20
+ },
21
+ {
22
+ "length": 128,
23
+ "max_absolute_logit_error": 0.0
24
+ },
25
+ {
26
+ "length": 256,
27
+ "max_absolute_logit_error": 0.0
28
+ }
29
+ ],
30
+ "optimizer": "reset for continued pretraining",
31
+ "method": "FFN expansion with zero new down columns; append three identity residual blocks"
32
+ },
33
+ "safetensors_roundtrip": "all model tensors exactly equal; tied weights restored",
34
+ "training_config": {
35
+ "model": {
36
+ "vocab": 12288,
37
+ "n_layer": 19,
38
+ "n_embd": 768,
39
+ "n_head": 12,
40
+ "block": 1024,
41
+ "ffn_mult": 2.8125,
42
+ "norm": "rmsnorm",
43
+ "norm_eps": 1e-06,
44
+ "pos": "rope",
45
+ "rope_theta": 100000.0,
46
+ "ffn": "swiglu",
47
+ "value_residual": true,
48
+ "qk_norm": true,
49
+ "logit_cap": 0.0,
50
+ "z_loss": 0.0
51
+ },
52
+ "train": {
53
+ "seed": 1337,
54
+ "micro_batch": 32,
55
+ "global_batch": 256,
56
+ "steps": 76294,
57
+ "warmup": 250,
58
+ "lr": 0.0002,
59
+ "min_lr": 2e-05,
60
+ "muon_lr": 0.006666666666666667,
61
+ "checkpoint_every": 1000,
62
+ "validate_every": 1000,
63
+ "log_every": 20,
64
+ "compile": true
65
+ },
66
+ "data": {
67
+ "train": "/data/final/train.bin",
68
+ "validation": "/data/final/val.bin",
69
+ "sha256": "ecfd0a4040a0ece6f07f728f092a1b373b6d253b4adfbe259c5107c1fc16681d"
70
+ }
71
+ }
72
+ }
training_metrics.jsonl ADDED
The diff for this file is too large to render. See raw diff