NovaAI commited on
Commit
b418a7a
·
verified ·
1 Parent(s): e7431db

Upload folder using huggingface_hub

Browse files
LICENSE.custom.md ADDED
@@ -0,0 +1,149 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # BaiHu-V1-Flash Custom License
2
+
3
+ **Version 1.0 | Effective date: 2026-09-29**
4
+
5
+ Copyright (c) 2026 NovaAI6868. All rights reserved.
6
+
7
+ This License governs the **BaiHu-V1-Flash** model (the "Model"), including its weights,
8
+ configuration, inference code and accompanying documentation. **By downloading, copying,
9
+ installing, invoking or otherwise using the Model, you confirm that you have read,
10
+ understood and agree to be bound by all terms of this License.** If you do not agree,
11
+ stop using the Model immediately and delete all copies.
12
+
13
+ ---
14
+
15
+ ## 1. Definitions
16
+
17
+ - **"Personal Use"** means use by a natural person for their own learning, research,
18
+ teaching, experimentation, hobby projects or other non-commercial purposes, where such
19
+ use does **not** directly or indirectly generate commercial revenue and does **not**
20
+ provide paid services to third parties based on the Model.
21
+ - **"Commercial Use"** means any use for profit, any use in the business operations of a
22
+ for-profit entity, or any use that generates commercial revenue. This includes, without
23
+ limitation: internal production systems of a company, paid APIs or SaaS offered to third
24
+ parties, integration into paid products, advertising or commercial analytics,
25
+ client-facing delivery work, and any paid service built on the Model.
26
+ - **"You"** means the natural person or legal entity exercising rights under this License.
27
+ - **"Derivative Model"** means any model obtained by fine-tuning, continued training,
28
+ distillation, quantization, pruning, merging or otherwise modifying the Model.
29
+
30
+ ## 2. Free Grant: Personal Use
31
+
32
+ Subject to this License, **Personal Use is free of charge**, with no fee and no prior
33
+ application required. You may:
34
+
35
+ 1. download, install, run and copy the Model;
36
+ 2. modify the Model and create or train Derivative Models;
37
+ 3. use the Model freely in personal, non-commercial projects, including publicly
38
+ releasing your research results and demos.
39
+
40
+ ## 3. Paid Grant: Commercial Use
41
+
42
+ **Any Commercial Use requires prior written commercial authorization.** Commercial Use
43
+ without such authorization is not permitted.
44
+
45
+ Commercial licenses are negotiated based on scale of use, deployment model and term, and
46
+ may include custom terms. To obtain one, contact:
47
+
48
+ > **Business contact: novaweb6868@outlook.com**
49
+ >
50
+ > Please briefly state: your entity name, intended use case, expected volume or deployment
51
+ > scale, and whether redistribution rights are required.
52
+
53
+ Until you and the copyright holder reach a written agreement on commercial licensing,
54
+ this License grants you **no** right of Commercial Use.
55
+
56
+ ## 4. General Restrictions (applies to both Personal and Commercial Use)
57
+
58
+ Whether your use is Personal Use or commercially licensed, you must not:
59
+
60
+ 1. remove, obscure or alter copyright notices, license identifiers or attribution in the Model;
61
+ 2. use the Model or any Derivative Model for any activity that violates applicable laws
62
+ or regulations;
63
+ 3. use the Model to generate or distribute malicious code, fraudulent content, or content
64
+ that infringes the lawful rights of others;
65
+ 4. sublicense the Model or any Derivative Model to third parties under terms that conflict
66
+ with this License (except where redistribution rights are expressly granted in a
67
+ commercial license);
68
+ 5. claim original authorship of the Model, or make any promise or warranty on behalf of
69
+ the copyright holder.
70
+
71
+ ## 5. Licensing of Derivative Models
72
+
73
+ Derivative Models you create **remain subject to this License**; your modifications do not
74
+ remove them from its scope. Sections 2, 3 and 4 therefore apply equally to Derivative
75
+ Models: Personal Use is free, Commercial Use requires a paid license. When distributing a
76
+ Derivative Model you must include the full text of this License and prominently state that
77
+ it is built upon BaiHu-V1-Flash.
78
+
79
+ ## 6. Upstream License and Third-Party Components
80
+
81
+ The Model is an architectural retrofit and continued-pretraining product of
82
+ `Qwen/Qwen3-0.6B-Base` (Apache License 2.0). This License **only** governs the rights the
83
+ copyright holder holds in the portions newly added to the Model, and **does not alter**
84
+ the original license terms of upstream components. All use of the Model must also comply
85
+ with the licenses applicable to those upstream components. In the event of a conflict,
86
+ the upstream license prevails with respect to the upstream components.
87
+
88
+ ## 7. What Counts as Commercial Use
89
+
90
+ The following **are** Commercial Use and require authorization (this list is not
91
+ exhaustive):
92
+
93
+ - use in the business of a company, studio, sole proprietorship or other for-profit
94
+ entity, even if no fee is charged directly;
95
+ - delivering projects or deliverables that incorporate the Model to clients;
96
+ - use to generate content, products or services sold externally;
97
+ - offering a free service built on the Model while monetizing through advertising,
98
+ traffic diversion, data monetization or similar.
99
+
100
+ The following are **generally not** Commercial Use (final determination rests with the
101
+ copyright holder):
102
+
103
+ - personal study, coursework, or academic research that is publicly published;
104
+ - use by non-profit organizations for non-commercial purposes.
105
+
106
+ ## 8. Disclaimer of Warranty
107
+
108
+ **The Model is provided "AS IS", without warranty of any kind**, express or implied,
109
+ including but not limited to the warranties of merchantability, fitness for a particular
110
+ purpose, accuracy and non-infringement. You bear all risks and consequences arising from
111
+ your use of the Model. In no event shall the copyright holder be liable for any direct,
112
+ indirect, incidental, special or consequential damages arising from the use of, or
113
+ inability to use, the Model. The Model may produce inaccurate, biased or otherwise
114
+ inappropriate output; you are responsible for evaluating it and for the consequences.
115
+
116
+ ## 9. Termination
117
+
118
+ If you breach any term of this License, the rights granted to you under it **terminate
119
+ automatically**, without notice. Upon termination you must immediately stop using the
120
+ Model and delete all copies and Derivative Models. Sections 4, 5, 8 and 10 survive
121
+ termination.
122
+
123
+ ## 10. Governing Law and Dispute Resolution
124
+
125
+ The formation, validity, interpretation and dispute resolution of this License are
126
+ governed by the laws of the copyright holder's jurisdiction. The parties shall first seek
127
+ an amicable resolution of any dispute arising from this License.
128
+
129
+ ## 11. Changes to This License
130
+
131
+ The copyright holder reserves the right to revise this License from time to time. A
132
+ revised License takes effect upon publication; your continued use of the Model after such
133
+ publication constitutes acceptance of the revised terms.
134
+
135
+ ---
136
+
137
+ ## Quick Summary (for convenience only; the full text above controls)
138
+
139
+ | Use case | Fee | Notes |
140
+ |---|---|---|
141
+ | Personal study / research / experimentation / hobby | **Free** | No application required |
142
+ | Academic research with public publication | **Free** | Please cite the source |
143
+ | Non-profit, non-commercial use by non-profits | **Free** | No application required |
144
+ | Internal use by a company (even without direct fees) | **License required** | Contact novaweb6868@outlook.com |
145
+ | Paid API / SaaS / product integration | **License required** | Contact novaweb6868@outlook.com |
146
+ | Client-facing deliverables | **License required** | Contact novaweb6868@outlook.com |
147
+ | Distributing fine-tuned / quantized versions | Same license | Free for personal, paid for commercial |
148
+
149
+ **Commercial licensing contact: novaweb6868@outlook.com**
README.md ADDED
@@ -0,0 +1,237 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: transformers
3
+ pipeline_tag: text-generation
4
+ language:
5
+ - zh
6
+ - en
7
+ license: other
8
+ license_name: baihu-custom-license
9
+ license_link: https://huggingface.co/NovaAI6868/BaiHu-V1-Flash/blob/main/LICENSE.custom.md
10
+ base_model: Qwen/Qwen3-0.6B-Base
11
+ tags:
12
+ - sparse-attention
13
+ - subq
14
+ - ssa
15
+ - long-context
16
+ - commercial-license-required
17
+ - text-generation
18
+ ---
19
+
20
+ # BaiHu-V1-Flash
21
+
22
+ **BaiHu-V1-Flash** is a retrofit of the dense-attention model `Qwen/Qwen3-0.6B-Base` into an
23
+ **SSA (Sparse-attention + SubQ)** architecture, obtained by continued pretraining.
24
+
25
+ - Base model: `Qwen/Qwen3-0.6B-Base` (28 layers / 16 Q heads / 8 KV heads / head_dim 128 / 32K context / tied embeddings)
26
+ - Parameters: 598.8M
27
+ - Training data: mixed Chinese + English (Fineweb-Edu-Chinese-V2.1 + fineweb-edu, 50/50)
28
+ - License: **free for personal use; a paid license is required for commercial use** (see "License" below)
29
+
30
+ ---
31
+
32
+ ## 1. Architecture: SSA (three paths, each with its own softmax, then summed)
33
+
34
+ Every layer keeps the base model's MLP / RMSNorm weights and replaces full attention with an
35
+ SSA layer built from three parallel paths:
36
+
37
+ | Path | Role | Complexity |
38
+ |---|---|---|
39
+ | `shared` | every query sees all **completed** blocks through one compressed vector per block | `O(T·T/B)` |
40
+ | `local` | dense causal attention over the most recent window | `O(T·w)` |
41
+ | `sparse` (**SubQ**) | only 4 of 16 query heads produce block scores, shared across the whole head group; real attention is computed only for the selected top-k blocks | `O(T·k·B)` |
42
+
43
+ ### Hyperparameters
44
+
45
+ | Parameter | Value | Meaning |
46
+ |---|---|---|
47
+ | `ssa_block_size` | 64 | block size B |
48
+ | `ssa_top_k` | 8 | number of blocks selected by the sparse path |
49
+ | `ssa_local_blocks` | 2 | local window = 3 × 64 = 192 tokens |
50
+ | `ssa_num_subq_heads` | 4 | SubQ heads, r = 16 / 4 = 4 |
51
+ | `ssa_router_dim` | 128 | router subspace dimension |
52
+ | `ssa_compress_dim` | 128 | block compression dimension |
53
+
54
+ Only **2.75M** parameters are new (≈0.46% of the model); all other weights are inherited
55
+ from the base model.
56
+
57
+ ---
58
+
59
+ ## 2. Training
60
+
61
+ | Item | Setting |
62
+ |---|---|
63
+ | Starting point | a conversion checkpoint that is **bit-exact** with the base model (`max\|Δlogit\| = 0.000e+00`) |
64
+ | Tokens seen | 5.0M (≈0.25 epoch of the corpus) |
65
+ | Sequence length | 512 (must be a multiple of `ssa_block_size = 64`) |
66
+ | Effective batch | 8192 tokens (batch 2 × grad_accum 8) |
67
+ | Precision | fp32 |
68
+ | Optimizer | SGD with momentum 0.9 |
69
+ | Learning rate | 5e-4, 50-step warmup, cosine decay to 10% |
70
+ | Hardware | single NVIDIA GTX TITAN X (Maxwell, sm_52, 12.9 GB) |
71
+ | Throughput | ≈274 tokens/s |
72
+
73
+ ### Critical issues found and fixed during this retrofit
74
+
75
+ Several defects silently break training and are worth documenting:
76
+
77
+ 1. **Both new branches had identically zero gradients (blocking).** To make the converted
78
+ model bit-exact with the base model, `compress_out` and `router_out` were initialized to
79
+ exactly zero, and the branches were skipped entirely by a gate. The branch output was
80
+ therefore always zero, so the back-propagated gradient was also always zero: all 2.75M
81
+ SSA parameters **stayed frozen for the entire run** and the sparse attention was dead
82
+ code. The fix is a small non-zero initialization.
83
+ 2. **Routing was non-differentiable.** `top-k` produces hard indices, and indexing is not
84
+ differentiable. If the routing scores are used only to decide *which* blocks to read and
85
+ never enter the softmax, the gradients of `router_q` / `router_k` are **exactly zero** —
86
+ the router can never learn to route. The fix is to feed the selected blocks' scores,
87
+ squashed through `tanh` and gently scaled, into the attention logits as an additive bias.
88
+ 3. **The shared summary was a sum, not a mean.** Its magnitude grew linearly with the
89
+ prefix, and because that branch is injected at full weight (a single-element softmax has
90
+ probability exactly 1), it swamped the residual stream: hidden states grew from 0.2 to
91
+ about 7 in layer 0 and to about 1900 by layer 27, and validation loss went 3.56 → 11.38.
92
+ The fix is to divide by the token count.
93
+ 4. **The shared branch needs an explicit gate.** With a single-element softmax the
94
+ probability is always 1, so the initialization scale of `compress_out` cannot control the
95
+ injection strength at all (measured: scales from 1e-4 to 0.03 all left the loss at
96
+ exactly 7.3526). A learnable scalar gate, initialized to a small positive value, lets the
97
+ optimizer decide how far to open it.
98
+
99
+ ---
100
+
101
+ ## 3. Evaluation
102
+
103
+ ### 3.1 Language modeling perplexity (validation set, identical windows)
104
+
105
+ | Sequence length | Qwen3-0.6B-Base | BaiHu-V1-Flash |
106
+ |---|---|---|
107
+ | 512 | 3.6726 / ppl 39.354 | 4.4108 / ppl 82.335 |
108
+ | 1024 | 3.4691 / ppl 32.109 | 4.2575 / ppl 70.631 |
109
+ | 2048 | 3.0801 / ppl 21.761 | 3.9213 / ppl 50.467 |
110
+
111
+ ### 3.2 Standard benchmarks (lm-evaluation-harness)
112
+
113
+ | Task | Qwen3-0.6B-Base | BaiHu-V1-Flash | Delta |
114
+ |---|---|---|---|
115
+ | arc_easy | 0.5550 | 0.6250 | +0.0700 |
116
+ | hellaswag | 0.5350 | 0.5200 | -0.0150 |
117
+ | piqa | 0.7050 | 0.6950 | -0.0100 |
118
+ | winogrande | 0.6300 | 0.6300 | +0.0000 |
119
+
120
+ ### 3.3 Inference compute and resource usage
121
+
122
+ | Metric | Qwen3-0.6B-Base | BaiHu-V1-Flash |
123
+ |---|---|---|
124
+ | Parameters (M) | 596.0500 | 598.8000 |
125
+ | Prefill peak memory (GB) | 9.6200 | 8.0210 |
126
+ | Generation peak memory (GB) | 9.9120 | 9.0050 |
127
+ | Prefill latency (s) | 0.5030 | 1.4310 |
128
+ | TTFT (ms) | 502.7 | 1431.2 |
129
+ | TPOT (ms) | 36.6 | 80.5 |
130
+ | Decode throughput (tok/s) | 27.3300 | 12.4200 |
131
+ | Attention FLOPs/token (GFLOPs) | 0.1176 | 0.1057 |
132
+ | Attention key accesses vs full attention | 1.0000 | 0.8993 |
133
+ | GPU utilization mean/max (%) | 93.9 | 43.3 |
134
+ | Power mean/max (W) | 179.1 | 124.3 |
135
+
136
+ Positive findings: peak inference memory is lower (generation 9.005 vs 9.912 GB, −9.2%), and the
137
+ model draws less power because it is not compute-bound.
138
+
139
+ Negative findings, stated plainly:
140
+
141
+ - **Decode is 2.2× slower** (12.42 vs 27.33 tok/s) and **prefill is 2.8× slower**
142
+ (TTFT 1431 vs 503 ms), despite the sparse path reading fewer keys. The current
143
+ implementation loops over query blocks in Python and issues many small kernels, so
144
+ launch overhead dominates the FLOPs saved. **The sparse attention does not yet pay off
145
+ on this hardware.**
146
+ - **Attention key accesses are still 89.9% of full attention** at this sequence length.
147
+ The reason is structural: the local window already covers 3 blocks (192 tokens) and the
148
+ sparse path reads up to `top_k + 1 = 9` blocks from a grid that only has 16 blocks at
149
+ 1024 tokens, so the selected set is almost the whole grid. Sparsity only becomes a real
150
+ saving once the sequence is long relative to `top_k × block_size` (i.e. well beyond
151
+ 10k tokens).
152
+ - **Language modeling perplexity is clearly worse than the base model** at every length
153
+ tested (ppl 82.3 vs 39.4 at 512; 50.5 vs 21.8 at 2048). This is the honest cost of
154
+ shrinking the dense local window from the full prefix to 192 tokens while the new
155
+ long-range branches are still very weakly trained.
156
+ - **Standard benchmarks are roughly neutral but not better**: arc_easy improves
157
+ (+0.070 acc_norm), winogrande is unchanged, while hellaswag (−0.015) and piqa (−0.010)
158
+ regress slightly.
159
+
160
+ ### 3.4 Sparsity
161
+
162
+ Two measurements are reported because they answer different questions and are **not**
163
+ interchangeable:
164
+
165
+ | Scenario | Average blocks selected per query | Key access ratio vs full attention |
166
+ |---|---|---|
167
+ | Chunked forward over a 512-token validation window | 0.38 | 0.0938 |
168
+ | Prefill of 1024 tokens (steady state) | up to `top_k` | 0.8993 |
169
+
170
+ The first number averages over *all* query blocks including the early ones, which have no
171
+ completed blocks available to select and therefore read nothing through the sparse path.
172
+ The second is the steady-state ratio for later queries, and it is the one that matters for
173
+ efficiency — see the note in section 3.3: at these sequence lengths the sparse path is not
174
+ yet saving meaningful work.
175
+
176
+ ---
177
+
178
+ ## 4. Known Limitations
179
+
180
+ 1. **Trained for very little.** Only 5.0M tokens (≈0.25 epoch). The new branches have not
181
+ converged; ppl is well above the base model and decode is slower (see 3.3). Continuing
182
+ to 200M tokens or more is required before the SSA layers can genuinely take over
183
+ long-range modeling.
184
+ 2. **The shared branch does not appear to help and was actively suppressed by the
185
+ optimizer.** Its learnable gate *decreased* over training (0.0100 → 0.0129 at step 100 →
186
+ 0.0122 at step 610) instead of growing, meaning the optimizer found the single
187
+ prefix-mean summary not worth injecting. This is the single most important thing to
188
+ change next: replace it with **one compressed vector per block** (same `O(T·T/B)` cost,
189
+ far more information retained).
190
+ 3. **Sparse attention is not yet a net win on this hardware.** It reads fewer keys but runs
191
+ 2.2× slower because the implementation loops over query blocks in Python and issues many
192
+ small kernels. It needs kernel-level batching (or a fused implementation) before the
193
+ sparsity can translate into speed.
194
+ 4. **Sparsity only pays off at long sequences.** With `top_k=8` and `block_size=64`, the
195
+ sparse path can read up to 9 blocks = 576 tokens; at 1024 tokens the grid only has 16
196
+ blocks, so the selected set covers most of the context and the local window already
197
+ covers the rest. Real savings require sequences well beyond 10k tokens.
198
+ 5. **Routing quality is not fully validated.** The distribution of selected top-k blocks
199
+ should be checked for degeneration (e.g. always selecting the same blocks). The
200
+ non-differentiable-routing bug that would have made this *impossible* to learn has been
201
+ fixed (see section 2), but the learned policy has not been analyzed in detail.
202
+ 6. **Block size B = 64 was not ablated.** Limited by memory and the Windows WDDM watchdog on
203
+ this machine; B ∈ {32, 64, 128} should be swept on a larger GPU.
204
+ 7. **Document-boundary packing.** The current packing strategy places multiple documents in
205
+ one sequence (separated by EOS).
206
+
207
+ ---
208
+
209
+ ## 5. License
210
+
211
+ **Free for personal use; a paid license is required for commercial use.**
212
+
213
+ - ✅ Personal study, research, teaching, hobby projects: **free**, no application required
214
+ - ✅ Academic research with public publication: **free** (please cite the source)
215
+ - 💰 Internal company/studio use, paid API/SaaS, product integration, client deliverables:
216
+ **commercial license required**
217
+
218
+ **Commercial licensing contact: novaweb6868@outlook.com**
219
+
220
+ Full terms: [LICENSE.custom.md](./LICENSE.custom.md).
221
+
222
+ This model is an architectural retrofit of `Qwen/Qwen3-0.6B-Base` (Apache License 2.0).
223
+ This license governs only the newly added portions and does not alter the upstream
224
+ component's original license.
225
+
226
+ ---
227
+
228
+ ## 6. Citation
229
+
230
+ ```bibtex
231
+ @misc{baihu-v1-flash,
232
+ title = {BaiHu-V1-Flash: An SSA (Sparse-attention + SubQ) Retrofit of Qwen3-0.6B-Base},
233
+ author = {NovaAI6868},
234
+ year = {2026},
235
+ url = {https://huggingface.co/NovaAI6868/BaiHu-V1-Flash}
236
+ }
237
+ ```
config.json ADDED
@@ -0,0 +1,79 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "BaiHuSSAForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 151643,
8
+ "dtype": "bfloat16",
9
+ "eos_token_id": 151643,
10
+ "head_dim": 128,
11
+ "hidden_act": "silu",
12
+ "hidden_size": 1024,
13
+ "initializer_range": 0.02,
14
+ "intermediate_size": 3072,
15
+ "layer_types": [
16
+ "full_attention",
17
+ "full_attention",
18
+ "full_attention",
19
+ "full_attention",
20
+ "full_attention",
21
+ "full_attention",
22
+ "full_attention",
23
+ "full_attention",
24
+ "full_attention",
25
+ "full_attention",
26
+ "full_attention",
27
+ "full_attention",
28
+ "full_attention",
29
+ "full_attention",
30
+ "full_attention",
31
+ "full_attention",
32
+ "full_attention",
33
+ "full_attention",
34
+ "full_attention",
35
+ "full_attention",
36
+ "full_attention",
37
+ "full_attention",
38
+ "full_attention",
39
+ "full_attention",
40
+ "full_attention",
41
+ "full_attention",
42
+ "full_attention",
43
+ "full_attention"
44
+ ],
45
+ "max_position_embeddings": 32768,
46
+ "max_window_layers": 28,
47
+ "model_type": "baihu_ssa",
48
+ "num_attention_heads": 16,
49
+ "num_hidden_layers": 28,
50
+ "num_key_value_heads": 8,
51
+ "pad_token_id": null,
52
+ "rms_norm_eps": 1e-06,
53
+ "rope_parameters": {
54
+ "rope_theta": 1000000,
55
+ "rope_type": "default"
56
+ },
57
+ "sliding_window": null,
58
+ "ssa_block_size": 64,
59
+ "ssa_component_init": 0.01,
60
+ "ssa_component_seed": 1234,
61
+ "ssa_compress_dim": 128,
62
+ "ssa_donor_model": "Qwen3-0.6B-Base",
63
+ "ssa_donor_revision": null,
64
+ "ssa_force_full_window": false,
65
+ "ssa_local_blocks": 2,
66
+ "ssa_num_subq_heads": 4,
67
+ "ssa_router_bias_scale": 0.1,
68
+ "ssa_router_dim": 128,
69
+ "ssa_router_init_scale": 0.01,
70
+ "ssa_shared_init_scale": 0.01,
71
+ "ssa_shared_kv": true,
72
+ "ssa_top_k": 8,
73
+ "ssa_use_router_key": true,
74
+ "tie_word_embeddings": true,
75
+ "transformers_version": "5.17.0",
76
+ "use_cache": true,
77
+ "use_sliding_window": false,
78
+ "vocab_size": 151936
79
+ }
generation_config.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 151643,
3
+ "do_sample": false,
4
+ "eos_token_id": 151643,
5
+ "max_new_tokens": 2048,
6
+ "transformers_version": "4.37.0"
7
+ }
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:40f5a9f413e8ce888f0e44e8bf3c18acd9d711cf2d0c76482096d43ec473c4c6
3
+ size 2395267832
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,239 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": false,
3
+ "add_prefix_space": false,
4
+ "added_tokens_decoder": {
5
+ "151643": {
6
+ "content": "<|endoftext|>",
7
+ "lstrip": false,
8
+ "normalized": false,
9
+ "rstrip": false,
10
+ "single_word": false,
11
+ "special": true
12
+ },
13
+ "151644": {
14
+ "content": "<|im_start|>",
15
+ "lstrip": false,
16
+ "normalized": false,
17
+ "rstrip": false,
18
+ "single_word": false,
19
+ "special": true
20
+ },
21
+ "151645": {
22
+ "content": "<|im_end|>",
23
+ "lstrip": false,
24
+ "normalized": false,
25
+ "rstrip": false,
26
+ "single_word": false,
27
+ "special": true
28
+ },
29
+ "151646": {
30
+ "content": "<|object_ref_start|>",
31
+ "lstrip": false,
32
+ "normalized": false,
33
+ "rstrip": false,
34
+ "single_word": false,
35
+ "special": true
36
+ },
37
+ "151647": {
38
+ "content": "<|object_ref_end|>",
39
+ "lstrip": false,
40
+ "normalized": false,
41
+ "rstrip": false,
42
+ "single_word": false,
43
+ "special": true
44
+ },
45
+ "151648": {
46
+ "content": "<|box_start|>",
47
+ "lstrip": false,
48
+ "normalized": false,
49
+ "rstrip": false,
50
+ "single_word": false,
51
+ "special": true
52
+ },
53
+ "151649": {
54
+ "content": "<|box_end|>",
55
+ "lstrip": false,
56
+ "normalized": false,
57
+ "rstrip": false,
58
+ "single_word": false,
59
+ "special": true
60
+ },
61
+ "151650": {
62
+ "content": "<|quad_start|>",
63
+ "lstrip": false,
64
+ "normalized": false,
65
+ "rstrip": false,
66
+ "single_word": false,
67
+ "special": true
68
+ },
69
+ "151651": {
70
+ "content": "<|quad_end|>",
71
+ "lstrip": false,
72
+ "normalized": false,
73
+ "rstrip": false,
74
+ "single_word": false,
75
+ "special": true
76
+ },
77
+ "151652": {
78
+ "content": "<|vision_start|>",
79
+ "lstrip": false,
80
+ "normalized": false,
81
+ "rstrip": false,
82
+ "single_word": false,
83
+ "special": true
84
+ },
85
+ "151653": {
86
+ "content": "<|vision_end|>",
87
+ "lstrip": false,
88
+ "normalized": false,
89
+ "rstrip": false,
90
+ "single_word": false,
91
+ "special": true
92
+ },
93
+ "151654": {
94
+ "content": "<|vision_pad|>",
95
+ "lstrip": false,
96
+ "normalized": false,
97
+ "rstrip": false,
98
+ "single_word": false,
99
+ "special": true
100
+ },
101
+ "151655": {
102
+ "content": "<|image_pad|>",
103
+ "lstrip": false,
104
+ "normalized": false,
105
+ "rstrip": false,
106
+ "single_word": false,
107
+ "special": true
108
+ },
109
+ "151656": {
110
+ "content": "<|video_pad|>",
111
+ "lstrip": false,
112
+ "normalized": false,
113
+ "rstrip": false,
114
+ "single_word": false,
115
+ "special": true
116
+ },
117
+ "151657": {
118
+ "content": "<tool_call>",
119
+ "lstrip": false,
120
+ "normalized": false,
121
+ "rstrip": false,
122
+ "single_word": false,
123
+ "special": false
124
+ },
125
+ "151658": {
126
+ "content": "</tool_call>",
127
+ "lstrip": false,
128
+ "normalized": false,
129
+ "rstrip": false,
130
+ "single_word": false,
131
+ "special": false
132
+ },
133
+ "151659": {
134
+ "content": "<|fim_prefix|>",
135
+ "lstrip": false,
136
+ "normalized": false,
137
+ "rstrip": false,
138
+ "single_word": false,
139
+ "special": false
140
+ },
141
+ "151660": {
142
+ "content": "<|fim_middle|>",
143
+ "lstrip": false,
144
+ "normalized": false,
145
+ "rstrip": false,
146
+ "single_word": false,
147
+ "special": false
148
+ },
149
+ "151661": {
150
+ "content": "<|fim_suffix|>",
151
+ "lstrip": false,
152
+ "normalized": false,
153
+ "rstrip": false,
154
+ "single_word": false,
155
+ "special": false
156
+ },
157
+ "151662": {
158
+ "content": "<|fim_pad|>",
159
+ "lstrip": false,
160
+ "normalized": false,
161
+ "rstrip": false,
162
+ "single_word": false,
163
+ "special": false
164
+ },
165
+ "151663": {
166
+ "content": "<|repo_name|>",
167
+ "lstrip": false,
168
+ "normalized": false,
169
+ "rstrip": false,
170
+ "single_word": false,
171
+ "special": false
172
+ },
173
+ "151664": {
174
+ "content": "<|file_sep|>",
175
+ "lstrip": false,
176
+ "normalized": false,
177
+ "rstrip": false,
178
+ "single_word": false,
179
+ "special": false
180
+ },
181
+ "151665": {
182
+ "content": "<tool_response>",
183
+ "lstrip": false,
184
+ "normalized": false,
185
+ "rstrip": false,
186
+ "single_word": false,
187
+ "special": false
188
+ },
189
+ "151666": {
190
+ "content": "</tool_response>",
191
+ "lstrip": false,
192
+ "normalized": false,
193
+ "rstrip": false,
194
+ "single_word": false,
195
+ "special": false
196
+ },
197
+ "151667": {
198
+ "content": "<think>",
199
+ "lstrip": false,
200
+ "normalized": false,
201
+ "rstrip": false,
202
+ "single_word": false,
203
+ "special": false
204
+ },
205
+ "151668": {
206
+ "content": "</think>",
207
+ "lstrip": false,
208
+ "normalized": false,
209
+ "rstrip": false,
210
+ "single_word": false,
211
+ "special": false
212
+ }
213
+ },
214
+ "additional_special_tokens": [
215
+ "<|im_start|>",
216
+ "<|im_end|>",
217
+ "<|object_ref_start|>",
218
+ "<|object_ref_end|>",
219
+ "<|box_start|>",
220
+ "<|box_end|>",
221
+ "<|quad_start|>",
222
+ "<|quad_end|>",
223
+ "<|vision_start|>",
224
+ "<|vision_end|>",
225
+ "<|vision_pad|>",
226
+ "<|image_pad|>",
227
+ "<|video_pad|>"
228
+ ],
229
+ "bos_token": null,
230
+ "chat_template": "{%- if tools %}\n {{- '<|im_start|>system\\n' }}\n {%- if messages[0].role == 'system' %}\n {{- messages[0].content + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n {%- if messages[0].role == 'system' %}\n {{- '<|im_start|>system\\n' + messages[0].content + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if (message.role == \"user\") or (message.role == \"system\" and not loop.first) %}\n {{- '<|im_start|>' + message.role + '\\n' + message.content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set content = message.content %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is defined and message.reasoning_content is not none %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- else %}\n {%- if '</think>' in message.content %}\n {%- set content = message.content.split('</think>')[-1].lstrip('\\n') %}\n {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\\n').split('<think>')[-1].lstrip('\\n') %}\n {%- endif %}\n {%- endif %}\n {%- if loop.index0 > ns.last_query_index %}\n {%- if loop.last or (not loop.last and reasoning_content) %}\n {{- '<|im_start|>' + message.role + '\\n<think>\\n' + reasoning_content.strip('\\n') + '\\n</think>\\n\\n' + content.lstrip('\\n') }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls %}\n {%- for tool_call in message.tool_calls %}\n {%- if (loop.first and content) or (not loop.first) %}\n {{- '\\n' }}\n {%- endif %}\n {%- if tool_call.function %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {{- '<tool_call>\\n{\"name\": \"' }}\n {{- tool_call.name }}\n {{- '\", \"arguments\": ' }}\n {%- if tool_call.arguments is string %}\n {{- tool_call.arguments }}\n {%- else %}\n {{- tool_call.arguments | tojson }}\n {%- endif %}\n {{- '}\\n</tool_call>' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.first or (messages[loop.index0 - 1].role != \"tool\") %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n<tool_response>\\n' }}\n {{- message.content }}\n {{- '\\n</tool_response>' }}\n {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '<think>\\n\\n</think>\\n\\n' }}\n {%- endif %}\n{%- endif %}",
231
+ "clean_up_tokenization_spaces": false,
232
+ "eos_token": "<|endoftext|>",
233
+ "errors": "replace",
234
+ "model_max_length": 131072,
235
+ "pad_token": "<|endoftext|>",
236
+ "split_special_tokens": false,
237
+ "tokenizer_class": "Qwen2Tokenizer",
238
+ "unk_token": null
239
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff