File size: 17,810 Bytes
f94f56b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
---
model_name: HyperS-10M-Base-v1.0
language:
  - en
pipeline_tag: text-generation
library_name: pytorch
datasets:
  - roneneldan/TinyStories
  - HuggingFaceFW/fineweb-edu
tags:
  - hypers
  - language-model
  - causal-language-model
  - recurrent-language-model
  - non-transformer
  - pytorch
  - geometry
  - hyperspherical
  - research
---

# HyperS-10M-Base-v1.0

**HyperS-10M-Base-v1.0** is a 9.96M-parameter experimental causal language model built around **HyperSpectrum Geometry (HSG)** and recurrent hyperspherical state updates rather than Transformer self-attention.

The model is intended as a compact research base model for studying geometry-driven sequence modeling, fixed-size recurrent context representations, adaptive metric learning, and non-Transformer language modeling.

> **Release status:** Base model only. This checkpoint is **not instruction-tuned**, **not a chat model**, and is not intended to behave like an assistant.

The public repository intentionally contains only the packaged PyTorch model artifact and this model card. The HyperS implementation, training source, and inference source are not included in this release.

## Model Details

### Model Description

HyperS processes language through a recurrent geometric state rather than a Transformer KV cache. Tokens are embedded onto a normalized hyperspherical representation, processed through stacked recurrent `HyperState` layers, and decoded with a geometry-aware HAPD output module.

Each recurrent layer maintains a small collection of learned memory tracks. The model uses an adaptive low-rank metric to change how angular similarity is measured as the recurrent hidden state evolves.

The design goal is to study whether useful linguistic history can be continually compressed into a fixed-size geometric state rather than retaining a growing set of key/value vectors for every previous token.

- **Developed by:** Sayak Mondal
- **Shared by:** [drelixer](https://huggingface.co/drelixer)
- **Model:** HyperS-10M-Base-v1.0
- **Model type:** Decoder-only causal recurrent language model; non-Transformer
- **Parameters:** 9,957,107
- **Language:** Primarily English
- **Vocabulary:** 8,192 tokens
- **Tokenizer:** HGT-8K v0.1
- **Pretraining exposure:** 25,000,960 tokens
- **Training precision:** BF16
- **Artifact format:** PyTorch `.pt`
- **Instruction tuned:** No
- **License:** Not yet specified

### Architecture

| Component | Configuration |
|---|---:|
| Vocabulary | 8,192 |
| Hidden dimension | 384 |
| Recurrent layers | 5 |
| Memory tracks per layer | 4 |
| Metric rank | 32 |
| Feed-forward dimension | 768 |
| Decoder rank | 32 |
| Parameters | 9,957,107 |
| Persistent recurrent state | 7,680 scalar values / sequence |
| Observed recurrent-state dtype | FP32 |
| Persistent-state footprint | ~30 KiB / sequence |

HyperS does **not** use:

- Transformer self-attention
- a Transformer KV cache
- positional embeddings as the mechanism for sequence order
- a per-token retained context representation

Instead, sequence history is carried forward through the recurrent geometric state.

### Model Sources

- **Model repository:** https://huggingface.co/drelixer/HyperS
- **Related research paper:** [HyperSpectrum geometry: a Riemannian hyperspherical framework for curvature-adaptive text classification](https://doi.org/10.1007/s00521-026-12332-4), *Neural Computing and Applications* (2026)

The published HyperSpectrum Geometry paper is the conceptual precursor to this project. It introduces the hyperspherical, metric-deformation perspective that motivated HyperS, but it does **not** describe this language-model architecture or this checkpoint.

## Release Artifact

The repository contains:

```text
hypers-10m-base-v1.0.pt
```

The `.pt` bundle contains:

- the HyperS `state_dict`
- the model configuration
- the complete HGT-8K tokenizer configuration
- release metadata

It intentionally excludes:

- optimizer state
- training-state checkpoints
- TBPTT runtime state
- RNG state
- training scripts
- inference scripts
- model source code

The artifact is approximately **40.6 MB** as uploaded to the Hugging Face Hub.

## Uses

### Direct Use

This checkpoint is primarily intended for:

- research on non-Transformer causal language modeling
- recurrent language-model experimentation
- geometric and hyperspherical representation learning research
- studying fixed-size recurrent context representations
- controlled continuation experiments
- architecture analysis and ablation studies
- downstream fine-tuning research

It is a **base language model**, so raw generations should be expected to behave like unconstrained next-token completion rather than instruction-following responses.

### Downstream Use

Potential downstream research directions include:

- continued pretraining
- domain adaptation
- task-specific fine-tuning
- instruction-tuning research
- recurrent-state analysis
- representation-geometry analysis
- long-history compression experiments
- alternative recurrent objectives
- alternative HGT tokenization studies

A compatible implementation of the HyperS architecture is required to execute the weights. This repository does not provide the implementation.

### Out-of-Scope Use

The released checkpoint should not be treated as:

- a production assistant
- an instruction-following model
- a factual knowledge system
- a safety-aligned conversational model
- a high-reliability coding model
- a replacement for substantially larger pretrained language models
- a system suitable for medical, legal, financial, or other high-stakes decision making

The checkpoint has only 25M tokens of pretraining exposure and should be understood as a research-scale base model.

## Bias, Risks, and Limitations

HyperS-10M-Base-v1.0 inherits limitations from both its training data and its small pretraining budget.

Important limitations include:

- The model is small at 9.96M parameters.
- Pretraining stopped at 25,000,960 exposed tokens, far below the scale normally used for modern general-purpose language models.
- The model is not instruction-tuned.
- Generated text can be repetitive, incoherent, factually incorrect, or unrelated to the prompt.
- The model may reproduce biases, stereotypes, inaccuracies, or undesirable patterns present in its training sources.
- Exact long-range retrieval has **not** been demonstrated.
- A synthetic long-range digit-recall benchmark was not learned by either HyperS or the matched Transformer control and therefore was inconclusive as a long-memory test.
- A public inference implementation is not included in this repository.
- Hugging Face `transformers.AutoModel` / `AutoModelForCausalLM` compatibility is not provided.
- The current implementation prioritizes the fixed-state research design rather than inference throughput.

### Recommendations

Use this checkpoint as an experimental research artifact rather than a production model. Claims about long-context behavior should distinguish between **predictive utility of recurrent history** and **exact retrieval of earlier tokens**.

The model's outputs should be independently verified before being used for any factual or consequential purpose.

## How to Get Started with the Model

This is a **weights-only research release**.

Download `hypers-10m-base-v1.0.pt` from the repository and load it using a compatible implementation of the HyperS architecture and HGT tokenizer.

The serialized bundle includes the model configuration and HGT-8K tokenizer data so that these do not need to be distributed as separate repository files.

No model or inference source code is distributed in this repository.

## Training Details

### Training Data

The pretraining stream was constructed from:

- [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories)
- [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu)

The encoded training corpus contained approximately **2.046B stored HGT tokens** across the prepared sources. HyperS-10M-Base-v1.0 was exposed to **25,000,960 tokens** from this stream; it was not trained over the entire encoded corpus.

TinyStories consists of synthetically generated short stories with a constrained vocabulary. FineWeb-Edu is a large collection of educational web text derived from FineWeb.

Users should consult the original dataset cards for dataset-specific provenance, licenses, filtering procedures, and limitations.

### Preprocessing

Text was encoded using **HGT-8K v0.1**, a custom 8,192-token tokenizer with UTF-8 byte fallback.

Special token IDs:

| Token | ID |
|---|---:|
| `<PAD>` | 0 |
| `<BOS>` | 1 |
| `<EOS>` | 2 |
| `<DOC>` | 3 |
| Byte fallback | 4–259 |
| Learned tokens | 260–8191 |

Documents were framed as:

```text
<BOS><DOC> content <EOS>
```

### Training Procedure

HyperS was trained with causal next-token prediction while recurrent states were carried between short TBPTT chunks and detached at chunk boundaries.

Core training configuration:

| Hyperparameter | Value |
|---|---:|
| Batch size | 64 |
| TBPTT length | 16 |
| Tokens / optimizer update | 1,024 |
| Optimizer | AdamW |
| Adam β1 | 0.9 |
| Adam β2 | 0.95 |
| Peak learning rate | 3e-4 |
| Minimum LR ratio | 0.1 |
| Weight decay | 0.1 |
| Gradient clipping | 1.0 |
| Precision | BF16 |
| LR schedule horizon | 200M tokens |
| Warmup ratio | 1% |
| Final exposed tokens | 25,000,960 |

The recurrent state was preserved between adjacent stream chunks during training while gradients were truncated at the TBPTT boundary.

### Speeds, Sizes, Times

- **Training/evaluation GPU:** NVIDIA GeForce RTX 4070 Laptop GPU
- **Public `.pt` artifact:** approximately 40.6 MB
- **Model parameters:** 9,957,107
- **Persistent recurrent state:** 7,680 FP32 values = approximately 30 KiB per sequence

A complete carbon-emissions estimate was not recorded.

## Evaluation

The evaluations below characterize this **25M-token base checkpoint**. They should not be interpreted as general-purpose language-model benchmark results.

### Causal-LM Evaluation

A matched stateless evaluation was performed with sequence length 16.

| Evaluation split | Targets | Cross-entropy | Perplexity |
|---|---:|---:|---:|
| General validation | 65,536 | 3.594227 | 36.3876 |
| TinyStories validation | 65,536 | 3.413734 | 30.3785 |

### Matched Transformer Control

A parameter-matched Transformer control contained **9,969,792 parameters**, only 12,685 more than HyperS. Both checkpoints were trained to the same 25,000,960-token exposure under the matched short-context protocol.

On the same stateless T=16 evaluation:

| Model | General CE | General PPL | TinyStories CE | TinyStories PPL |
|---|---:|---:|---:|---:|
| HyperS-10M | 3.594227 | 36.3876 | 3.413734 | 30.3785 |
| Transformer-10M control | 3.247609 | 25.7288 | 2.912044 | 18.3944 |

Under this specific matched stateless short-context evaluation, the Transformer control achieved lower cross-entropy.

This comparison should not be generalized beyond the evaluated protocol.

### Recurrent-History Utility

HyperS was also evaluated on natural validation sequences while varying how much earlier history was allowed to enter the recurrent state before measuring the same 16-token future.

Selected results:

| Available history | Future-token CE |
|---:|---:|
| 0 | 3.597356 |
| 16 | 3.378207 |
| 64 | 3.352380 |
| 128 | **3.348234** |
| 256 | 3.349865 |
| 512 | 3.351987 |
| 1,024 | 3.357652 |

The results show that recurrent history provided predictive utility on natural language, with the point estimate saturating around 64–128 history tokens in this experiment.

A separate history-specificity test found that replacing earlier same-document history with unrelated natural-language history increasingly harmed prediction as the history window grew. This supports the conclusion that the fixed recurrent state carries document-specific information.

These tests **do not establish exact retrieval of a token 1,024 positions earlier**.

### Generation Evaluation

On a 14-prompt sampled generation suite using temperature 0.8 and top-k 50:

| Metric | HyperS-10M |
|---|---:|
| Generated tokens | 1,851 |
| EOS reached | 2 / 14 |
| Valid UTF-8 | 14 / 14 |
| Bigram repetition | 0.1698 |
| Trigram repetition | 0.0703 |
| Distinct-token ratio | 0.4942 |
| Generation throughput | 41.36 tok/s |

These aggregate statistics measure generation mechanics and repetition, not semantic answer quality.

### Inference-State Scaling

HyperS retains a fixed recurrent state of:

```text
5 layers × 4 tracks × 384 dimensions = 7,680 scalar values
```

In the evaluated implementation the retained recurrent state was FP32, corresponding to approximately **30 KiB**, independent of processed context length.

For comparison, a conventional 5-layer, 384-dimensional multi-head Transformer KV cache requires:

```text
2 × 5 × 384 = 3,840 cached scalar values per retained token
```

At 2,048 tokens and BF16 this corresponds to approximately **15 MiB**, a **512× byte ratio** relative to the measured HyperS recurrent state.

This comparison concerns **retained context state**, not total peak CUDA allocation. Temporary activations and workspace memory are separate.

The current HyperS implementation is substantially slower than the matched Transformer implementation. The fixed-state design should therefore be understood as a memory/representation trade-off, not a demonstrated speed advantage.

## Instruction-Tuning Status

An instruction-tuned model is **not included in this release**.

Exploratory instruction-tuning experiments showed that the 25M-token base could learn some simple prompt-conditioned mappings, but the resulting checkpoints did not meet the project's release threshold for reliable autoregressive instruction following, copying, and context lookup.

Those experimental checkpoints are not part of the public model release.

## Model Examination

The core experimental observation behind the release is that HyperS can maintain a constant-size recurrent state while still showing measurable, document-specific predictive benefit from earlier natural-language history.

This is evidence for **history compression into the recurrent state**. It should not be interpreted as evidence of perfect memory, exact long-range retrieval, or superiority over attention-based language models.

## Environmental Impact

A formal carbon-footprint measurement was not performed.

- **Hardware:** NVIDIA GeForce RTX 4070 Laptop GPU
- **Cloud provider:** None / local hardware
- **Compute region:** Not reported
- **Carbon emitted:** Not measured

No numerical emissions estimate is provided because the necessary power-consumption and carbon-intensity measurements were not recorded.

## Technical Specifications

### Model Architecture and Objective

HyperS is a decoder-only recurrent causal language model built around adaptive hyperspherical geometry.

Its primary components are:

1. **HGT tokenizer** — fixed 8K vocabulary with UTF-8 byte fallback.
2. **HyperEmbedding** — token representations normalized onto a hyperspherical space.
3. **Adaptive metric** — low-rank state-conditioned metric deformation.
4. **HyperState** — recurrent spherical memory with four tracks per layer.
5. **Geometric state updates** — recurrent updates constructed around hyperspherical/geodesic operations.
6. **HAPD decoder** — geometry-aware token decoding using tied token prototypes.
7. **Causal next-token objective** — standard language-model cross-entropy.

The architecture contains no Transformer self-attention or Transformer KV cache.

### Compute Infrastructure

#### Hardware

- NVIDIA GeForce RTX 4070 Laptop GPU

#### Software

- PyTorch
- CUDA-capable NVIDIA environment
- Custom HyperS model implementation
- Custom HGT tokenizer implementation

Exact source code is intentionally not part of this weights-only release.

## Citation

If you use this model artifact directly, cite the Hugging Face repository:

```bibtex
@misc{mondal2026hypers10m,
  author       = {Sayak Mondal},
  title        = {HyperS-10M-Base-v1.0},
  year         = {2026},
  howpublished = {Hugging Face model repository},
  url          = {https://huggingface.co/drelixer/HyperS}
}
```

For the geometric framework that motivated HyperS:

```bibtex
@article{mondal2026hyperspectrum,
  author    = {Sayak Mondal},
  title     = {HyperSpectrum geometry: a Riemannian hyperspherical framework for curvature-adaptive text classification},
  journal   = {Neural Computing and Applications},
  year      = {2026},
  publisher = {Springer Nature},
  doi       = {10.1007/s00521-026-12332-4}
}
```

## Glossary

- **HSG:** HyperSpectrum Geometry.
- **HGT:** Hyper-Geometric Tokenizer used by HyperS.
- **HAPD:** HyperS geometry-aware prediction/decoding component.
- **HyperState:** Recurrent geometric memory layer.
- **TBPTT:** Truncated backpropagation through time.
- **Persistent recurrent state:** The state carried between processed chunks; for this model it contains 7,680 scalar values per sequence.
- **Stateless T=16 evaluation:** Evaluation in which each 16-token segment begins without earlier recurrent history.

## More Information

- **Hugging Face:** https://huggingface.co/drelixer/HyperS
- **Related paper:** https://doi.org/10.1007/s00521-026-12332-4
- **Paper venue:** *Neural Computing and Applications*, Springer Nature, 2026

HyperS is an ongoing research project. Future work includes larger-scale pretraining, improved recurrent-state optimization, value-binding and copying mechanisms, instruction tuning, and expanded evaluation.

## Model Card Authors

**Sayak Mondal**

## Model Card Contact

For questions, issues, or research discussions, use the Hugging Face repository community/discussion interface associated with `drelixer/HyperS`.