- HyperS-10M-Base-v1.0
HyperS-10M-Base-v1.0
HyperS-10M-Base-v1.0 is a 9.96M-parameter experimental causal language model built around HyperSpectrum Geometry (HSG) and recurrent hyperspherical state updates rather than Transformer self-attention.
The model is intended as a compact research base model for studying geometry-driven sequence modeling, fixed-size recurrent context representations, adaptive metric learning, and non-Transformer language modeling.
Release status: Base model only. This checkpoint is not instruction-tuned, not a chat model, and is not intended to behave like an assistant.
The public repository intentionally contains only the packaged PyTorch model artifact and this model card. The HyperS implementation, training source, and inference source are not included in this release.
Model Details
Model Description
HyperS processes language through a recurrent geometric state rather than a Transformer KV cache. Tokens are embedded onto a normalized hyperspherical representation, processed through stacked recurrent HyperState layers, and decoded with a geometry-aware HAPD output module.
Each recurrent layer maintains a small collection of learned memory tracks. The model uses an adaptive low-rank metric to change how angular similarity is measured as the recurrent hidden state evolves.
The design goal is to study whether useful linguistic history can be continually compressed into a fixed-size geometric state rather than retaining a growing set of key/value vectors for every previous token.
- Developed by: Sayak Mondal
- Shared by: drelixer
- Model: HyperS-10M-Base-v1.0
- Model type: Decoder-only causal recurrent language model; non-Transformer
- Parameters: 9,957,107
- Language: Primarily English
- Vocabulary: 8,192 tokens
- Tokenizer: HGT-8K v0.1
- Pretraining exposure: 25,000,960 tokens
- Training precision: BF16
- Artifact format: PyTorch
.pt - Instruction tuned: No
- License: Not yet specified
Architecture
| Component | Configuration |
|---|---|
| Vocabulary | 8,192 |
| Hidden dimension | 384 |
| Recurrent layers | 5 |
| Memory tracks per layer | 4 |
| Metric rank | 32 |
| Feed-forward dimension | 768 |
| Decoder rank | 32 |
| Parameters | 9,957,107 |
| Persistent recurrent state | 7,680 scalar values / sequence |
| Observed recurrent-state dtype | FP32 |
| Persistent-state footprint | ~30 KiB / sequence |
HyperS does not use:
- Transformer self-attention
- a Transformer KV cache
- positional embeddings as the mechanism for sequence order
- a per-token retained context representation
Instead, sequence history is carried forward through the recurrent geometric state.
Model Sources
- Model repository: https://huggingface.co/drelixer/HyperS
- Related research paper: HyperSpectrum geometry: a Riemannian hyperspherical framework for curvature-adaptive text classification, Neural Computing and Applications (2026)
The published HyperSpectrum Geometry paper is the conceptual precursor to this project. It introduces the hyperspherical, metric-deformation perspective that motivated HyperS, but it does not describe this language-model architecture or this checkpoint.
Release Artifact
The repository contains:
hypers-10m-base-v1.0.pt
The .pt bundle contains:
- the HyperS
state_dict - the model configuration
- the complete HGT-8K tokenizer configuration
- release metadata
It intentionally excludes:
- optimizer state
- training-state checkpoints
- TBPTT runtime state
- RNG state
- training scripts
- inference scripts
- model source code
The artifact is approximately 40.6 MB as uploaded to the Hugging Face Hub.
Uses
Direct Use
This checkpoint is primarily intended for:
- research on non-Transformer causal language modeling
- recurrent language-model experimentation
- geometric and hyperspherical representation learning research
- studying fixed-size recurrent context representations
- controlled continuation experiments
- architecture analysis and ablation studies
- downstream fine-tuning research
It is a base language model, so raw generations should be expected to behave like unconstrained next-token completion rather than instruction-following responses.
Downstream Use
Potential downstream research directions include:
- continued pretraining
- domain adaptation
- task-specific fine-tuning
- instruction-tuning research
- recurrent-state analysis
- representation-geometry analysis
- long-history compression experiments
- alternative recurrent objectives
- alternative HGT tokenization studies
A compatible implementation of the HyperS architecture is required to execute the weights. This repository does not provide the implementation.
Out-of-Scope Use
The released checkpoint should not be treated as:
- a production assistant
- an instruction-following model
- a factual knowledge system
- a safety-aligned conversational model
- a high-reliability coding model
- a replacement for substantially larger pretrained language models
- a system suitable for medical, legal, financial, or other high-stakes decision making
The checkpoint has only 25M tokens of pretraining exposure and should be understood as a research-scale base model.
Bias, Risks, and Limitations
HyperS-10M-Base-v1.0 inherits limitations from both its training data and its small pretraining budget.
Important limitations include:
- The model is small at 9.96M parameters.
- Pretraining stopped at 25,000,960 exposed tokens, far below the scale normally used for modern general-purpose language models.
- The model is not instruction-tuned.
- Generated text can be repetitive, incoherent, factually incorrect, or unrelated to the prompt.
- The model may reproduce biases, stereotypes, inaccuracies, or undesirable patterns present in its training sources.
- Exact long-range retrieval has not been demonstrated.
- A synthetic long-range digit-recall benchmark was not learned by either HyperS or the matched Transformer control and therefore was inconclusive as a long-memory test.
- A public inference implementation is not included in this repository.
- Hugging Face
transformers.AutoModel/AutoModelForCausalLMcompatibility is not provided. - The current implementation prioritizes the fixed-state research design rather than inference throughput.
Recommendations
Use this checkpoint as an experimental research artifact rather than a production model. Claims about long-context behavior should distinguish between predictive utility of recurrent history and exact retrieval of earlier tokens.
The model's outputs should be independently verified before being used for any factual or consequential purpose.
How to Get Started with the Model
This is a weights-only research release.
Download hypers-10m-base-v1.0.pt from the repository and load it using a compatible implementation of the HyperS architecture and HGT tokenizer.
The serialized bundle includes the model configuration and HGT-8K tokenizer data so that these do not need to be distributed as separate repository files.
No model or inference source code is distributed in this repository.
Training Details
Training Data
The pretraining stream was constructed from:
The encoded training corpus contained approximately 2.046B stored HGT tokens across the prepared sources. HyperS-10M-Base-v1.0 was exposed to 25,000,960 tokens from this stream; it was not trained over the entire encoded corpus.
TinyStories consists of synthetically generated short stories with a constrained vocabulary. FineWeb-Edu is a large collection of educational web text derived from FineWeb.
Users should consult the original dataset cards for dataset-specific provenance, licenses, filtering procedures, and limitations.
Preprocessing
Text was encoded using HGT-8K v0.1, a custom 8,192-token tokenizer with UTF-8 byte fallback.
Special token IDs:
| Token | ID |
|---|---|
<PAD> |
0 |
<BOS> |
1 |
<EOS> |
2 |
<DOC> |
3 |
| Byte fallback | 4β259 |
| Learned tokens | 260β8191 |
Documents were framed as:
<BOS><DOC> content <EOS>
Training Procedure
HyperS was trained with causal next-token prediction while recurrent states were carried between short TBPTT chunks and detached at chunk boundaries.
Core training configuration:
| Hyperparameter | Value |
|---|---|
| Batch size | 64 |
| TBPTT length | 16 |
| Tokens / optimizer update | 1,024 |
| Optimizer | AdamW |
| Adam Ξ²1 | 0.9 |
| Adam Ξ²2 | 0.95 |
| Peak learning rate | 3e-4 |
| Minimum LR ratio | 0.1 |
| Weight decay | 0.1 |
| Gradient clipping | 1.0 |
| Precision | BF16 |
| LR schedule horizon | 200M tokens |
| Warmup ratio | 1% |
| Final exposed tokens | 25,000,960 |
The recurrent state was preserved between adjacent stream chunks during training while gradients were truncated at the TBPTT boundary.
Speeds, Sizes, Times
- Training/evaluation GPU: NVIDIA GeForce RTX 4070 Laptop GPU
- Public
.ptartifact: approximately 40.6 MB - Model parameters: 9,957,107
- Persistent recurrent state: 7,680 FP32 values = approximately 30 KiB per sequence
A complete carbon-emissions estimate was not recorded.
Evaluation
The evaluations below characterize this 25M-token base checkpoint. They should not be interpreted as general-purpose language-model benchmark results.
Causal-LM Evaluation
A matched stateless evaluation was performed with sequence length 16.
| Evaluation split | Targets | Cross-entropy | Perplexity |
|---|---|---|---|
| General validation | 65,536 | 3.594227 | 36.3876 |
| TinyStories validation | 65,536 | 3.413734 | 30.3785 |
Matched Transformer Control
A parameter-matched Transformer control contained 9,969,792 parameters, only 12,685 more than HyperS. Both checkpoints were trained to the same 25,000,960-token exposure under the matched short-context protocol.
On the same stateless T=16 evaluation:
| Model | General CE | General PPL | TinyStories CE | TinyStories PPL |
|---|---|---|---|---|
| HyperS-10M | 3.594227 | 36.3876 | 3.413734 | 30.3785 |
| Transformer-10M control | 3.247609 | 25.7288 | 2.912044 | 18.3944 |
Under this specific matched stateless short-context evaluation, the Transformer control achieved lower cross-entropy.
This comparison should not be generalized beyond the evaluated protocol.
Recurrent-History Utility
HyperS was also evaluated on natural validation sequences while varying how much earlier history was allowed to enter the recurrent state before measuring the same 16-token future.
Selected results:
| Available history | Future-token CE |
|---|---|
| 0 | 3.597356 |
| 16 | 3.378207 |
| 64 | 3.352380 |
| 128 | 3.348234 |
| 256 | 3.349865 |
| 512 | 3.351987 |
| 1,024 | 3.357652 |
The results show that recurrent history provided predictive utility on natural language, with the point estimate saturating around 64β128 history tokens in this experiment.
A separate history-specificity test found that replacing earlier same-document history with unrelated natural-language history increasingly harmed prediction as the history window grew. This supports the conclusion that the fixed recurrent state carries document-specific information.
These tests do not establish exact retrieval of a token 1,024 positions earlier.
Generation Evaluation
On a 14-prompt sampled generation suite using temperature 0.8 and top-k 50:
| Metric | HyperS-10M |
|---|---|
| Generated tokens | 1,851 |
| EOS reached | 2 / 14 |
| Valid UTF-8 | 14 / 14 |
| Bigram repetition | 0.1698 |
| Trigram repetition | 0.0703 |
| Distinct-token ratio | 0.4942 |
| Generation throughput | 41.36 tok/s |
These aggregate statistics measure generation mechanics and repetition, not semantic answer quality.
Inference-State Scaling
HyperS retains a fixed recurrent state of:
5 layers Γ 4 tracks Γ 384 dimensions = 7,680 scalar values
In the evaluated implementation the retained recurrent state was FP32, corresponding to approximately 30 KiB, independent of processed context length.
For comparison, a conventional 5-layer, 384-dimensional multi-head Transformer KV cache requires:
2 Γ 5 Γ 384 = 3,840 cached scalar values per retained token
At 2,048 tokens and BF16 this corresponds to approximately 15 MiB, a 512Γ byte ratio relative to the measured HyperS recurrent state.
This comparison concerns retained context state, not total peak CUDA allocation. Temporary activations and workspace memory are separate.
The current HyperS implementation is substantially slower than the matched Transformer implementation. The fixed-state design should therefore be understood as a memory/representation trade-off, not a demonstrated speed advantage.
Instruction-Tuning Status
An instruction-tuned model is not included in this release.
Exploratory instruction-tuning experiments showed that the 25M-token base could learn some simple prompt-conditioned mappings, but the resulting checkpoints did not meet the project's release threshold for reliable autoregressive instruction following, copying, and context lookup.
Those experimental checkpoints are not part of the public model release.
Model Examination
The core experimental observation behind the release is that HyperS can maintain a constant-size recurrent state while still showing measurable, document-specific predictive benefit from earlier natural-language history.
This is evidence for history compression into the recurrent state. It should not be interpreted as evidence of perfect memory, exact long-range retrieval, or superiority over attention-based language models.
Environmental Impact
A formal carbon-footprint measurement was not performed.
- Hardware: NVIDIA GeForce RTX 4070 Laptop GPU
- Cloud provider: None / local hardware
- Compute region: Not reported
- Carbon emitted: Not measured
No numerical emissions estimate is provided because the necessary power-consumption and carbon-intensity measurements were not recorded.
Technical Specifications
Model Architecture and Objective
HyperS is a decoder-only recurrent causal language model built around adaptive hyperspherical geometry.
Its primary components are:
- HGT tokenizer β fixed 8K vocabulary with UTF-8 byte fallback.
- HyperEmbedding β token representations normalized onto a hyperspherical space.
- Adaptive metric β low-rank state-conditioned metric deformation.
- HyperState β recurrent spherical memory with four tracks per layer.
- Geometric state updates β recurrent updates constructed around hyperspherical/geodesic operations.
- HAPD decoder β geometry-aware token decoding using tied token prototypes.
- Causal next-token objective β standard language-model cross-entropy.
The architecture contains no Transformer self-attention or Transformer KV cache.
Compute Infrastructure
Hardware
- NVIDIA GeForce RTX 4070 Laptop GPU
Software
- PyTorch
- CUDA-capable NVIDIA environment
- Custom HyperS model implementation
- Custom HGT tokenizer implementation
Exact source code is intentionally not part of this weights-only release.
Citation
If you use this model artifact directly, cite the Hugging Face repository:
@misc{mondal2026hypers10m,
author = {Sayak Mondal},
title = {HyperS-10M-Base-v1.0},
year = {2026},
howpublished = {Hugging Face model repository},
url = {https://huggingface.co/drelixer/HyperS}
}
For the geometric framework that motivated HyperS:
@article{mondal2026hyperspectrum,
author = {Sayak Mondal},
title = {HyperSpectrum geometry: a Riemannian hyperspherical framework for curvature-adaptive text classification},
journal = {Neural Computing and Applications},
year = {2026},
publisher = {Springer Nature},
doi = {10.1007/s00521-026-12332-4}
}
Glossary
- HSG: HyperSpectrum Geometry.
- HGT: Hyper-Geometric Tokenizer used by HyperS.
- HAPD: HyperS geometry-aware prediction/decoding component.
- HyperState: Recurrent geometric memory layer.
- TBPTT: Truncated backpropagation through time.
- Persistent recurrent state: The state carried between processed chunks; for this model it contains 7,680 scalar values per sequence.
- Stateless T=16 evaluation: Evaluation in which each 16-token segment begins without earlier recurrent history.
More Information
- Hugging Face: https://huggingface.co/drelixer/HyperS
- Related paper: https://doi.org/10.1007/s00521-026-12332-4
- Paper venue: Neural Computing and Applications, Springer Nature, 2026
HyperS is an ongoing research project. Future work includes larger-scale pretraining, improved recurrent-state optimization, value-binding and copying mechanisms, instruction tuning, and expanded evaluation.
Model Card Authors
Sayak Mondal
Model Card Contact
For questions, issues, or research discussions, use the Hugging Face repository community/discussion interface associated with drelixer/HyperS.