|
Download README.md from drelixer/HyperS: direct link, hf CLI and curl.
- Browser
- Download file 17.8 kB
-
https://huggingface.co/drelixer/HyperS/resolve/main/README.md
- Command line
-
hf download hf://drelixer/HyperS/README.md
-
curl -L -o README.md https://huggingface.co/drelixer/HyperS/resolve/main/README.md
17.8 kB
| model_name: HyperS-10M-Base-v1.0 | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| library_name: pytorch | |
| datasets: | |
| - roneneldan/TinyStories | |
| - HuggingFaceFW/fineweb-edu | |
| tags: | |
| - hypers | |
| - language-model | |
| - causal-language-model | |
| - recurrent-language-model | |
| - non-transformer | |
| - pytorch | |
| - geometry | |
| - hyperspherical | |
| - research | |
| # HyperS-10M-Base-v1.0 | |
| **HyperS-10M-Base-v1.0** is a 9.96M-parameter experimental causal language model built around **HyperSpectrum Geometry (HSG)** and recurrent hyperspherical state updates rather than Transformer self-attention. | |
| The model is intended as a compact research base model for studying geometry-driven sequence modeling, fixed-size recurrent context representations, adaptive metric learning, and non-Transformer language modeling. | |
| > **Release status:** Base model only. This checkpoint is **not instruction-tuned**, **not a chat model**, and is not intended to behave like an assistant. | |
| The public repository intentionally contains only the packaged PyTorch model artifact and this model card. The HyperS implementation, training source, and inference source are not included in this release. | |
| ## Model Details | |
| ### Model Description | |
| HyperS processes language through a recurrent geometric state rather than a Transformer KV cache. Tokens are embedded onto a normalized hyperspherical representation, processed through stacked recurrent `HyperState` layers, and decoded with a geometry-aware HAPD output module. | |
| Each recurrent layer maintains a small collection of learned memory tracks. The model uses an adaptive low-rank metric to change how angular similarity is measured as the recurrent hidden state evolves. | |
| The design goal is to study whether useful linguistic history can be continually compressed into a fixed-size geometric state rather than retaining a growing set of key/value vectors for every previous token. | |
| - **Developed by:** Sayak Mondal | |
| - **Shared by:** [drelixer](https://huggingface.co/drelixer) | |
| - **Model:** HyperS-10M-Base-v1.0 | |
| - **Model type:** Decoder-only causal recurrent language model; non-Transformer | |
| - **Parameters:** 9,957,107 | |
| - **Language:** Primarily English | |
| - **Vocabulary:** 8,192 tokens | |
| - **Tokenizer:** HGT-8K v0.1 | |
| - **Pretraining exposure:** 25,000,960 tokens | |
| - **Training precision:** BF16 | |
| - **Artifact format:** PyTorch `.pt` | |
| - **Instruction tuned:** No | |
| - **License:** Not yet specified | |
| ### Architecture | |
| | Component | Configuration | | |
| |---|---:| | |
| | Vocabulary | 8,192 | | |
| | Hidden dimension | 384 | | |
| | Recurrent layers | 5 | | |
| | Memory tracks per layer | 4 | | |
| | Metric rank | 32 | | |
| | Feed-forward dimension | 768 | | |
| | Decoder rank | 32 | | |
| | Parameters | 9,957,107 | | |
| | Persistent recurrent state | 7,680 scalar values / sequence | | |
| | Observed recurrent-state dtype | FP32 | | |
| | Persistent-state footprint | ~30 KiB / sequence | | |
| HyperS does **not** use: | |
| - Transformer self-attention | |
| - a Transformer KV cache | |
| - positional embeddings as the mechanism for sequence order | |
| - a per-token retained context representation | |
| Instead, sequence history is carried forward through the recurrent geometric state. | |
| ### Model Sources | |
| - **Model repository:** https://huggingface.co/drelixer/HyperS | |
| - **Related research paper:** [HyperSpectrum geometry: a Riemannian hyperspherical framework for curvature-adaptive text classification](https://doi.org/10.1007/s00521-026-12332-4), *Neural Computing and Applications* (2026) | |
| The published HyperSpectrum Geometry paper is the conceptual precursor to this project. It introduces the hyperspherical, metric-deformation perspective that motivated HyperS, but it does **not** describe this language-model architecture or this checkpoint. | |
| ## Release Artifact | |
| The repository contains: | |
| ```text | |
| hypers-10m-base-v1.0.pt | |
| ``` | |
| The `.pt` bundle contains: | |
| - the HyperS `state_dict` | |
| - the model configuration | |
| - the complete HGT-8K tokenizer configuration | |
| - release metadata | |
| It intentionally excludes: | |
| - optimizer state | |
| - training-state checkpoints | |
| - TBPTT runtime state | |
| - RNG state | |
| - training scripts | |
| - inference scripts | |
| - model source code | |
| The artifact is approximately **40.6 MB** as uploaded to the Hugging Face Hub. | |
| ## Uses | |
| ### Direct Use | |
| This checkpoint is primarily intended for: | |
| - research on non-Transformer causal language modeling | |
| - recurrent language-model experimentation | |
| - geometric and hyperspherical representation learning research | |
| - studying fixed-size recurrent context representations | |
| - controlled continuation experiments | |
| - architecture analysis and ablation studies | |
| - downstream fine-tuning research | |
| It is a **base language model**, so raw generations should be expected to behave like unconstrained next-token completion rather than instruction-following responses. | |
| ### Downstream Use | |
| Potential downstream research directions include: | |
| - continued pretraining | |
| - domain adaptation | |
| - task-specific fine-tuning | |
| - instruction-tuning research | |
| - recurrent-state analysis | |
| - representation-geometry analysis | |
| - long-history compression experiments | |
| - alternative recurrent objectives | |
| - alternative HGT tokenization studies | |
| A compatible implementation of the HyperS architecture is required to execute the weights. This repository does not provide the implementation. | |
| ### Out-of-Scope Use | |
| The released checkpoint should not be treated as: | |
| - a production assistant | |
| - an instruction-following model | |
| - a factual knowledge system | |
| - a safety-aligned conversational model | |
| - a high-reliability coding model | |
| - a replacement for substantially larger pretrained language models | |
| - a system suitable for medical, legal, financial, or other high-stakes decision making | |
| The checkpoint has only 25M tokens of pretraining exposure and should be understood as a research-scale base model. | |
| ## Bias, Risks, and Limitations | |
| HyperS-10M-Base-v1.0 inherits limitations from both its training data and its small pretraining budget. | |
| Important limitations include: | |
| - The model is small at 9.96M parameters. | |
| - Pretraining stopped at 25,000,960 exposed tokens, far below the scale normally used for modern general-purpose language models. | |
| - The model is not instruction-tuned. | |
| - Generated text can be repetitive, incoherent, factually incorrect, or unrelated to the prompt. | |
| - The model may reproduce biases, stereotypes, inaccuracies, or undesirable patterns present in its training sources. | |
| - Exact long-range retrieval has **not** been demonstrated. | |
| - A synthetic long-range digit-recall benchmark was not learned by either HyperS or the matched Transformer control and therefore was inconclusive as a long-memory test. | |
| - A public inference implementation is not included in this repository. | |
| - Hugging Face `transformers.AutoModel` / `AutoModelForCausalLM` compatibility is not provided. | |
| - The current implementation prioritizes the fixed-state research design rather than inference throughput. | |
| ### Recommendations | |
| Use this checkpoint as an experimental research artifact rather than a production model. Claims about long-context behavior should distinguish between **predictive utility of recurrent history** and **exact retrieval of earlier tokens**. | |
| The model's outputs should be independently verified before being used for any factual or consequential purpose. | |
| ## How to Get Started with the Model | |
| This is a **weights-only research release**. | |
| Download `hypers-10m-base-v1.0.pt` from the repository and load it using a compatible implementation of the HyperS architecture and HGT tokenizer. | |
| The serialized bundle includes the model configuration and HGT-8K tokenizer data so that these do not need to be distributed as separate repository files. | |
| No model or inference source code is distributed in this repository. | |
| ## Training Details | |
| ### Training Data | |
| The pretraining stream was constructed from: | |
| - [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) | |
| - [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) | |
| The encoded training corpus contained approximately **2.046B stored HGT tokens** across the prepared sources. HyperS-10M-Base-v1.0 was exposed to **25,000,960 tokens** from this stream; it was not trained over the entire encoded corpus. | |
| TinyStories consists of synthetically generated short stories with a constrained vocabulary. FineWeb-Edu is a large collection of educational web text derived from FineWeb. | |
| Users should consult the original dataset cards for dataset-specific provenance, licenses, filtering procedures, and limitations. | |
| ### Preprocessing | |
| Text was encoded using **HGT-8K v0.1**, a custom 8,192-token tokenizer with UTF-8 byte fallback. | |
| Special token IDs: | |
| | Token | ID | | |
| |---|---:| | |
| | `<PAD>` | 0 | | |
| | `<BOS>` | 1 | | |
| | `<EOS>` | 2 | | |
| | `<DOC>` | 3 | | |
| | Byte fallback | 4–259 | | |
| | Learned tokens | 260–8191 | | |
| Documents were framed as: | |
| ```text | |
| <BOS><DOC> content <EOS> | |
| ``` | |
| ### Training Procedure | |
| HyperS was trained with causal next-token prediction while recurrent states were carried between short TBPTT chunks and detached at chunk boundaries. | |
| Core training configuration: | |
| | Hyperparameter | Value | | |
| |---|---:| | |
| | Batch size | 64 | | |
| | TBPTT length | 16 | | |
| | Tokens / optimizer update | 1,024 | | |
| | Optimizer | AdamW | | |
| | Adam β1 | 0.9 | | |
| | Adam β2 | 0.95 | | |
| | Peak learning rate | 3e-4 | | |
| | Minimum LR ratio | 0.1 | | |
| | Weight decay | 0.1 | | |
| | Gradient clipping | 1.0 | | |
| | Precision | BF16 | | |
| | LR schedule horizon | 200M tokens | | |
| | Warmup ratio | 1% | | |
| | Final exposed tokens | 25,000,960 | | |
| The recurrent state was preserved between adjacent stream chunks during training while gradients were truncated at the TBPTT boundary. | |
| ### Speeds, Sizes, Times | |
| - **Training/evaluation GPU:** NVIDIA GeForce RTX 4070 Laptop GPU | |
| - **Public `.pt` artifact:** approximately 40.6 MB | |
| - **Model parameters:** 9,957,107 | |
| - **Persistent recurrent state:** 7,680 FP32 values = approximately 30 KiB per sequence | |
| A complete carbon-emissions estimate was not recorded. | |
| ## Evaluation | |
| The evaluations below characterize this **25M-token base checkpoint**. They should not be interpreted as general-purpose language-model benchmark results. | |
| ### Causal-LM Evaluation | |
| A matched stateless evaluation was performed with sequence length 16. | |
| | Evaluation split | Targets | Cross-entropy | Perplexity | | |
| |---|---:|---:|---:| | |
| | General validation | 65,536 | 3.594227 | 36.3876 | | |
| | TinyStories validation | 65,536 | 3.413734 | 30.3785 | | |
| ### Matched Transformer Control | |
| A parameter-matched Transformer control contained **9,969,792 parameters**, only 12,685 more than HyperS. Both checkpoints were trained to the same 25,000,960-token exposure under the matched short-context protocol. | |
| On the same stateless T=16 evaluation: | |
| | Model | General CE | General PPL | TinyStories CE | TinyStories PPL | | |
| |---|---:|---:|---:|---:| | |
| | HyperS-10M | 3.594227 | 36.3876 | 3.413734 | 30.3785 | | |
| | Transformer-10M control | 3.247609 | 25.7288 | 2.912044 | 18.3944 | | |
| Under this specific matched stateless short-context evaluation, the Transformer control achieved lower cross-entropy. | |
| This comparison should not be generalized beyond the evaluated protocol. | |
| ### Recurrent-History Utility | |
| HyperS was also evaluated on natural validation sequences while varying how much earlier history was allowed to enter the recurrent state before measuring the same 16-token future. | |
| Selected results: | |
| | Available history | Future-token CE | | |
| |---:|---:| | |
| | 0 | 3.597356 | | |
| | 16 | 3.378207 | | |
| | 64 | 3.352380 | | |
| | 128 | **3.348234** | | |
| | 256 | 3.349865 | | |
| | 512 | 3.351987 | | |
| | 1,024 | 3.357652 | | |
| The results show that recurrent history provided predictive utility on natural language, with the point estimate saturating around 64–128 history tokens in this experiment. | |
| A separate history-specificity test found that replacing earlier same-document history with unrelated natural-language history increasingly harmed prediction as the history window grew. This supports the conclusion that the fixed recurrent state carries document-specific information. | |
| These tests **do not establish exact retrieval of a token 1,024 positions earlier**. | |
| ### Generation Evaluation | |
| On a 14-prompt sampled generation suite using temperature 0.8 and top-k 50: | |
| | Metric | HyperS-10M | | |
| |---|---:| | |
| | Generated tokens | 1,851 | | |
| | EOS reached | 2 / 14 | | |
| | Valid UTF-8 | 14 / 14 | | |
| | Bigram repetition | 0.1698 | | |
| | Trigram repetition | 0.0703 | | |
| | Distinct-token ratio | 0.4942 | | |
| | Generation throughput | 41.36 tok/s | | |
| These aggregate statistics measure generation mechanics and repetition, not semantic answer quality. | |
| ### Inference-State Scaling | |
| HyperS retains a fixed recurrent state of: | |
| ```text | |
| 5 layers × 4 tracks × 384 dimensions = 7,680 scalar values | |
| ``` | |
| In the evaluated implementation the retained recurrent state was FP32, corresponding to approximately **30 KiB**, independent of processed context length. | |
| For comparison, a conventional 5-layer, 384-dimensional multi-head Transformer KV cache requires: | |
| ```text | |
| 2 × 5 × 384 = 3,840 cached scalar values per retained token | |
| ``` | |
| At 2,048 tokens and BF16 this corresponds to approximately **15 MiB**, a **512× byte ratio** relative to the measured HyperS recurrent state. | |
| This comparison concerns **retained context state**, not total peak CUDA allocation. Temporary activations and workspace memory are separate. | |
| The current HyperS implementation is substantially slower than the matched Transformer implementation. The fixed-state design should therefore be understood as a memory/representation trade-off, not a demonstrated speed advantage. | |
| ## Instruction-Tuning Status | |
| An instruction-tuned model is **not included in this release**. | |
| Exploratory instruction-tuning experiments showed that the 25M-token base could learn some simple prompt-conditioned mappings, but the resulting checkpoints did not meet the project's release threshold for reliable autoregressive instruction following, copying, and context lookup. | |
| Those experimental checkpoints are not part of the public model release. | |
| ## Model Examination | |
| The core experimental observation behind the release is that HyperS can maintain a constant-size recurrent state while still showing measurable, document-specific predictive benefit from earlier natural-language history. | |
| This is evidence for **history compression into the recurrent state**. It should not be interpreted as evidence of perfect memory, exact long-range retrieval, or superiority over attention-based language models. | |
| ## Environmental Impact | |
| A formal carbon-footprint measurement was not performed. | |
| - **Hardware:** NVIDIA GeForce RTX 4070 Laptop GPU | |
| - **Cloud provider:** None / local hardware | |
| - **Compute region:** Not reported | |
| - **Carbon emitted:** Not measured | |
| No numerical emissions estimate is provided because the necessary power-consumption and carbon-intensity measurements were not recorded. | |
| ## Technical Specifications | |
| ### Model Architecture and Objective | |
| HyperS is a decoder-only recurrent causal language model built around adaptive hyperspherical geometry. | |
| Its primary components are: | |
| 1. **HGT tokenizer** — fixed 8K vocabulary with UTF-8 byte fallback. | |
| 2. **HyperEmbedding** — token representations normalized onto a hyperspherical space. | |
| 3. **Adaptive metric** — low-rank state-conditioned metric deformation. | |
| 4. **HyperState** — recurrent spherical memory with four tracks per layer. | |
| 5. **Geometric state updates** — recurrent updates constructed around hyperspherical/geodesic operations. | |
| 6. **HAPD decoder** — geometry-aware token decoding using tied token prototypes. | |
| 7. **Causal next-token objective** — standard language-model cross-entropy. | |
| The architecture contains no Transformer self-attention or Transformer KV cache. | |
| ### Compute Infrastructure | |
| #### Hardware | |
| - NVIDIA GeForce RTX 4070 Laptop GPU | |
| #### Software | |
| - PyTorch | |
| - CUDA-capable NVIDIA environment | |
| - Custom HyperS model implementation | |
| - Custom HGT tokenizer implementation | |
| Exact source code is intentionally not part of this weights-only release. | |
| ## Citation | |
| If you use this model artifact directly, cite the Hugging Face repository: | |
| ```bibtex | |
| @misc{mondal2026hypers10m, | |
| author = {Sayak Mondal}, | |
| title = {HyperS-10M-Base-v1.0}, | |
| year = {2026}, | |
| howpublished = {Hugging Face model repository}, | |
| url = {https://huggingface.co/drelixer/HyperS} | |
| } | |
| ``` | |
| For the geometric framework that motivated HyperS: | |
| ```bibtex | |
| @article{mondal2026hyperspectrum, | |
| author = {Sayak Mondal}, | |
| title = {HyperSpectrum geometry: a Riemannian hyperspherical framework for curvature-adaptive text classification}, | |
| journal = {Neural Computing and Applications}, | |
| year = {2026}, | |
| publisher = {Springer Nature}, | |
| doi = {10.1007/s00521-026-12332-4} | |
| } | |
| ``` | |
| ## Glossary | |
| - **HSG:** HyperSpectrum Geometry. | |
| - **HGT:** Hyper-Geometric Tokenizer used by HyperS. | |
| - **HAPD:** HyperS geometry-aware prediction/decoding component. | |
| - **HyperState:** Recurrent geometric memory layer. | |
| - **TBPTT:** Truncated backpropagation through time. | |
| - **Persistent recurrent state:** The state carried between processed chunks; for this model it contains 7,680 scalar values per sequence. | |
| - **Stateless T=16 evaluation:** Evaluation in which each 16-token segment begins without earlier recurrent history. | |
| ## More Information | |
| - **Hugging Face:** https://huggingface.co/drelixer/HyperS | |
| - **Related paper:** https://doi.org/10.1007/s00521-026-12332-4 | |
| - **Paper venue:** *Neural Computing and Applications*, Springer Nature, 2026 | |
| HyperS is an ongoing research project. Future work includes larger-scale pretraining, improved recurrent-state optimization, value-binding and copying mechanisms, instruction tuning, and expanded evaluation. | |
| ## Model Card Authors | |
| **Sayak Mondal** | |
| ## Model Card Contact | |
| For questions, issues, or research discussions, use the Hugging Face repository community/discussion interface associated with `drelixer/HyperS`. | |