HyperS-10M-Base-v1.0

HyperS-10M-Base-v1.0 is a 9.96M-parameter experimental causal language model built around HyperSpectrum Geometry (HSG) and recurrent hyperspherical state updates rather than Transformer self-attention.

The model is intended as a compact research base model for studying geometry-driven sequence modeling, fixed-size recurrent context representations, adaptive metric learning, and non-Transformer language modeling.

Release status: Base model only. This checkpoint is not instruction-tuned, not a chat model, and is not intended to behave like an assistant.

The public repository intentionally contains only the packaged PyTorch model artifact and this model card. The HyperS implementation, training source, and inference source are not included in this release.

Model Details

Model Description

HyperS processes language through a recurrent geometric state rather than a Transformer KV cache. Tokens are embedded onto a normalized hyperspherical representation, processed through stacked recurrent HyperState layers, and decoded with a geometry-aware HAPD output module.

Each recurrent layer maintains a small collection of learned memory tracks. The model uses an adaptive low-rank metric to change how angular similarity is measured as the recurrent hidden state evolves.

The design goal is to study whether useful linguistic history can be continually compressed into a fixed-size geometric state rather than retaining a growing set of key/value vectors for every previous token.

  • Developed by: Sayak Mondal
  • Shared by: drelixer
  • Model: HyperS-10M-Base-v1.0
  • Model type: Decoder-only causal recurrent language model; non-Transformer
  • Parameters: 9,957,107
  • Language: Primarily English
  • Vocabulary: 8,192 tokens
  • Tokenizer: HGT-8K v0.1
  • Pretraining exposure: 25,000,960 tokens
  • Training precision: BF16
  • Artifact format: PyTorch .pt
  • Instruction tuned: No
  • License: Not yet specified

Architecture

Component Configuration
Vocabulary 8,192
Hidden dimension 384
Recurrent layers 5
Memory tracks per layer 4
Metric rank 32
Feed-forward dimension 768
Decoder rank 32
Parameters 9,957,107
Persistent recurrent state 7,680 scalar values / sequence
Observed recurrent-state dtype FP32
Persistent-state footprint ~30 KiB / sequence

HyperS does not use:

  • Transformer self-attention
  • a Transformer KV cache
  • positional embeddings as the mechanism for sequence order
  • a per-token retained context representation

Instead, sequence history is carried forward through the recurrent geometric state.

Model Sources

The published HyperSpectrum Geometry paper is the conceptual precursor to this project. It introduces the hyperspherical, metric-deformation perspective that motivated HyperS, but it does not describe this language-model architecture or this checkpoint.

Release Artifact

The repository contains:

hypers-10m-base-v1.0.pt

The .pt bundle contains:

  • the HyperS state_dict
  • the model configuration
  • the complete HGT-8K tokenizer configuration
  • release metadata

It intentionally excludes:

  • optimizer state
  • training-state checkpoints
  • TBPTT runtime state
  • RNG state
  • training scripts
  • inference scripts
  • model source code

The artifact is approximately 40.6 MB as uploaded to the Hugging Face Hub.

Uses

Direct Use

This checkpoint is primarily intended for:

  • research on non-Transformer causal language modeling
  • recurrent language-model experimentation
  • geometric and hyperspherical representation learning research
  • studying fixed-size recurrent context representations
  • controlled continuation experiments
  • architecture analysis and ablation studies
  • downstream fine-tuning research

It is a base language model, so raw generations should be expected to behave like unconstrained next-token completion rather than instruction-following responses.

Downstream Use

Potential downstream research directions include:

  • continued pretraining
  • domain adaptation
  • task-specific fine-tuning
  • instruction-tuning research
  • recurrent-state analysis
  • representation-geometry analysis
  • long-history compression experiments
  • alternative recurrent objectives
  • alternative HGT tokenization studies

A compatible implementation of the HyperS architecture is required to execute the weights. This repository does not provide the implementation.

Out-of-Scope Use

The released checkpoint should not be treated as:

  • a production assistant
  • an instruction-following model
  • a factual knowledge system
  • a safety-aligned conversational model
  • a high-reliability coding model
  • a replacement for substantially larger pretrained language models
  • a system suitable for medical, legal, financial, or other high-stakes decision making

The checkpoint has only 25M tokens of pretraining exposure and should be understood as a research-scale base model.

Bias, Risks, and Limitations

HyperS-10M-Base-v1.0 inherits limitations from both its training data and its small pretraining budget.

Important limitations include:

  • The model is small at 9.96M parameters.
  • Pretraining stopped at 25,000,960 exposed tokens, far below the scale normally used for modern general-purpose language models.
  • The model is not instruction-tuned.
  • Generated text can be repetitive, incoherent, factually incorrect, or unrelated to the prompt.
  • The model may reproduce biases, stereotypes, inaccuracies, or undesirable patterns present in its training sources.
  • Exact long-range retrieval has not been demonstrated.
  • A synthetic long-range digit-recall benchmark was not learned by either HyperS or the matched Transformer control and therefore was inconclusive as a long-memory test.
  • A public inference implementation is not included in this repository.
  • Hugging Face transformers.AutoModel / AutoModelForCausalLM compatibility is not provided.
  • The current implementation prioritizes the fixed-state research design rather than inference throughput.

Recommendations

Use this checkpoint as an experimental research artifact rather than a production model. Claims about long-context behavior should distinguish between predictive utility of recurrent history and exact retrieval of earlier tokens.

The model's outputs should be independently verified before being used for any factual or consequential purpose.

How to Get Started with the Model

This is a weights-only research release.

Download hypers-10m-base-v1.0.pt from the repository and load it using a compatible implementation of the HyperS architecture and HGT tokenizer.

The serialized bundle includes the model configuration and HGT-8K tokenizer data so that these do not need to be distributed as separate repository files.

No model or inference source code is distributed in this repository.

Training Details

Training Data

The pretraining stream was constructed from:

The encoded training corpus contained approximately 2.046B stored HGT tokens across the prepared sources. HyperS-10M-Base-v1.0 was exposed to 25,000,960 tokens from this stream; it was not trained over the entire encoded corpus.

TinyStories consists of synthetically generated short stories with a constrained vocabulary. FineWeb-Edu is a large collection of educational web text derived from FineWeb.

Users should consult the original dataset cards for dataset-specific provenance, licenses, filtering procedures, and limitations.

Preprocessing

Text was encoded using HGT-8K v0.1, a custom 8,192-token tokenizer with UTF-8 byte fallback.

Special token IDs:

Token ID
<PAD> 0
<BOS> 1
<EOS> 2
<DOC> 3
Byte fallback 4–259
Learned tokens 260–8191

Documents were framed as:

<BOS><DOC> content <EOS>

Training Procedure

HyperS was trained with causal next-token prediction while recurrent states were carried between short TBPTT chunks and detached at chunk boundaries.

Core training configuration:

Hyperparameter Value
Batch size 64
TBPTT length 16
Tokens / optimizer update 1,024
Optimizer AdamW
Adam Ξ²1 0.9
Adam Ξ²2 0.95
Peak learning rate 3e-4
Minimum LR ratio 0.1
Weight decay 0.1
Gradient clipping 1.0
Precision BF16
LR schedule horizon 200M tokens
Warmup ratio 1%
Final exposed tokens 25,000,960

The recurrent state was preserved between adjacent stream chunks during training while gradients were truncated at the TBPTT boundary.

Speeds, Sizes, Times

  • Training/evaluation GPU: NVIDIA GeForce RTX 4070 Laptop GPU
  • Public .pt artifact: approximately 40.6 MB
  • Model parameters: 9,957,107
  • Persistent recurrent state: 7,680 FP32 values = approximately 30 KiB per sequence

A complete carbon-emissions estimate was not recorded.

Evaluation

The evaluations below characterize this 25M-token base checkpoint. They should not be interpreted as general-purpose language-model benchmark results.

Causal-LM Evaluation

A matched stateless evaluation was performed with sequence length 16.

Evaluation split Targets Cross-entropy Perplexity
General validation 65,536 3.594227 36.3876
TinyStories validation 65,536 3.413734 30.3785

Matched Transformer Control

A parameter-matched Transformer control contained 9,969,792 parameters, only 12,685 more than HyperS. Both checkpoints were trained to the same 25,000,960-token exposure under the matched short-context protocol.

On the same stateless T=16 evaluation:

Model General CE General PPL TinyStories CE TinyStories PPL
HyperS-10M 3.594227 36.3876 3.413734 30.3785
Transformer-10M control 3.247609 25.7288 2.912044 18.3944

Under this specific matched stateless short-context evaluation, the Transformer control achieved lower cross-entropy.

This comparison should not be generalized beyond the evaluated protocol.

Recurrent-History Utility

HyperS was also evaluated on natural validation sequences while varying how much earlier history was allowed to enter the recurrent state before measuring the same 16-token future.

Selected results:

Available history Future-token CE
0 3.597356
16 3.378207
64 3.352380
128 3.348234
256 3.349865
512 3.351987
1,024 3.357652

The results show that recurrent history provided predictive utility on natural language, with the point estimate saturating around 64–128 history tokens in this experiment.

A separate history-specificity test found that replacing earlier same-document history with unrelated natural-language history increasingly harmed prediction as the history window grew. This supports the conclusion that the fixed recurrent state carries document-specific information.

These tests do not establish exact retrieval of a token 1,024 positions earlier.

Generation Evaluation

On a 14-prompt sampled generation suite using temperature 0.8 and top-k 50:

Metric HyperS-10M
Generated tokens 1,851
EOS reached 2 / 14
Valid UTF-8 14 / 14
Bigram repetition 0.1698
Trigram repetition 0.0703
Distinct-token ratio 0.4942
Generation throughput 41.36 tok/s

These aggregate statistics measure generation mechanics and repetition, not semantic answer quality.

Inference-State Scaling

HyperS retains a fixed recurrent state of:

5 layers Γ— 4 tracks Γ— 384 dimensions = 7,680 scalar values

In the evaluated implementation the retained recurrent state was FP32, corresponding to approximately 30 KiB, independent of processed context length.

For comparison, a conventional 5-layer, 384-dimensional multi-head Transformer KV cache requires:

2 Γ— 5 Γ— 384 = 3,840 cached scalar values per retained token

At 2,048 tokens and BF16 this corresponds to approximately 15 MiB, a 512Γ— byte ratio relative to the measured HyperS recurrent state.

This comparison concerns retained context state, not total peak CUDA allocation. Temporary activations and workspace memory are separate.

The current HyperS implementation is substantially slower than the matched Transformer implementation. The fixed-state design should therefore be understood as a memory/representation trade-off, not a demonstrated speed advantage.

Instruction-Tuning Status

An instruction-tuned model is not included in this release.

Exploratory instruction-tuning experiments showed that the 25M-token base could learn some simple prompt-conditioned mappings, but the resulting checkpoints did not meet the project's release threshold for reliable autoregressive instruction following, copying, and context lookup.

Those experimental checkpoints are not part of the public model release.

Model Examination

The core experimental observation behind the release is that HyperS can maintain a constant-size recurrent state while still showing measurable, document-specific predictive benefit from earlier natural-language history.

This is evidence for history compression into the recurrent state. It should not be interpreted as evidence of perfect memory, exact long-range retrieval, or superiority over attention-based language models.

Environmental Impact

A formal carbon-footprint measurement was not performed.

  • Hardware: NVIDIA GeForce RTX 4070 Laptop GPU
  • Cloud provider: None / local hardware
  • Compute region: Not reported
  • Carbon emitted: Not measured

No numerical emissions estimate is provided because the necessary power-consumption and carbon-intensity measurements were not recorded.

Technical Specifications

Model Architecture and Objective

HyperS is a decoder-only recurrent causal language model built around adaptive hyperspherical geometry.

Its primary components are:

  1. HGT tokenizer β€” fixed 8K vocabulary with UTF-8 byte fallback.
  2. HyperEmbedding β€” token representations normalized onto a hyperspherical space.
  3. Adaptive metric β€” low-rank state-conditioned metric deformation.
  4. HyperState β€” recurrent spherical memory with four tracks per layer.
  5. Geometric state updates β€” recurrent updates constructed around hyperspherical/geodesic operations.
  6. HAPD decoder β€” geometry-aware token decoding using tied token prototypes.
  7. Causal next-token objective β€” standard language-model cross-entropy.

The architecture contains no Transformer self-attention or Transformer KV cache.

Compute Infrastructure

Hardware

  • NVIDIA GeForce RTX 4070 Laptop GPU

Software

  • PyTorch
  • CUDA-capable NVIDIA environment
  • Custom HyperS model implementation
  • Custom HGT tokenizer implementation

Exact source code is intentionally not part of this weights-only release.

Citation

If you use this model artifact directly, cite the Hugging Face repository:

@misc{mondal2026hypers10m,
  author       = {Sayak Mondal},
  title        = {HyperS-10M-Base-v1.0},
  year         = {2026},
  howpublished = {Hugging Face model repository},
  url          = {https://huggingface.co/drelixer/HyperS}
}

For the geometric framework that motivated HyperS:

@article{mondal2026hyperspectrum,
  author    = {Sayak Mondal},
  title     = {HyperSpectrum geometry: a Riemannian hyperspherical framework for curvature-adaptive text classification},
  journal   = {Neural Computing and Applications},
  year      = {2026},
  publisher = {Springer Nature},
  doi       = {10.1007/s00521-026-12332-4}
}

Glossary

  • HSG: HyperSpectrum Geometry.
  • HGT: Hyper-Geometric Tokenizer used by HyperS.
  • HAPD: HyperS geometry-aware prediction/decoding component.
  • HyperState: Recurrent geometric memory layer.
  • TBPTT: Truncated backpropagation through time.
  • Persistent recurrent state: The state carried between processed chunks; for this model it contains 7,680 scalar values per sequence.
  • Stateless T=16 evaluation: Evaluation in which each 16-token segment begins without earlier recurrent history.

More Information

HyperS is an ongoing research project. Future work includes larger-scale pretraining, improved recurrent-state optimization, value-binding and copying mechanisms, instruction tuning, and expanded evaluation.

Model Card Authors

Sayak Mondal

Model Card Contact

For questions, issues, or research discussions, use the Hugging Face repository community/discussion interface associated with drelixer/HyperS.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train drelixer/HyperS