TinyChat-2M

TinyChat-2M is a 1.99M parameter, ultra-compact causal language model trained on conversational dialogue (starhopp3r/TinyChat). It is based off my custom architecture named Freeformer, a sub-quadratic, hardware-efficient architecture engineered to combine spherical vector-quantized sparse attention, parallel prefix-scan associative state memory, and factorized Kronecker matrix contractions.

Freeformer replaces conventional $O(N^2)$ dot-product attention and dense feed-forward networks with compute primitives designed for modern tensor-core and memory-coalescing GPU hierarchies, achieving training throughputs exceeding 120,000 tokens/sec on a cloud T4 GPU instance while maintaining monotonic convergence.


Architecture: Freeformer

Freeformer departs from standard Transformer and Mamba/SSM blocks by hybridizing linear causal state mixing, hardware-coalesced sparse token indexing, and Kronecker-decomposed FFNs.

Input Tokens [B, N]
       β”‚
  Embedding Layer (4096 -> 256)
       β”‚
β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Freeformer Block (x2)                                         β”‚
β”‚                                                               β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚   β”‚ Multi-Scale Causal State Mixer (MSCSM)                β”‚   β”‚
β”‚   β”‚   β”œβ”€ In-projection (d -> 2d) with Gated Split         β”‚   β”‚
β”‚   β”‚   β”œβ”€ Dilated Depthwise Conv1D (d=1, 2, 5; RF=20)      β”‚   β”‚
β”‚   β”‚   └─ SiLU Activation & Residual Injection             β”‚   β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                               β”‚                               β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚   β”‚ Residual Pyramid Attention (RPA)                      β”‚   β”‚
β”‚   β”‚   β”œβ”€ Causal Value Context Injector (DW Conv1D, k=3)   β”‚   β”‚
β”‚   β”‚   β”œβ”€ Multi-Octave Precomputed RoPE (Base 64 to 262k)   β”‚   β”‚
β”‚   β”‚   β”œβ”€ Spherical Residual VQ (SRVQ, L=2, M=4 -> 16 bkts)β”‚   β”‚
β”‚   β”‚   β”œβ”€ Triton Coalesced Causal Inverted Index Kernel    β”‚   β”‚
β”‚   β”‚   β”‚     └─ Gathers dynamic top-16 causal key/values   β”‚   β”‚
β”‚   β”‚   β”œβ”€ Vectorized Global Register Memory (G=4 Prefix)   β”‚   β”‚
β”‚   β”‚   β”œβ”€ Unified Softmax over Local (B_max) + Global (G)  β”‚   β”‚
β”‚   β”‚   └─ Swish/SiLU Multiplicative Output Gating          β”‚   β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                               β”‚                               β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚   β”‚ Factorized Kronecker FFN (FRK-FFN)                    β”‚   β”‚
β”‚   β”‚   β”œβ”€ Input reshape to 16 x 16 tensor matrix           β”‚   β”‚
β”‚   β”‚   β”œβ”€ Bilinear contraction: T_k = B_k Β· X Β· A_k^T      β”‚   β”‚
β”‚   β”‚   β”œβ”€ Dynamic Learned Router (u, g, v projections)     β”‚   β”‚
β”‚   β”‚   └─ Zero-parameter overhead rank-decomposition       β”‚   β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                β”‚
                  Final LayerNorm (RMS style)
                                β”‚
                 Logit Softcapping (C = 30.0)
                                β”‚
                        Tied LM Head Output

Key Components

1. Spherical Residual Vector Quantization (SRVQ) + Triton Inverted Index

Instead of calculating full pairwise attention matrices, queries and keys are projected onto a multi-codebook unit hypersphere ($L=2, M=4$, creating 16 discrete centroids). A custom Triton GPU kernel dynamically builds an inverted causal index in a single burst-coalesced sweep over memory. For every query token, only the causal tokens mapping to identical geometric clusters are retrieved ($B_{\text{max}} = 16$).

2. Vectorized Global Register Memory

Sparse local attention is prone to context fragmentation. Freeformer remedies this with a parallel prefix-scan associative memory ($G=4$ global registers). The registers perform a causal cumulative sum scan across sequence history: vreg[t]=βˆ‘Ο„<texp⁑(sΟ„)βŠ™vΟ„βˆ‘Ο„<texp⁑(sΟ„)+Ο΅\mathbf{v}_{\text{reg}}[t] = \frac{\sum_{\tau < t} \exp(\mathbf{s}_\tau) \odot \mathbf{v}_\tau}{\sum_{\tau < t} \exp(\mathbf{s}_\tau) + \epsilon} Local dynamic candidates and global register summaries are unified in a single, well-calibrated softmax partition function.

3. Factorized Kronecker FFN (FRK-FFN)

Standard MLPs account for ~66% of parameters in conventional LLMs. FRK-FFN restructures the hidden state vector $x \in \mathbb{R}^{256}$ into a matrix $X \in \mathbb{R}^{16 \times 16}$. Forward projections compute bilinear matrix contractions $T_k = B_k X A_k^T$ over small factor matrices ($A_k, B_k \in \mathbb{R}^{16 \times 16}$), executed as tensor-core GEMMs without scalar loop overhead.

4. Multi-Octave RoPE

Rotary position frequencies vary logarithmically across attention heads: bh=bminβ‹…(bmaxbmin)hHβˆ’1,bmin=64.0,bmax=262144.0b_h = b_{\text{min}} \cdot \left(\frac{b_{\text{max}}}{b_{\text{min}}}\right)^{\frac{h}{H - 1}}, \quad b_{\text{min}}=64.0, \quad b_{\text{max}}=262144.0 Heads with small base frequencies focus on local grammar and token adjacencies, while heads with large bases maintain phase stability across long context horizons.


Model Specifications

Parameter Value
Total Parameters 1,993,472 (~1.99M)
Layers 2
Hidden Dimension ($d_{\text{model}}$) 256
Attention Heads 8 (Head Dim: 32)
Vocabulary Size 4,096 (Trained BPE tokenizer)
Max Context Length 512 tokens
Attention Mechanism RPA (Dynamic Local $B_{\text{max}}=16$ + Global $G=4$)
Feed Forward FRK-FFN ($K=2$ experts, $16 \times 16$ Kronecker factors)
Position Embeddings Multi-Octave Rotary Position Embeddings (RoPE)
Logit Softcapping 30.0 ($\text{logits} = 30 \cdot \tanh(\text{logits} / 30)$)
Weight Tying Yes (Embedding $\leftrightarrow$ LM Head)

Training Details & Hyperparameters

  • Dataset: starhopp3r/TinyChat (Streaming causal multi-turn conversational dialog)
  • Optimization Strategy:
    • Matrix weights ($\ge 16 \times 16$): Muon Optimizer (Learning rate: $3.0 \times 10^{-2}$, Momentum: 0.95, 5-step Quintic Newton-Schulz semi-orthogonalization)
    • Vectors, biases, embeddings, LM head: AdamW (Learning rate: $1.5 \times 10^{-3}$, $\beta_1=0.9, \beta_2=0.95$, Weight Decay: 0.01)
    • Polyak Averaging: Fast Foreach Exponential Moving Average (EMA decay: 0.98)
  • Precision: Mixed Precision FP16 via PyTorch autocast + GradScaler with Tensor Core TF32 enabled
  • Compiler: torch.compile (inductor backend) + Custom Triton Causal Inverted Index
  • Batch Size: 32 sequences $\times$ 2 gradient accumulation steps = 64 effective batch size (32,768 tokens per step)
  • Total Training Steps: 20,000 steps (~655 Million tokens processed, approximately 3–4 epochs)

Training

Training demonstrated steady monotonic decrease in cross-entropy loss alongside consistent clustering balance within the spherical quantization codebook.

Step       Learning Rate   Train Loss   Val Loss    Aux Loss     Throughput
─────────────────────────────────────────────────────────────────────────────
10         3.00e-05        8.3921       β€”           0.0043         4,046 tok/s (Compilation phase)
250        7.50e-04        3.0237       3.2639      0.0034       131,943 tok/s
500        1.50e-03        2.4245       2.3908      0.0035       124,337 tok/s
1,000      1.50e-03        2.1781       2.0870      0.0028       113,223 tok/s
2,500      1.47e-03        1.9691       1.9004      0.0013       118,106 tok/s
5,000      1.33e-03        1.8911       1.8304      0.0008       117,996 tok/s
10,000     8.52e-04        1.8238       1.7821      0.0006       117,516 tok/s
15,000     3.57e-04        1.7842       1.7549      0.0005       118,726 tok/s
20,000     1.50e-04        1.7630       1.7428      0.0005       118,547 tok/s

Generation Samples across Checkpoints

  • Step 500 (Early semantic alignment):

    Prompt: [INST] Why do people always think about how their words change and feel scared. [/INST] Output: It is understandable to feel that way and fear can feel a bit overwhelming. [INST] Yes, it is hard to stay calm when others do not understand my feelings. [/INST] Finding ways to cope can be strong, but sometimes

  • Step 5,500 (Coherent multi-turn dialog):

    Prompt: [INST] How are you feeling today with your happy too? [/INST] Output: I feel the same way about the sunny weather and its impact on our plans. [INST] I wanted to wear my card to the park with my friends, it is so nice. [/INST] That sounds great, spending time with friends is always a

  • Step 20,000 (Final model):

    Prompt: [INST] Can you say something to the manager about the new manager about the project? [/INST] Output: I have not yet want to talk to him about it before they are important. [/INST] That sounds like a good idea, but will you really want to improve his work? [INST] I have not, but I

Training Code

The training code is available in this repository and is released under the same license as the model.


Intended Use and Limitations

Intended Use Cases

  • Architecture Research: Testbed for examining a novel architecture.

Limitations

  • World Knowledge: Due to its 2M parameter budget, TinyChat-2M cannot serve as a reliable knowledge base or factual encyclopedia. It excels at local grammar, conversational tone, and syntactic reasoning, but is prone to hallucinating factual details.
  • Context Bound: Optimal generation performance occurs within its 512-token sequence window.
  • Format Dependency: The model is trained on conversational prompt structures adhering to [INST] ... [/INST] (see starhopp3r/TinyChat). Prompting without instruction markers may degrade reply coherence.

Citation & Authors

If you utilize the Freeformer architecture or the TinyChat-2M checkpoint in your research or applications, please cite:

@misc{tinychat2m_freeformer2026,
  author = {Oscar Lo},
  title = {TinyChat-2M: High-Throughput Sub-Quadratic Language Modeling via Freeformer},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train oscar128372/TinyChat-2M