TinyChat-2M
TinyChat-2M is a 1.99M parameter, ultra-compact causal language model trained on conversational dialogue (starhopp3r/TinyChat). It is based off my custom architecture named Freeformer, a sub-quadratic, hardware-efficient architecture engineered to combine spherical vector-quantized sparse attention, parallel prefix-scan associative state memory, and factorized Kronecker matrix contractions.
Freeformer replaces conventional $O(N^2)$ dot-product attention and dense feed-forward networks with compute primitives designed for modern tensor-core and memory-coalescing GPU hierarchies, achieving training throughputs exceeding 120,000 tokens/sec on a cloud T4 GPU instance while maintaining monotonic convergence.
Architecture: Freeformer
Freeformer departs from standard Transformer and Mamba/SSM blocks by hybridizing linear causal state mixing, hardware-coalesced sparse token indexing, and Kronecker-decomposed FFNs.
Input Tokens [B, N]
β
Embedding Layer (4096 -> 256)
β
ββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Freeformer Block (x2) β
β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Multi-Scale Causal State Mixer (MSCSM) β β
β β ββ In-projection (d -> 2d) with Gated Split β β
β β ββ Dilated Depthwise Conv1D (d=1, 2, 5; RF=20) β β
β β ββ SiLU Activation & Residual Injection β β
β βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββ β
β β β
β βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββ β
β β Residual Pyramid Attention (RPA) β β
β β ββ Causal Value Context Injector (DW Conv1D, k=3) β β
β β ββ Multi-Octave Precomputed RoPE (Base 64 to 262k) β β
β β ββ Spherical Residual VQ (SRVQ, L=2, M=4 -> 16 bkts)β β
β β ββ Triton Coalesced Causal Inverted Index Kernel β β
β β β ββ Gathers dynamic top-16 causal key/values β β
β β ββ Vectorized Global Register Memory (G=4 Prefix) β β
β β ββ Unified Softmax over Local (B_max) + Global (G) β β
β β ββ Swish/SiLU Multiplicative Output Gating β β
β βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββ β
β β β
β βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββ β
β β Factorized Kronecker FFN (FRK-FFN) β β
β β ββ Input reshape to 16 x 16 tensor matrix β β
β β ββ Bilinear contraction: T_k = B_k Β· X Β· A_k^T β β
β β ββ Dynamic Learned Router (u, g, v projections) β β
β β ββ Zero-parameter overhead rank-decomposition β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββ
β
Final LayerNorm (RMS style)
β
Logit Softcapping (C = 30.0)
β
Tied LM Head Output
Key Components
1. Spherical Residual Vector Quantization (SRVQ) + Triton Inverted Index
Instead of calculating full pairwise attention matrices, queries and keys are projected onto a multi-codebook unit hypersphere ($L=2, M=4$, creating 16 discrete centroids). A custom Triton GPU kernel dynamically builds an inverted causal index in a single burst-coalesced sweep over memory. For every query token, only the causal tokens mapping to identical geometric clusters are retrieved ($B_{\text{max}} = 16$).
2. Vectorized Global Register Memory
Sparse local attention is prone to context fragmentation. Freeformer remedies this with a parallel prefix-scan associative memory ($G=4$ global registers). The registers perform a causal cumulative sum scan across sequence history: Local dynamic candidates and global register summaries are unified in a single, well-calibrated softmax partition function.
3. Factorized Kronecker FFN (FRK-FFN)
Standard MLPs account for ~66% of parameters in conventional LLMs. FRK-FFN restructures the hidden state vector $x \in \mathbb{R}^{256}$ into a matrix $X \in \mathbb{R}^{16 \times 16}$. Forward projections compute bilinear matrix contractions $T_k = B_k X A_k^T$ over small factor matrices ($A_k, B_k \in \mathbb{R}^{16 \times 16}$), executed as tensor-core GEMMs without scalar loop overhead.
4. Multi-Octave RoPE
Rotary position frequencies vary logarithmically across attention heads: Heads with small base frequencies focus on local grammar and token adjacencies, while heads with large bases maintain phase stability across long context horizons.
Model Specifications
| Parameter | Value |
|---|---|
| Total Parameters | 1,993,472 (~1.99M) |
| Layers | 2 |
| Hidden Dimension ($d_{\text{model}}$) | 256 |
| Attention Heads | 8 (Head Dim: 32) |
| Vocabulary Size | 4,096 (Trained BPE tokenizer) |
| Max Context Length | 512 tokens |
| Attention Mechanism | RPA (Dynamic Local $B_{\text{max}}=16$ + Global $G=4$) |
| Feed Forward | FRK-FFN ($K=2$ experts, $16 \times 16$ Kronecker factors) |
| Position Embeddings | Multi-Octave Rotary Position Embeddings (RoPE) |
| Logit Softcapping | 30.0 ($\text{logits} = 30 \cdot \tanh(\text{logits} / 30)$) |
| Weight Tying | Yes (Embedding $\leftrightarrow$ LM Head) |
Training Details & Hyperparameters
- Dataset:
starhopp3r/TinyChat(Streaming causal multi-turn conversational dialog) - Optimization Strategy:
- Matrix weights ($\ge 16 \times 16$): Muon Optimizer (Learning rate: $3.0 \times 10^{-2}$, Momentum: 0.95, 5-step Quintic Newton-Schulz semi-orthogonalization)
- Vectors, biases, embeddings, LM head: AdamW (Learning rate: $1.5 \times 10^{-3}$, $\beta_1=0.9, \beta_2=0.95$, Weight Decay: 0.01)
- Polyak Averaging: Fast Foreach Exponential Moving Average (EMA decay: 0.98)
- Precision: Mixed Precision FP16 via PyTorch
autocast+GradScalerwith Tensor Core TF32 enabled - Compiler:
torch.compile(inductor backend) + Custom Triton Causal Inverted Index - Batch Size: 32 sequences $\times$ 2 gradient accumulation steps = 64 effective batch size (32,768 tokens per step)
- Total Training Steps: 20,000 steps (~655 Million tokens processed, approximately 3β4 epochs)
Training
Training demonstrated steady monotonic decrease in cross-entropy loss alongside consistent clustering balance within the spherical quantization codebook.
Step Learning Rate Train Loss Val Loss Aux Loss Throughput
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
10 3.00e-05 8.3921 β 0.0043 4,046 tok/s (Compilation phase)
250 7.50e-04 3.0237 3.2639 0.0034 131,943 tok/s
500 1.50e-03 2.4245 2.3908 0.0035 124,337 tok/s
1,000 1.50e-03 2.1781 2.0870 0.0028 113,223 tok/s
2,500 1.47e-03 1.9691 1.9004 0.0013 118,106 tok/s
5,000 1.33e-03 1.8911 1.8304 0.0008 117,996 tok/s
10,000 8.52e-04 1.8238 1.7821 0.0006 117,516 tok/s
15,000 3.57e-04 1.7842 1.7549 0.0005 118,726 tok/s
20,000 1.50e-04 1.7630 1.7428 0.0005 118,547 tok/s
Generation Samples across Checkpoints
Step 500 (Early semantic alignment):
Prompt:
[INST] Why do people always think about how their words change and feel scared. [/INST]Output:It is understandable to feel that way and fear can feel a bit overwhelming. [INST] Yes, it is hard to stay calm when others do not understand my feelings. [/INST] Finding ways to cope can be strong, but sometimesStep 5,500 (Coherent multi-turn dialog):
Prompt:
[INST] How are you feeling today with your happy too? [/INST]Output:I feel the same way about the sunny weather and its impact on our plans. [INST] I wanted to wear my card to the park with my friends, it is so nice. [/INST] That sounds great, spending time with friends is always aStep 20,000 (Final model):
Prompt:
[INST] Can you say something to the manager about the new manager about the project? [/INST]Output:I have not yet want to talk to him about it before they are important. [/INST] That sounds like a good idea, but will you really want to improve his work? [INST] I have not, but I
Training Code
The training code is available in this repository and is released under the same license as the model.
Intended Use and Limitations
Intended Use Cases
- Architecture Research: Testbed for examining a novel architecture.
Limitations
- World Knowledge: Due to its 2M parameter budget, TinyChat-2M cannot serve as a reliable knowledge base or factual encyclopedia. It excels at local grammar, conversational tone, and syntactic reasoning, but is prone to hallucinating factual details.
- Context Bound: Optimal generation performance occurs within its 512-token sequence window.
- Format Dependency: The model is trained on conversational prompt structures adhering to
[INST] ... [/INST](seestarhopp3r/TinyChat). Prompting without instruction markers may degrade reply coherence.
Citation & Authors
If you utilize the Freeformer architecture or the TinyChat-2M checkpoint in your research or applications, please cite:
@misc{tinychat2m_freeformer2026,
author = {Oscar Lo},
title = {TinyChat-2M: High-Throughput Sub-Quadratic Language Modeling via Freeformer},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/}}
}