HuggingFaceFW/fineweb-edu
Viewer • Updated • 3.5B • 420k • 1.31k
FS-SSA-LM-100M is an autoregressive causal language model (~93.9M parameters) powered by Softmax-Free Spiking Self-Attention (FS-SSA) operating at an ultra-low latency of K=2 timesteps.
Trained from scratch on ~655M tokens of FineWeb-Edu, the model eliminates quadratic floating-point Softmax attention and signed discrete spike accumulation, substituting dense floating-point Multiply-Accumulates (MACs) with sparse synaptic additions (ACs / SOPs).
| Parameter | Value |
|---|---|
| Total Parameters | 93,884,544 (~93.9M) |
| Layers | 16 |
| Hidden Dimension | 576 |
| Attention Heads | 9 |
| Sequence Length (T) | 1024 tokens |
| Vocabulary Size | 50,257 (GPT-2 BPE Tokenizer) |
| Temporal Latency (K) | K=2 timesteps |
| Spike Coding | Signed bipolar ternary pairs (+-) |
| Decay Structure | Causal per-head geometric ladder ( gamma in [0.875, 0.999]) |
| Sustained Firing Rate | ~13.1% (sparse event-driven regime) |
Pretrained across 10,000 steps with an effective batch size of 64 sequences (65,536 tokens/step):
| Model Architecture | Precision / Attention | Best Val Loss | Best PPL | Δ Loss (vs Dense) |
|---|---|---|---|---|
| Dense Baseline Control | FP16 / Softmax + GELU | 3.5399 | 34.46 | REF (Control) |
| FS-SSA-LM-100M (Verified) | Discrete K=2 / Softmax-Free | 3.6147 | 37.14 | +0.0748 nats (+2.68 PPL) |
| FS-SSA-LM-100M (Conservative) | Discrete K=2 / Softmax-Free | 3.6228 | 37.44 | +0.0829 nats (+2.98 PPL) |
The model achieves near-parity with the compute-matched dense Softmax Transformer (within 7.8% relative gap), while replacing dense continuous matrix multiplications with sparse integer event additions.