Dean Byrne PRO
Quazim0t0
AI & ML interests
DaisyChainAI🌼 / SmallLM's / San Francisco / Open Source
Recent Activity
updated a model about 7 hours ago
Quazim0t0/On-Fly-Jev reacted to DedeProGames's post with 🔥 about 15 hours ago
🧱 SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?
I built an arena where tiny decoder-only LMs (50K–250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.
How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack low…").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).
Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.
First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r ≈ 0.06). Survival does (r ≈ 0.9): the models that avoid holes and keep the stack low are the ones that win.
Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.
▶ Play: https://huggingface.co/spaces/DedeProGames/SLM-Tetris-Arena
📊 Results: https://huggingface.co/datasets/DedeProGames/lm-tetris-arena-results
Want your model in the Ranked pool? Drop it in the comments! repliedto their post 1 day ago
On-Fly-Jev is up: https://huggingface.co/Quazim0t0/On-Fly-Jev
Second Jev-style model. First was Byrne-Jev (70M SpikeWhale). This one is a 96M spiking trunk - every unit is a copy of one of 100 real MaleCNS fly neurons - plus a typed-decision head. One forward pass. No generated text.
I built it for speed.
Typed-decisions test split, 400 cases, 2,000 decisions, 5 questions per case, local GPU:
• 7.3 ms p50 per decision (~137 / s)
• 36.6 ms p50 / 53.5 ms p95 per case of 5
• ~27 cases / s
Same protocol vs the others:
• On-Fly-Jev: 36.6 ms / case, ~137 decisions / s
• Byrne-Jev: 110.7 ms / case, ~45 / s
• ModernBERT-base: 349 ms / case, ~14 / s
• TypeSafe Jev 1.13 (hosted, so network is in it): 710 ms / case, ~7 / s
About 3x Byrne-Jev, 9.5x ModernBERT, 19x TypeSafe Jev per case.
Live ViZDoom, 1 question per tick including game I/O: 23.4 ms (~43 / s).
• Accuracy 0.666 (Byrne-Jev 0.630, Jev 1.13 0.727)
• ECE 0.045, same as Byrne-Jev, about a third of Jev 1.13
50/50 merge of two checkpoints from one run. Research artifact, not a chatbot. More videos are on the card.