Instructions to use hopski/seat-intelligence-1.7b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use hopski/seat-intelligence-1.7b with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download hopski/seat-intelligence-1.7b --local-dir seat-intelligence-1.7b
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Seat Intelligence (SI) 1.7B
A tiny model that stops AI leaders from causing the apocalypse. It does this by fixing the seating chart.
It reads a seating rule written in plain English and writes a short instruction program. It never seats anyone itself: a separate, tested seating program follows the instructions exactly. Play with it at seat-intelligence.com; code, data and every experiment are at github.com/ho3h/seat-intelligence.
This repository holds a LoRA adapter (35 MB, MLX format) for mlx-community/Qwen3-1.7B-4bit. It runs on Apple-silicon Macs with mlx-lm.
The instruction words
| Word | Meaning |
|---|---|
size N |
at most N guests per section (1 to 6) |
together company |
guests from the same company sit in one section |
together CAT |
all guests of a category sit together |
limit CAT [CAT ...] K |
at most K guests from the listed categories per section |
apart CATA CATB |
no guest of CATA shares a section with a guest of CATB |
order CAT [CAT ...] |
seat these categories first, in this order |
avoid X Y |
the named guest or company X never shares a section with Y |
pair X Y |
X and Y sit in the same section |
Categories: ai_lab, big_tech, chips, software_security, investor, government, unlabelled. Names are written with underscores, for example Elon_Musk or OpenAI.
Use
from huggingface_hub import snapshot_download
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
path = snapshot_download("hopski/seat-intelligence-1.7b")
model, tok = load("mlx-community/Qwen3-1.7B-4bit", adapter_path=path)
def prompt(text):
return ("Convert the seating policy into a stage program (one word per line), or `none` if it states no rules.\n\n"
f"## Policy\n\n{text}\n\n## Reply format\n\nReply with the stage program only.\n")
p = tok.apply_chat_template([{"role": "user", "content": prompt("Keep Elon Musk away from OpenAI and from Mark Zuckerberg.")}],
add_generation_prompt=True, enable_thinking=False)
print(generate(model, tok, p, max_tokens=120, sampler=make_sampler(temp=0.0)))
# avoid Elon_Musk OpenAI
# avoid Elon_Musk Mark_Zuckerberg
The seating program that runs these instructions, and the checker that counts broken rules, are in the GitHub repository (demo/luncheon/seating.js, genome/hero6/lang6.py).
How well it works
Measured on this laptop-sized model with the prompt above, plain decoding (pass = the instructions it writes give exactly the intended seating):
| Test | This version (v2) | v1 |
|---|---|---|
| 100 fresh rules by a separate AI writer it never saw, 5 tries each (the clean test: written, with their answers frozen, before this version was trained) | 442 of 500 (88%) under the strict check, 449 (90%) under the lenient one | 392 (78%) strict, 405 (81%) lenient |
| 72 rules naming guests, worded by a separate AI writer, 5 tries each | 336 of 360 (93%) strict, 347 (96%) lenient | 309 (86%) strict, 322 (89%) lenient |
| 72 earlier rules by category, 5 tries each | 355 of 360 (99%) | 334 (93%) |
| 72 rules in 14 writing styles, 5 tries each | 325 of 360 (90%) | 315 (88%) |
| Time to write the instructions | 0.1 to 0.3 seconds on an Apple M5 Max |
The last three test sets were around while earlier versions were built, so their gains are probably a little flattering; the first row is the number to quote. Both versions were re-scored the same way for this table. Restricting the output to valid instruction words while it writes ("constrained decoding", in the repository) adds about two points on the fresh test (451 of 500); voting over five tries adds nothing. A 4B version trained the same way scored within about a point on the fresh test, so this stays tiny.
For comparison, an untuned 4B chatbot asked to seat the guests directly gave 0 fully correct answers in 400 tries on 20 such rules.
Where it goes wrong
Of its 11 misses on the fresh test (one try each):
- Six involve "White House staff", a group the chart never labels (the guests are
unlabelled), which it reads asgovernmentor invents a category for. - Three involve guests added on the spot who aren't on the chart, where it mixes up which name goes with which company.
- Similar or partial names can pick the wrong guest: asked to keep Arora apart from Daniels, it chose Director Clayton.
- One shorthand ("Palantir duo: same section") came out as an instruction naming the company instead of its two guests.
- A request outside its words (for example "at least two") can turn into the wrong instruction instead of an error.
- It always shows the instructions it wrote, so a person can check them before anyone sits down.
Training
LoRA on mlx-community/Qwen3-1.7B-4bit: 15,912 sentences (v1's 8,948 plus replayed earlier examples, terse phrasings, rarer rules and paraphrases), with anything matching a test sentence, or sharing any six-word phrase with one, removed; rank 8, scale 20, batch 16, 900 iterations, learning rate 5e-5. The training data generator and the frozen test sets are in the GitHub repository. Write-ups: docs/HERO-1.md, docs/HERO-6.md and docs/HERO-7.md.
Versions: v2 (1 October 2026) is this one. v1 (30 September 2026) is in this repository's history.
A game: the rules are made up and say nothing about any real person. Not affiliated with anyone at the table.
Quantized