rapidchat âš¡
Frontier scores: official τ²-bench leaderboard (user simulator gpt-5.2, standard harness). rapidchat: our runs, user simulator gpt-6.1-sol, custom harness, train split seen in training. Different setups, not a like-for-like ranking.
a 9B that handles airline customer support on τ²-bench. built in about a day and a half. ngl, that's the point.
the full training + practice data is open too: rapidchat-data 📦
the numbers
τ²-bench airline, all 50 tasks × 4 trials each. User simulator: gpt-6.1-sol (low reasoning). Served with vLLM, bf16, temperature 0.6, top_p 0.95, top_k 20.
| setup | pass^1 | pass^2 | pass^3 | pass^4 | test split (20) | train split (30) |
|---|---|---|---|---|---|---|
| with working notes | 0.850 | 0.803 | 0.775 | 0.760 | 0.787 | 0.892 |
| without working notes | 0.865 | 0.813 | 0.780 | 0.760 | 0.825 | 0.892 |
the hot take 🔥
rapidchat is a receipt. The whole pipeline (synthetic scenarios, self-play conversations, three rounds of LoRA, the eval harness, even this card) was built and run end to end by an AI agent, Claude Opus 5.5 on medium effort, on one rented A100. It wasn't even that hard.
So if the pitch is "we fine-tuned an open-weights model for our vertical", lowkey that's not a moat, it's a weekend. Stop trying to fight the frontier labs on their own turf, and stop shipping OSS fine-tunes as a marketing gimmick. People can tell, fr. The builders you actually need to scale can tell too, and it's not pulling them in. Put that energy into the product, the distribution and the data nobody else has.
what's inside
- base: Qwen/Qwen3.5-9B, thinking mode on
- training: LoRA rank 32, three rounds, merged into the weights
- data: Training data: the model's own conversations that scored reward 1.0, on (a) synthetic airline scenarios generated over a re-skinned copy of the tau2-bench airline database (renamed users, reservations and flights; users referenced by any official task excluded) and (b) the 30 official airline train-split tasks. No official test-split task, and no conversation on one, was used as training data. The round-3 scenario families and the working notes were designed after reviewing failure modes on official tasks, including test-split tasks; the scenarios are new look-alike cases, not copies of test tasks. 30 of the 50 airline tasks (the train split) were seen in training.
- harness: reasoning stays in context within a chain of tool calls, one retry on an empty reply. The "with working notes" row also appends
working_notes.txt(short airline-policy reminders, in this repo) after the standard τ²-bench agent instruction.
run it
vllm serve badr7/rapidchat --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder --language-model-only
run it on your phone 📱
4-bit builds for phones and laptops:
- rapidchat-GGUF: Q4_K_M (5.6 GB) and Q8_0 (9.5 GB) for llama.cpp, LM Studio, Ollama and GGUF phone apps
- rapidchat-MLX-4bit: for iPhone, iPad and Mac (Apple silicon)
The quantized builds were not benchmarked separately; the scores above are for the full bf16 model.
fine print
- Specialised for the τ²-bench airline domain; not evaluated on retail, telecom or banking.
- Scores move with the user simulator; ours is listed above.
- 30 of the 50 airline tasks (the train split) were seen in training, so the test-split column is the cleaner signal.
- It's a benchmark model. Please don't let it rebook your actual flights lol.
- Downloads last month
- 20
