SmolLM2-135M SQL Drafter
A draft model for speculative decoding, fine-tuned to predict what
HuggingFaceTB/SmolLM2-360M-Instruct outputs on text-to-SQL prompts.
The point
A drafter's job is not to be good at SQL. It is to predict what the target model says next. So this was trained by on-policy distillation: SQL prompts were run through the 360M target, and this 135M model was fine-tuned on the target's own completions โ not on the dataset's ground-truth SQL.
Train on ground truth and you get a model that writes good SQL and still gets rejected by the verifier.
Measured accept length
Target SmolLM2-360M-Instruct, 200 held-out text-to-SQL prompts.
Accept length = tokens emitted per target forward pass. Ceiling is k+1.
| k | generic 135M draft | this SQL drafter | gain |
|---|---|---|---|
| 4 | 3.751 | 4.239 | +13% |
| 7 | 5.067 | 6.078 | +20% |
| 9 | 5.791 | 7.043 | +22% |
The gain grows with k: at shallow depth the ceiling caps both drafters, but the deeper you speculate the more a generic drafter drifts on its own guesses while a target-matched one holds.
Use with vLLM
vllm serve HuggingFaceTB/SmolLM2-360M-Instruct \
--speculative-config '{"method":"draft_model","model":"DEXX23/smollm2-135m-sql-drafter","num_speculative_tokens":7}'
Speculative decoding is lossless โ the target verifies every token, so a draft model can only change speed, never output.
Training
b-mc2/sql-create-context prompts โ completions generated by the target โ
3 epochs, lr 1e-4, bf16. About 10 minutes on one GPU; fits a 4GB card.
- Downloads last month
- 11
Model tree for DEXX23/smollm2-135m-sql-drafter
Base model
HuggingFaceTB/SmolLM2-135M