SmolLM2-135M SQL Drafter

A draft model for speculative decoding, fine-tuned to predict what HuggingFaceTB/SmolLM2-360M-Instruct outputs on text-to-SQL prompts.

The point

A drafter's job is not to be good at SQL. It is to predict what the target model says next. So this was trained by on-policy distillation: SQL prompts were run through the 360M target, and this 135M model was fine-tuned on the target's own completions โ€” not on the dataset's ground-truth SQL.

Train on ground truth and you get a model that writes good SQL and still gets rejected by the verifier.

Measured accept length

Target SmolLM2-360M-Instruct, 200 held-out text-to-SQL prompts. Accept length = tokens emitted per target forward pass. Ceiling is k+1.

k generic 135M draft this SQL drafter gain
4 3.751 4.239 +13%
7 5.067 6.078 +20%
9 5.791 7.043 +22%

The gain grows with k: at shallow depth the ceiling caps both drafters, but the deeper you speculate the more a generic drafter drifts on its own guesses while a target-matched one holds.

Use with vLLM

vllm serve HuggingFaceTB/SmolLM2-360M-Instruct \
  --speculative-config '{"method":"draft_model","model":"DEXX23/smollm2-135m-sql-drafter","num_speculative_tokens":7}'

Speculative decoding is lossless โ€” the target verifies every token, so a draft model can only change speed, never output.

Training

b-mc2/sql-create-context prompts โ†’ completions generated by the target โ†’ 3 epochs, lr 1e-4, bf16. About 10 minutes on one GPU; fits a 4GB card.

Downloads last month
11
Safetensors
Model size
0.1B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for DEXX23/smollm2-135m-sql-drafter

Finetuned
(389)
this model