Instructions to use jtatman/kev-0.8b-code-verify-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use jtatman/kev-0.8b-code-verify-v2 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
kev-0.8b-code-verify-v2
A fine-tune of jaredpalmer/kev-0.8b (current best local judge: weak-coder data added): a Jev-style decision model that judges whether code written by a cheap coding model is correct, so a router can keep the answer local or escalate it. One forward pass, a calibrated probability, no generated text.
Trained on execution-labelled data: coder attempts at HumanEvalPack (Python) tasks, labelled by running each task's hidden test suite. The model never sees the hidden tests; its input is what a router can compute itself.
How to ask it
The fine-tune binds these exact strings. One noul question (probability that the statement holds):
{"type": "noul", "instructions": "Does the code correctly and completely implement the request, so it would pass a thorough hidden test suite including edge cases? Judge from the request, the code and the checks shown."}
State (an object; Kev renders it as key: value lines):
{"request": "<the task, verbatim>",
"code": "<the model's code>",
"checks": {"compiles": true, "defines_requested_function": true, "passes_examples_in_request": true,
"generated_edge_case_tests": "7 of 8 passed",
"note": "edge-case tests were written by a small model from the request alone; some may be wrong"}}
generated_edge_case_tests comes from a small model writing ~8 asserts from the request only, then running them.
Serve with Kev's runtime (python -m kev.serve --run jtatman/kev-0.8b-code-verify-v2, a TypeSafe /v1/systemone endpoint) or
kev.predictors.LocalPredictor. The checkpoint carries its fitted temperature.
Results
Unseen coder (llama-3.1-8b-instruct, never in training): 488 attempts, pass rate 0.57. "Kept local @X% err" = share of attempts a router can accept, highest score first, with at most X% of
them wrong. evidence score = (request examples passed + generated-assert pass rate) / 2.
| score | AUROC | Brier | ECE | kept local @2% err | @5% | @10% |
|---|---|---|---|---|---|---|
| baseline kev | 0.935 | 0.096 | 0.11 | 0.23 | 0.28 | 0.48 |
| fine-tuned kev | 0.970 | 0.055 | 0.04 | 0.22 | 0.51 | 0.64 |
| evidence score | 0.944 | 0.091 | 0.09 | 0.00 | 0.21 | 0.62 |
| mean(evidence, baseline) | 0.952 | 0.089 | 0.10 | 0.11 | 0.38 | 0.61 |
| mean(evidence, fine-tuned) | 0.961 | 0.063 | 0.07 | 0.18 | 0.39 | 0.64 |
Development (coders seen in training, tasks not seen):
| score | AUROC | Brier | ECE | kept local @2% err | @5% | @10% |
|---|---|---|---|---|---|---|
| baseline kev | 0.894 | 0.047 | 0.05 | 0.13 | 0.84 | 0.89 |
| fine-tuned kev | 0.913 | 0.043 | 0.03 | 0.22 | 0.83 | 0.89 |
| evidence score | 0.959 | 0.047 | 0.08 | 0.69 | 0.81 | 0.88 |
| mean(evidence, baseline) | 0.964 | 0.043 | 0.07 | 0.74 | 0.81 | 0.89 |
| mean(evidence, fine-tuned) | 0.962 | 0.041 | 0.05 | 0.72 | 0.81 | 0.89 |
Numbers are on small sets; see the project repo for task-bootstrap confidence intervals.
Recommended use
Average this model's probability with the execution evidence score, and tune the accept threshold per cheap-tier coder on a few hundred of that coder's execution-labelled attempts: a fixed threshold's error rate depends on how often the coder fails.
Training
- Init:
jaredpalmer/kev-0.8b@9a45d25(LoRA r=16 + pointer head, baseQwen/Qwen3.5-0.8B-Base), Kev's own trainer (kev.train, commit f2bb629d). - Data: 2290 execution-labelled records (task-grouped split;
llama-3.1-8bheld out) + 2000 public decision-v7 replay records against forgetting. - One epoch, lr 2e-05, batch 4 x accum 2, gradient checkpointing, bf16 autocast, state limit 1664 tokens.
- Temperature fitted on a held-out calibration split (min NLL).
- Hardware: one NVIDIA L4 (Google Colab).
Training data provenance
Code attempts whose correctness this model learned to judge (labels from executing hidden tests):
| source | model | license | training rows |
|---|---|---|---|
| or:gemma-3-4b | google/gemma-3-4b-it | Gemma Terms of Use | 275 |
| or:ministral-3b | mistralai/Ministral-3-3B-Instruct-2512 | Apache-2.0 | 242 |
| or:mistral-small-3.2-24b | mistralai/Mistral-Small-3.2-24B-Instruct-2506 | Apache-2.0 | 297 |
| or:qwen2.5-7b | Qwen/Qwen2.5-7B-Instruct | Apache-2.0 | 267 |
| or:qwen3-235b-a22b | Qwen/Qwen3-235B-A22B-Instruct-2507 | Apache-2.0 | 292 |
| or:qwen3-coder-30b-a3b | Qwen/Qwen3-Coder-30B-A3B-Instruct | Apache-2.0 | 238 |
| qwen2.5-coder:3b | Qwen/Qwen2.5-Coder-3B-Instruct | Qwen Research License | 344 |
| reference-buggy | bigcode/humanevalpack buggy solutions | MIT | 115 |
| reference-canonical | bigcode/humanevalpack canonical solutions | MIT | 115 |
| ternary-qwen3.8-27b | prism-ml/Ternary-Bonsai-2-27B-gguf (PTQ1_0; base Qwen/Qwen3.8-27B) | Apache-2.0 | 105 |
Built with Qwen. Training data includes outputs of Qwen2.5-Coder-3B-Instruct, used under the Qwen Research License Agreement (section 4b). Training data also includes code written by Gemma 3 (4B), used as labelled examples of correct and incorrect code; this model judges that code and is not trained to reproduce Gemma. Gemma is provided under and subject to the Gemma Terms of Use at ai.google.dev/gemma/terms. The held-out test coder (Llama 3.1 8B Instruct) was never used for training.
Limitations
- Python function-level tasks only (HumanEvalPack); other languages, repos and multi-file changes are untested.
- The coder pool is 8-10 models of 3B-235B parameters; accept thresholds do not transfer across coders.
- Trained on this question and state layout; other phrasings work less well.
Credits
Kev by Jared Palmer (jaredpalmer/kev, Apache-2.0); Qwen3.5 base (Apache-2.0); HumanEvalPack (bigcode). Project: https://github.com/jtatman/kev-code-verify
- Downloads last month
- 25