kev-4b-code-verify-v2

A fine-tune of jaredpalmer/kev-4b (best judge: 4B with weak-coder data, trained on SaladCloud (RTX 3090 Ti)): a Jev-style decision model that judges whether code written by a cheap coding model is correct, so a router can keep the answer local or escalate it. One forward pass, a calibrated probability, no generated text.

Trained on execution-labelled data: coder attempts at HumanEvalPack (Python) tasks, labelled by running each task's hidden test suite. The model never sees the hidden tests; its input is what a router can compute itself.

How to ask it

The fine-tune binds these exact strings. One noul question (probability that the statement holds):

{"type": "noul", "instructions": "Does the code correctly and completely implement the request, so it would pass a thorough hidden test suite including edge cases? Judge from the request, the code and the checks shown."}

State (an object; Kev renders it as key: value lines):

{"request": "<the task, verbatim>",
  "code": "<the model's code>",
  "checks": {"compiles": true, "defines_requested_function": true, "passes_examples_in_request": true,
             "generated_edge_case_tests": "7 of 8 passed",
             "note": "edge-case tests were written by a small model from the request alone; some may be wrong"}}

generated_edge_case_tests comes from a small model writing ~8 asserts from the request only, then running them. Serve with Kev's runtime (python -m kev.serve --run jtatman/kev-4b-code-verify-v2, a TypeSafe /v1/systemone endpoint) or kev.predictors.LocalPredictor. The checkpoint carries its fitted temperature.

Results

Unseen coder (llama-3.1-8b-instruct, never in training): 488 attempts, pass rate 0.57. "Kept local @X% err" = share of attempts a router can accept, highest score first, with at most X% of them wrong. evidence score = (request examples passed + generated-assert pass rate) / 2.

score AUROC Brier ECE kept local @2% err @5% @10%
baseline kev 0.971 0.060 0.03 0.31 0.51 0.62
fine-tuned kev 0.984 0.040 0.03 0.47 0.59 0.64
evidence score 0.944 0.091 0.09 0.00 0.21 0.62
mean(evidence, baseline) 0.970 0.066 0.06 0.31 0.51 0.63
mean(evidence, fine-tuned) 0.980 0.054 0.09 0.34 0.58 0.64

Second unseen coder (local Qwen3.5-9B fine-tune, IQ2_M GGUF, never in training): 486 attempts, pass rate 0.64.

score AUROC Brier ECE kept local @2% err @5% @10%
baseline kev 0.966 0.053 0.04 0.35 0.62 0.71
fine-tuned kev 0.984 0.035 0.03 0.48 0.67 0.71
evidence score 0.965 0.068 0.07 0.00 0.49 0.69
mean(evidence, baseline) 0.976 0.053 0.07 0.52 0.62 0.71
mean(evidence, fine-tuned) 0.984 0.043 0.08 0.52 0.67 0.71

Development (coders seen in training, tasks not seen):

score AUROC Brier ECE kept local @2% err @5% @10%
baseline kev 0.971 0.049 0.07 0.76 0.82 0.89
fine-tuned kev 0.968 0.042 0.06 0.66 0.84 0.89
evidence score 0.959 0.047 0.08 0.69 0.81 0.88
mean(evidence, baseline) 0.976 0.039 0.07 0.78 0.84 0.89
mean(evidence, fine-tuned) 0.977 0.038 0.07 0.73 0.84 0.89

Numbers are on small sets; see the project repo for task-bootstrap confidence intervals.

Recommended use

Average this model's probability with the execution evidence score, and tune the accept threshold per cheap-tier coder on a few hundred of that coder's execution-labelled attempts: a fixed threshold's error rate depends on how often the coder fails.

Training

  • Init: jaredpalmer/kev-4b@6cfce5c (LoRA r=16 + pointer head, base Qwen/Qwen3.5-4B-Base), Kev's own trainer (kev.train, commit f2bb629d).
  • Data: 2290 execution-labelled records (task-grouped split; llama-3.1-8b and the local 9B coder held out) + 2000 public decision-v7 replay records against forgetting.
  • One epoch, lr 2e-05, batch 1 x accum 8, gradient checkpointing, bf16 autocast, state limit 1664 tokens, frozen backbone in bf16.
  • Temperature fitted on a held-out calibration split (min NLL).
  • Hardware: one NVIDIA RTX 3090 Ti (SaladCloud, ghcr.io/jtatman/kev-trainer:0.1.2).

Training data provenance

Code attempts whose correctness this model learned to judge (labels from executing hidden tests):

source model license training rows
or:gemma-3-4b google/gemma-3-4b-it Gemma Terms of Use 275
or:ministral-3b mistralai/Ministral-3-3B-Instruct-2512 Apache-2.0 242
or:mistral-small-3.2-24b mistralai/Mistral-Small-3.2-24B-Instruct-2506 Apache-2.0 297
or:qwen2.5-7b Qwen/Qwen2.5-7B-Instruct Apache-2.0 267
or:qwen3-235b-a22b Qwen/Qwen3-235B-A22B-Instruct-2507 Apache-2.0 292
or:qwen3-coder-30b-a3b Qwen/Qwen3-Coder-30B-A3B-Instruct Apache-2.0 238
qwen2.5-coder:3b Qwen/Qwen2.5-Coder-3B-Instruct Qwen Research License 344
reference-buggy bigcode/humanevalpack buggy solutions MIT 115
reference-canonical bigcode/humanevalpack canonical solutions MIT 115
ternary-qwen3.8-27b prism-ml/Ternary-Bonsai-2-27B-gguf (PTQ1_0; base Qwen/Qwen3.8-27B) Apache-2.0 105

Built with Qwen. Training data includes outputs of Qwen2.5-Coder-3B-Instruct, used under the Qwen Research License Agreement (section 4b). Training data also includes code written by Gemma 3 (4B), used as labelled examples of correct and incorrect code; this model judges that code and is not trained to reproduce Gemma. Gemma is provided under and subject to the Gemma Terms of Use at ai.google.dev/gemma/terms. The held-out test coder (Llama 3.1 8B Instruct) was never used for training.

Limitations

  • Python function-level tasks only (HumanEvalPack); other languages, repos and multi-file changes are untested.
  • The coder pool is 8-10 models of 3B-235B parameters; accept thresholds do not transfer across coders.
  • Trained on this question and state layout; other phrasings work less well.

Credits

Kev by Jared Palmer (jaredpalmer/kev, Apache-2.0); Qwen3.5 base (Apache-2.0); HumanEvalPack (bigcode). Project: https://github.com/jtatman/kev-code-verify

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jtatman/kev-4b-code-verify-v2

Adapter
(2)
this model

Collection including jtatman/kev-4b-code-verify-v2