mverify
A CPU classifier for the mcode verification pass. It scores one pair: does this output satisfy this prompt?
It is not an LLM. It is not a general correctness judge. It does not run tests or builds.
- Code: https://git.simonharms.com/thesimonharms/mverify
- Weights: this repo
- License: MIT
Files
| file | role |
|---|---|
mverify.npz |
TF-IDF weights and logistic-regression coefficients |
vocab.json |
term vocabulary |
mverify.json |
threshold, pair-feature names, train metadata |
eval.json |
last measured scores |
Download into a local artifacts folder:
hf download thesimonharms/mverify --local-dir artifacts
Use
Install the code from git, then check a pair:
uv add "mverify @ git+https://git.simonharms.com/thesimonharms/mverify"
uv run mverify check --prompt "What is 6*7?" --output "42"
Or from a clone:
git clone https://git.simonharms.com/thesimonharms/mverify.git
cd mverify
uv sync --extra dev
uv run mverify check --prompt "Write a Python add function" --output "def add(a, b): return a + b"
Stdin JSON is also valid:
echo '{"prompt":"...","output":"..."}' | uv run mverify check --stdin
{
"pass": true,
"confidence": 0.93,
"reasons": ["output has a code block in the requested language"]
}
pass: did the output do the thing the prompt askedconfidence: P(output satisfies prompt), 0 to 1- Exit code 0 means the CLI ran.
passcan still be false
mcode should treat a hit as pass && confidence >= confidenceThreshold (default 0.7).
Model
TF-IDF plus L2 logistic regression on packed prompt/output text, plus pair features (overlap, format gaps, refusal, language). Inference is numpy on CPU. There is no GPU path.
Scores
Measured after scripts/train.py (seed 42) on the held-out eval set:
| slice | n | accuracy | fpr on good | blatant fail recall | p95 ms |
|---|---|---|---|---|---|
| eval | 500 | 0.860 | 0.030 | 0.956 | 0.22 |
| same-data Bayes | 500 | 0.794 | 0.758 | 0.918 | — |
Targets: blatant-fail recall at or above 0.95. False-fail rate on the good slice below 0.05. p95 latency under 200 ms on CPU.