StandardOne-3B / server /benchmarks /METRIC_VALIDATION.md
MyeongHoJeong's picture
Add files using upload-large-folder tool
ea84b46 verified
|
Raw History Blame Contribute Delete
3.8 kB

KEV metric parity validation

On 2026-09-21, the adapter's scalar quality metrics were compared numerically with the original KEV implementations at commit 4f8110a3f8620cc3a182ae9a708e4398492c4b1a. All common scalar metrics matched within 1e-12 absolute error. The largest observed difference was 8.881784197001252e-16 (score MAE). Complete contrastive pair metrics matched exactly.

This validates the metric calculations. It does not validate H200 execution, inference probabilities, model accuracy, or serving latency.

Method

The optional verify_metrics.py script reads these original files from a local KEV checkout and verifies their full SHA256 before execution:

File Extracted functions SHA256
kev/evaluate.py ece 1b20e3f9edf417aa8dae924b1526e52f74b710cadf7213c5ec68334f6e7f8fe1
kev/benchmark.py coverage_at_error, metrics 4192ec3b26b065452f84bde38a091e6854a28fe185d2e0f39a0c91b2efe69df7
kev/contrastive.py paired_flip cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf

Python AST extraction keeps only those function definitions. KEV's package, PyTorch, Transformers and model-loading code are not imported. The audit uses NumPy in its own optional environment; NumPy is not an adapter dependency. The original successful audit used NumPy 2.3.5.

Input generation uses numpy.random.default_rng(483) for 100 batches of 50 rows, totaling 5,000 rows. Rows mix choice, boolean and ordinal score questions, with 2–10 options. Probabilities come from Dirichlet distributions. The first 10 rows in each batch additionally cover binary probability endpoints and decimal boundaries, including zero and one. Both implementations receive the same rows in the same order, including confidence ties.

Compared metrics include accuracy, NLL, multiclass Brier score, ten-bin ECE, mean confidence, confidence bias, confident errors, coverage/accuracy at 0.9, coverage at 1%/5% empirical error, score MAE and ranked probability score. Ten complete two-sibling pairs additionally check relevant and invariant pair metrics. Five pairs have changed gold labels and five have unchanged labels.

Reproduce

Use a Python environment that already has NumPy installed:

git clone https://github.com/jaredpalmer/kev.git /tmp/kev-reference
git -C /tmp/kev-reference checkout 4f8110a3f8620cc3a182ae9a708e4398492c4b1a
python benchmarks/verify_metrics.py --kev-root /tmp/kev-reference

The JSON output includes the reference hashes, current adapter metric-code hash, NumPy version, sample counts, per-metric maximum differences and pair results. A source mismatch or numerical discrepancy exits unsuccessfully. Reference files are checked even if the local checkout has uncommitted changes.

Scope differences

  • Accuracy headlines use clean question rows, not all submitted records. Source unknowable is excluded from accuracy and scored for confidence.
  • Incomplete pairs are counted explicitly for smoke subsets; KEV's original pair function rejects incomplete pairs.
  • The adapter omits meaningless unknowable accuracy from grouped reports; KEV's original code includes it in some diagnostic subreports.
  • This audit compares common scalar metrics and complete pairs. It does not certify every report field, input conversion, HTTP behavior, or SemIf's separate family-balanced aggregation.

Original definitions: KEV benchmark.py, evaluate.py, contrastive.py.