learning-report
Two tools for making an honest claim about a model.
A model card says "our model achieves 97% accuracy." That number can
be true and still hide four separate failures: no held-out evaluation,
no OOD split, a temperature that hit the edge of the search grid, and
an output distribution that has collapsed to one class. None of these
show up in the single accuracy number.
learning-report provides the missing pieces:
- A diagnostic checklist that produces a report you cannot summarise in one number: four splits, mean confidence and ECE on each, temperature-boundary detection, output-uniformity canary, and calibration thresholds. The verdict counts the issues.
- A training-data-volume calculator that answers "how many examples do I need?" before you run the experiment. Two estimates: a formula and an empirical ladder with two curve families (exponential saturation and Hill cooperativity).
Both are pure numpy, train in seconds, and ship as code β there are no weights to download.
The claim in one sentence
A model claim that can be summarised in a single accuracy number is an incomplete claim; a complete claim names the splits, reports calibration on each, and states how many examples the target required in the first place.
What the diagnostic reports
Every LearningReport contains:
parametersβ model sizetraining examplesβ dataset sizetraining accuracyβ fit, not evidence of learningtemperatureβ fitted value, with a[BOUNDARY]flag if it landed on the edge of the search gridoutput uniformityβ normalized entropy of the model's prediction histogram; low values mean the model has collapsed to one classmean max-probβ average top-class probability; near-uniform values are not calibrationverdictβ[ok],[?],[??], or[???]by issue countissuesβ plain-English list of every problem detectedsplitsβ four rows, each withn,accuracy,mean_confidence,ece, and an OOD marker
The four splits are:
| split | operands | operation | what it tests |
|---|---|---|---|
| interpolation | in training range | same | fit + generalization on distribution |
| near-OOD | one axis extended | same | partial transfer |
| far-OOD | all outside training range | same | generalization to new inputs |
| structural-OOD | in training range | different | transfer to a new operation |
The last split is the load-bearing one and it is almost never reported. It distinguishes "learned the sum rule" from "learned the input-output table."
What the volume calculator reports
Two independent estimates of training-set size:
Formula. N β k Β· P / log2(C) Β· 1 / (1 - target_acc) with
k = 0.05 calibrated on the demo task. Fast, no training, useful for
the first pass.
Empirical ladder. Train the model at a ladder of dataset sizes, fit a saturating curve to the resulting accuracy, and invert to recommend an N. Two curve families:
| family | form | when it fits |
|---|---|---|
| exponential | acc(N) = acc_max Β· (1 - exp(-N / N_half)) |
smooth saturation |
| Hill | acc(N) = acc_max Β· N^h / (N_half^h + N^h) |
step-shaped ladder |
The Hill exponent h is a diagnostic: h > 1 means the transition
is cooperative, and the exponential fit is systematically
overestimating the required N.
Install
pip install numpy
python learning_report.py
No downloads. No tokenizer files. No model hub. Runs in under a minute on CPU.
Usage
Diagnostic checklist
from learning_report import (
MLP, train, fit_temperature, diagnose, print_report,
sample_pairs, encode_pairs, labels_add, labels_mul,
)
# Train a model (or bring your own that exposes .predict_proba).
model = MLP(in_dim=40, hidden=32, out_dim=40, seed=0)
# Build four splits as lists of raw (a, b) tuples.
interp = sample_pairs(0, 9, 0, 9, 80, seed=200)
near_ood = sample_pairs(0, 9, 10, 14, 60, seed=201)
far_ood = sample_pairs(10, 19, 10, 19, 60, seed=202)
struct = sample_pairs(0, 5, 0, 5, 36, seed=203)
# diagnose() encodes internally; pass raw pairs, not arrays.
report = diagnose(
model,
train_pairs, y_train,
interp, labels_add(interp),
near_ood, labels_add(near_ood),
far_ood, labels_add(far_ood),
struct, labels_mul(struct),
)
print_report(report)
Volume calculator
from learning_report import (
estimate_volume_formula,
estimate_volume_empirical,
make_add_data,
make_add_model,
)
# Formula only.
n = estimate_volume_formula(n_params=2632, n_classes=40,
target_acc=0.80) # β 124
# Empirical ladder with exponential fit.
recommended_exp, diag_exp = estimate_volume_empirical(
make_data=make_add_data,
model_factory=make_add_model,
N_grid=(16, 32, 64, 128, 256, 512, 1024),
epochs=500,
seeds=2,
target_acc=0.80,
form="exp",
)
# Empirical ladder with Hill fit.
recommended_hill, diag_hill = estimate_volume_empirical(
make_data=make_add_model and make_add_data,
model_factory=make_add_model,
target_acc=0.80,
form="hill",
)
Results from the demo
The demo trains a 2,632-parameter MLP on a + b for a, b β [0, 9]
with 60 examples.
Self-test
21/21 checks pass. The load-bearing ones:
| check | guards against |
|---|---|
threshold: ece fires when ece>0.20 |
high ECE going unreported |
threshold: underconf fires when conf<0.40 and acc>0.5 |
temperature over-correction going unreported |
curve fit hill beats exp on step data |
the fit family being chosen for the wrong shape |
boundary: T=20 flagged |
temperature hitting the grid edge |
Diagnostic report
parameters : 2632
training examples : 60
training accuracy : 100.0%
temperature : 4.45
output uniformity : 0.658
mean max-prob : 0.244
verdict : [???]
issues:
- cannot generalise to OOD inputs: far-OOD accuracy 0.00
- cannot transfer to a different operation: structural-OOD accuracy 0.11
- poor calibration on interpolation: ECE=0.33 (threshold 0.20)
- underconfident on interpolation (acc=0.60, conf=0.38); temperature may be over-corrected
split n acc mean_conf ECE
--------------------------------------------------------
interpolation 80 0.60 0.38 0.33
near-OOD 50 0.00 0.17 0.17 (OOD)
far-OOD 60 0.00 0.07 0.07 (OOD)
structural-OOD 36 0.11 0.39 0.28 (OOD)
A model that fits its training data perfectly (100.0% training
accuracy) and generalizes to 60% on-distribution, 0% far-OOD,
and 11% structural-OOD. The report names four issues. A model card
that reported only the training accuracy would be technically true
and completely misleading.
Volume calculator
Formula:
| target acc | estimated N |
|---|---|
| 0.30 | 36 |
| 0.50 | 50 |
| 0.70 | 83 |
| 0.80 | 124 |
| 0.90 | 248 |
Empirical ladder:
| N | val_acc |
|---|---|
| 16 | 0.165 |
| 32 | 0.320 |
| 64 | 0.640 |
| 128 | 1.000 |
| 256 | 1.000 |
| 512 | 1.000 |
| 1024 | 1.000 |
Two fits:
| family | ceiling | N_half | h | MSE | recommended N for 0.80 |
|---|---|---|---|---|---|
| exp | 1.000 | 67.7 | β | 0.0042 | 110 |
| hill | 1.000 | 40.8 | 2.04 | 0.0027 | 81 |
The Hill fit has lower MSE (0.0027 vs 0.0042), and the fitted h = 2.04
is the signature of a step, not a smooth curve. The ladder jumps from
0.640 at N=64 to 1.000 at N=128; the exponential fit smooths over
that and recommends 110, while the Hill fit recommends 81.
Verification:
| N | held-out accuracy | target met? |
|---|---|---|
| 81 (hill) | 0.810 | yes, exactly |
| 110 (exp) | 1.000 | yes, overshot |
| 124 (formula) | 1.000 | yes, overshot |
All three recommendations land above the target. The Hill recommendation is the most economical β it hits the target without overshooting β because the fit correctly identified the ladder as a step rather than a smooth curve.
The bug class
"Our model achieves X% accuracy."
The single number hides:
- which split the X% was measured on
- whether any OOD split was tested
- whether confidence is calibrated
- whether the temperature fit degenerated
- whether the output distribution collapsed
- how many examples the target required in the first place
Every one of these can be wrong while X% is true.
The six defenses
| level | mechanism | action |
|---|---|---|
| 1. split | four evaluation axes | OOD is measured, not assumed |
| 2. calibrate | mean confidence + ECE on each split | confidence claims are checked |
| 3. canary | output uniformity | degenerate models are flagged |
| 4. bound | temperature boundary check | degenerate calibration is flagged |
| 5. threshold | ECE and underconfidence thresholds | calibration issues become issues |
| 6. plan | volume calculator (formula + empirical) | the target N is known before training |
Any evaluation can choose its level. The choice is explicit in the report.
When to use it
- Before publishing any accuracy number. Run the diagnostic and quote the full report, not the top line.
- Before running an expensive training job. The volume calculator tells you how many examples the target requires. If you do not have them, you have a data problem, not a model problem.
- When comparing two models. Same splits, same calibration metric, same OOD axes. A model that wins on interpolation and loses on structural-OOD is a different model.
- When a model's accuracy looks too good. The diagnostic
includes a formula check: if
N * log2(C) / Pis under 0.05, the model is over-parameterized for the data.
When not to use it
- When the task has no meaningful OOD axis. Some tasks are bounded by construction (a fixed vocabulary, a closed label set). The structural-OOD split requires two different operations on the same inputs; if your task has only one, skip that split.
- When the model is already validated against a public benchmark. Public benchmarks handle splitting. Use the diagnostic for the cases the benchmark does not cover.
- As a substitute for domain expertise. The report flags problems; it does not tell you which ones matter for your application.
Honest limitations
- The formula constant
k = 0.05is calibrated on one demo. It is a rule of thumb, not a law. Re-calibrate on your own task before trusting the number. - The empirical ladder is expensive. Seven training runs at two seeds each. If training is expensive, use the formula and validate with a single held-out training run.
- The Hill fit is chosen by MSE, not by hypothesis. A Hill fit beating an exp fit is evidence that the ladder has a step, not proof. The curve families are approximations.
- The four splits do not cover every OOD axis. Adversarial inputs, distribution shift in the feature distribution, label noise, and length-based generalization are separate axes not measured here.
- The diagnostic does not know about your loss function. ECE is meaningful for classification. For regression, use a proper scoring rule and report residuals.
temperature_hit_boundaryis a heuristic. A temperature of 19.9 on a grid with upper bound 20.0 is technically inside the grid but suspiciously close. The check uses>= upper - 1e-6, which catches only exact boundary hits.- The volume calculator assumes the curve shape is fixed. Real learning curves can have plateaus, jumps, and non-monotonic regions. The two fitted families cover two shapes; other shapes are silently mis-modelled.
What this tool demonstrates
Four findings from the demo, each visible in the report:
1. Training accuracy is not evidence. The model reaches 100% on training data and 0% on far-OOD.
2. Calibration is not accuracy. Mean confidence is 0.38 on interpolation where accuracy is 0.60 β the model is underconfident, not overconfident, and the temperature-scaled ECE is 0.33.
3. A single number is not a report. Four issues, all invisible in the training accuracy figure.
4. The fit family matters. Exponential and Hill fits produce recommendations of 110 and 81 examples for the same target. The Hill fit is more economical because it correctly identifies the ladder as a step.
Version history
| version | change |
|---|---|
| 0.1.0 | LearningReport, diagnose(), print_report(), four splits |
| 0.2.0 | estimate_volume_formula() |
| 0.3.0 | estimate_volume_empirical(), exponential saturation fit |
| 0.4.0 | temperature-boundary check, output-uniformity canary |
| 0.5.0 | ECE and underconfidence thresholds in diagnose() |
| 0.6.0 | Hill cooperativity fit; both curve families supported |
Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
hv-manifold |
a corpus | the geometry of style space |
hv-reader |
one text | how it reads |
anomaly-or-bug |
a number and a matrix | is this a bug or a discovery? |
frontier-check |
one claim | where does it sit relative to the frontier? |
numerical-provenance |
a numeric pipeline | can I trust this number? |
token-provenance |
an LLM pipeline | can I trust this token count? |
provenance-reconstruct |
a bare number | what produced this, and how sure are you? |
math-ood |
a small model | did it learn a rule or memorize a table? |
learning-report |
a trained model | can I make an honest claim about what it learned? |
The design principle β that a model claim should be as structured as the provenance of a single number β came out of a session in which the same diagnostic (four splits, calibration, OOD) was applied by hand to five different models, and then written down.
License
Apache-2.0