learning-report

Two tools for making an honest claim about a model.

A model card says "our model achieves 97% accuracy." That number can be true and still hide four separate failures: no held-out evaluation, no OOD split, a temperature that hit the edge of the search grid, and an output distribution that has collapsed to one class. None of these show up in the single accuracy number.

learning-report provides the missing pieces:

  • A diagnostic checklist that produces a report you cannot summarise in one number: four splits, mean confidence and ECE on each, temperature-boundary detection, output-uniformity canary, and calibration thresholds. The verdict counts the issues.
  • A training-data-volume calculator that answers "how many examples do I need?" before you run the experiment. Two estimates: a formula and an empirical ladder with two curve families (exponential saturation and Hill cooperativity).

Both are pure numpy, train in seconds, and ship as code β€” there are no weights to download.

The claim in one sentence

A model claim that can be summarised in a single accuracy number is an incomplete claim; a complete claim names the splits, reports calibration on each, and states how many examples the target required in the first place.

What the diagnostic reports

Every LearningReport contains:

  • parameters β€” model size
  • training examples β€” dataset size
  • training accuracy β€” fit, not evidence of learning
  • temperature β€” fitted value, with a [BOUNDARY] flag if it landed on the edge of the search grid
  • output uniformity β€” normalized entropy of the model's prediction histogram; low values mean the model has collapsed to one class
  • mean max-prob β€” average top-class probability; near-uniform values are not calibration
  • verdict β€” [ok], [?], [??], or [???] by issue count
  • issues β€” plain-English list of every problem detected
  • splits β€” four rows, each with n, accuracy, mean_confidence, ece, and an OOD marker

The four splits are:

split operands operation what it tests
interpolation in training range same fit + generalization on distribution
near-OOD one axis extended same partial transfer
far-OOD all outside training range same generalization to new inputs
structural-OOD in training range different transfer to a new operation

The last split is the load-bearing one and it is almost never reported. It distinguishes "learned the sum rule" from "learned the input-output table."

What the volume calculator reports

Two independent estimates of training-set size:

Formula. N β‰ˆ k Β· P / log2(C) Β· 1 / (1 - target_acc) with k = 0.05 calibrated on the demo task. Fast, no training, useful for the first pass.

Empirical ladder. Train the model at a ladder of dataset sizes, fit a saturating curve to the resulting accuracy, and invert to recommend an N. Two curve families:

family form when it fits
exponential acc(N) = acc_max Β· (1 - exp(-N / N_half)) smooth saturation
Hill acc(N) = acc_max Β· N^h / (N_half^h + N^h) step-shaped ladder

The Hill exponent h is a diagnostic: h > 1 means the transition is cooperative, and the exponential fit is systematically overestimating the required N.

Install

pip install numpy
python learning_report.py

No downloads. No tokenizer files. No model hub. Runs in under a minute on CPU.

Usage

Diagnostic checklist

from learning_report import (
    MLP, train, fit_temperature, diagnose, print_report,
    sample_pairs, encode_pairs, labels_add, labels_mul,
)

# Train a model (or bring your own that exposes .predict_proba).
model = MLP(in_dim=40, hidden=32, out_dim=40, seed=0)

# Build four splits as lists of raw (a, b) tuples.
interp   = sample_pairs(0, 9, 0, 9, 80, seed=200)
near_ood = sample_pairs(0, 9, 10, 14, 60, seed=201)
far_ood  = sample_pairs(10, 19, 10, 19, 60, seed=202)
struct   = sample_pairs(0, 5, 0, 5, 36, seed=203)

# diagnose() encodes internally; pass raw pairs, not arrays.
report = diagnose(
    model,
    train_pairs, y_train,
    interp, labels_add(interp),
    near_ood, labels_add(near_ood),
    far_ood, labels_add(far_ood),
    struct, labels_mul(struct),
)
print_report(report)

Volume calculator

from learning_report import (
    estimate_volume_formula,
    estimate_volume_empirical,
    make_add_data,
    make_add_model,
)

# Formula only.
n = estimate_volume_formula(n_params=2632, n_classes=40,
                            target_acc=0.80)  # β†’ 124

# Empirical ladder with exponential fit.
recommended_exp, diag_exp = estimate_volume_empirical(
    make_data=make_add_data,
    model_factory=make_add_model,
    N_grid=(16, 32, 64, 128, 256, 512, 1024),
    epochs=500,
    seeds=2,
    target_acc=0.80,
    form="exp",
)

# Empirical ladder with Hill fit.
recommended_hill, diag_hill = estimate_volume_empirical(
    make_data=make_add_model and make_add_data,
    model_factory=make_add_model,
    target_acc=0.80,
    form="hill",
)

Results from the demo

The demo trains a 2,632-parameter MLP on a + b for a, b ∈ [0, 9] with 60 examples.

Self-test

21/21 checks pass. The load-bearing ones:

check guards against
threshold: ece fires when ece>0.20 high ECE going unreported
threshold: underconf fires when conf<0.40 and acc>0.5 temperature over-correction going unreported
curve fit hill beats exp on step data the fit family being chosen for the wrong shape
boundary: T=20 flagged temperature hitting the grid edge

Diagnostic report

parameters                : 2632
training examples         : 60
training accuracy         : 100.0%
temperature               : 4.45
output uniformity         : 0.658
mean max-prob             : 0.244
verdict                   : [???]

issues:
  - cannot generalise to OOD inputs: far-OOD accuracy 0.00
  - cannot transfer to a different operation: structural-OOD accuracy 0.11
  - poor calibration on interpolation: ECE=0.33 (threshold 0.20)
  - underconfident on interpolation (acc=0.60, conf=0.38); temperature may be over-corrected

split                  n     acc   mean_conf     ECE
--------------------------------------------------------
interpolation         80    0.60        0.38    0.33
near-OOD              50    0.00        0.17    0.17  (OOD)
far-OOD               60    0.00        0.07    0.07  (OOD)
structural-OOD        36    0.11        0.39    0.28  (OOD)

A model that fits its training data perfectly (100.0% training accuracy) and generalizes to 60% on-distribution, 0% far-OOD, and 11% structural-OOD. The report names four issues. A model card that reported only the training accuracy would be technically true and completely misleading.

Volume calculator

Formula:

target acc estimated N
0.30 36
0.50 50
0.70 83
0.80 124
0.90 248

Empirical ladder:

N val_acc
16 0.165
32 0.320
64 0.640
128 1.000
256 1.000
512 1.000
1024 1.000

Two fits:

family ceiling N_half h MSE recommended N for 0.80
exp 1.000 67.7 β€” 0.0042 110
hill 1.000 40.8 2.04 0.0027 81

The Hill fit has lower MSE (0.0027 vs 0.0042), and the fitted h = 2.04 is the signature of a step, not a smooth curve. The ladder jumps from 0.640 at N=64 to 1.000 at N=128; the exponential fit smooths over that and recommends 110, while the Hill fit recommends 81.

Verification:

N held-out accuracy target met?
81 (hill) 0.810 yes, exactly
110 (exp) 1.000 yes, overshot
124 (formula) 1.000 yes, overshot

All three recommendations land above the target. The Hill recommendation is the most economical β€” it hits the target without overshooting β€” because the fit correctly identified the ladder as a step rather than a smooth curve.

The bug class

"Our model achieves X% accuracy."

The single number hides:

  • which split the X% was measured on
  • whether any OOD split was tested
  • whether confidence is calibrated
  • whether the temperature fit degenerated
  • whether the output distribution collapsed
  • how many examples the target required in the first place

Every one of these can be wrong while X% is true.

The six defenses

level mechanism action
1. split four evaluation axes OOD is measured, not assumed
2. calibrate mean confidence + ECE on each split confidence claims are checked
3. canary output uniformity degenerate models are flagged
4. bound temperature boundary check degenerate calibration is flagged
5. threshold ECE and underconfidence thresholds calibration issues become issues
6. plan volume calculator (formula + empirical) the target N is known before training

Any evaluation can choose its level. The choice is explicit in the report.

When to use it

  • Before publishing any accuracy number. Run the diagnostic and quote the full report, not the top line.
  • Before running an expensive training job. The volume calculator tells you how many examples the target requires. If you do not have them, you have a data problem, not a model problem.
  • When comparing two models. Same splits, same calibration metric, same OOD axes. A model that wins on interpolation and loses on structural-OOD is a different model.
  • When a model's accuracy looks too good. The diagnostic includes a formula check: if N * log2(C) / P is under 0.05, the model is over-parameterized for the data.

When not to use it

  • When the task has no meaningful OOD axis. Some tasks are bounded by construction (a fixed vocabulary, a closed label set). The structural-OOD split requires two different operations on the same inputs; if your task has only one, skip that split.
  • When the model is already validated against a public benchmark. Public benchmarks handle splitting. Use the diagnostic for the cases the benchmark does not cover.
  • As a substitute for domain expertise. The report flags problems; it does not tell you which ones matter for your application.

Honest limitations

  • The formula constant k = 0.05 is calibrated on one demo. It is a rule of thumb, not a law. Re-calibrate on your own task before trusting the number.
  • The empirical ladder is expensive. Seven training runs at two seeds each. If training is expensive, use the formula and validate with a single held-out training run.
  • The Hill fit is chosen by MSE, not by hypothesis. A Hill fit beating an exp fit is evidence that the ladder has a step, not proof. The curve families are approximations.
  • The four splits do not cover every OOD axis. Adversarial inputs, distribution shift in the feature distribution, label noise, and length-based generalization are separate axes not measured here.
  • The diagnostic does not know about your loss function. ECE is meaningful for classification. For regression, use a proper scoring rule and report residuals.
  • temperature_hit_boundary is a heuristic. A temperature of 19.9 on a grid with upper bound 20.0 is technically inside the grid but suspiciously close. The check uses >= upper - 1e-6, which catches only exact boundary hits.
  • The volume calculator assumes the curve shape is fixed. Real learning curves can have plateaus, jumps, and non-monotonic regions. The two fitted families cover two shapes; other shapes are silently mis-modelled.

What this tool demonstrates

Four findings from the demo, each visible in the report:

1. Training accuracy is not evidence. The model reaches 100% on training data and 0% on far-OOD.

2. Calibration is not accuracy. Mean confidence is 0.38 on interpolation where accuracy is 0.60 β€” the model is underconfident, not overconfident, and the temperature-scaled ECE is 0.33.

3. A single number is not a report. Four issues, all invisible in the training accuracy figure.

4. The fit family matters. Exponential and Hill fits produce recommendations of 110 and 81 examples for the same target. The Hill fit is more economical because it correctly identifies the ladder as a step.

Version history

version change
0.1.0 LearningReport, diagnose(), print_report(), four splits
0.2.0 estimate_volume_formula()
0.3.0 estimate_volume_empirical(), exponential saturation fit
0.4.0 temperature-boundary check, output-uniformity canary
0.5.0 ECE and underconfidence thresholds in diagnose()
0.6.0 Hill cooperativity fit; both curve families supported

Reference

Part of a series of small tools built in one session:

tool reads answers
hv-manifold a corpus the geometry of style space
hv-reader one text how it reads
anomaly-or-bug a number and a matrix is this a bug or a discovery?
frontier-check one claim where does it sit relative to the frontier?
numerical-provenance a numeric pipeline can I trust this number?
token-provenance an LLM pipeline can I trust this token count?
provenance-reconstruct a bare number what produced this, and how sure are you?
math-ood a small model did it learn a rule or memorize a table?
learning-report a trained model can I make an honest claim about what it learned?

The design principle β€” that a model claim should be as structured as the provenance of a single number β€” came out of a session in which the same diagnostic (four splits, calibration, OOD) was applied by hand to five different models, and then written down.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support