--- license: apache-2.0 library_name: learning-report tags: - evaluation - calibration - out-of-distribution - model-diagnostics - training-data-volume - reproducibility - numpy - cpu pipeline_tag: other language: - en --- # learning-report **Two tools for making an honest claim about a model.** A model card says `"our model achieves 97% accuracy."` That number can be true and still hide four separate failures: no held-out evaluation, no OOD split, a temperature that hit the edge of the search grid, and an output distribution that has collapsed to one class. None of these show up in the single accuracy number. `learning-report` provides the missing pieces: - **A diagnostic checklist** that produces a report you cannot summarise in one number: four splits, mean confidence and ECE on each, temperature-boundary detection, output-uniformity canary, and calibration thresholds. The verdict counts the issues. - **A training-data-volume calculator** that answers "how many examples do I need?" before you run the experiment. Two estimates: a formula and an empirical ladder with two curve families (exponential saturation and Hill cooperativity). Both are pure numpy, train in seconds, and ship as code — there are no weights to download. ## The claim in one sentence A model claim that can be summarised in a single accuracy number is an incomplete claim; a complete claim names the splits, reports calibration on each, and states how many examples the target required in the first place. ## What the diagnostic reports Every `LearningReport` contains: - `parameters` — model size - `training examples` — dataset size - `training accuracy` — fit, not evidence of learning - `temperature` — fitted value, with a `[BOUNDARY]` flag if it landed on the edge of the search grid - `output uniformity` — normalized entropy of the model's prediction histogram; low values mean the model has collapsed to one class - `mean max-prob` — average top-class probability; near-uniform values are not calibration - `verdict` — `[ok]`, `[?]`, `[??]`, or `[???]` by issue count - `issues` — plain-English list of every problem detected - `splits` — four rows, each with `n`, `accuracy`, `mean_confidence`, `ece`, and an OOD marker The four splits are: | split | operands | operation | what it tests | |---|---|---|---| | interpolation | in training range | same | fit + generalization on distribution | | near-OOD | one axis extended | same | partial transfer | | far-OOD | all outside training range | same | generalization to new inputs | | structural-OOD | in training range | different | transfer to a new operation | The last split is the load-bearing one and it is almost never reported. It distinguishes "learned the sum rule" from "learned the input-output table." ## What the volume calculator reports Two independent estimates of training-set size: **Formula.** `N ≈ k · P / log2(C) · 1 / (1 - target_acc)` with `k = 0.05` calibrated on the demo task. Fast, no training, useful for the first pass. **Empirical ladder.** Train the model at a ladder of dataset sizes, fit a saturating curve to the resulting accuracy, and invert to recommend an N. Two curve families: | family | form | when it fits | |---|---|---| | exponential | `acc(N) = acc_max · (1 - exp(-N / N_half))` | smooth saturation | | Hill | `acc(N) = acc_max · N^h / (N_half^h + N^h)` | step-shaped ladder | The Hill exponent `h` is a diagnostic: `h > 1` means the transition is cooperative, and the exponential fit is systematically overestimating the required N. ## Install ```bash pip install numpy python learning_report.py ``` No downloads. No tokenizer files. No model hub. Runs in under a minute on CPU. ## Usage ### Diagnostic checklist ```python from learning_report import ( MLP, train, fit_temperature, diagnose, print_report, sample_pairs, encode_pairs, labels_add, labels_mul, ) # Train a model (or bring your own that exposes .predict_proba). model = MLP(in_dim=40, hidden=32, out_dim=40, seed=0) # Build four splits as lists of raw (a, b) tuples. interp = sample_pairs(0, 9, 0, 9, 80, seed=200) near_ood = sample_pairs(0, 9, 10, 14, 60, seed=201) far_ood = sample_pairs(10, 19, 10, 19, 60, seed=202) struct = sample_pairs(0, 5, 0, 5, 36, seed=203) # diagnose() encodes internally; pass raw pairs, not arrays. report = diagnose( model, train_pairs, y_train, interp, labels_add(interp), near_ood, labels_add(near_ood), far_ood, labels_add(far_ood), struct, labels_mul(struct), ) print_report(report) ``` ### Volume calculator ```python from learning_report import ( estimate_volume_formula, estimate_volume_empirical, make_add_data, make_add_model, ) # Formula only. n = estimate_volume_formula(n_params=2632, n_classes=40, target_acc=0.80) # → 124 # Empirical ladder with exponential fit. recommended_exp, diag_exp = estimate_volume_empirical( make_data=make_add_data, model_factory=make_add_model, N_grid=(16, 32, 64, 128, 256, 512, 1024), epochs=500, seeds=2, target_acc=0.80, form="exp", ) # Empirical ladder with Hill fit. recommended_hill, diag_hill = estimate_volume_empirical( make_data=make_add_model and make_add_data, model_factory=make_add_model, target_acc=0.80, form="hill", ) ``` ## Results from the demo The demo trains a 2,632-parameter MLP on `a + b` for `a, b ∈ [0, 9]` with 60 examples. ### Self-test **21/21 checks pass.** The load-bearing ones: | check | guards against | |---|---| | `threshold: ece fires when ece>0.20` | high ECE going unreported | | `threshold: underconf fires when conf<0.40 and acc>0.5` | temperature over-correction going unreported | | `curve fit hill beats exp on step data` | the fit family being chosen for the wrong shape | | `boundary: T=20 flagged` | temperature hitting the grid edge | ### Diagnostic report ``` parameters : 2632 training examples : 60 training accuracy : 100.0% temperature : 4.45 output uniformity : 0.658 mean max-prob : 0.244 verdict : [???] issues: - cannot generalise to OOD inputs: far-OOD accuracy 0.00 - cannot transfer to a different operation: structural-OOD accuracy 0.11 - poor calibration on interpolation: ECE=0.33 (threshold 0.20) - underconfident on interpolation (acc=0.60, conf=0.38); temperature may be over-corrected split n acc mean_conf ECE -------------------------------------------------------- interpolation 80 0.60 0.38 0.33 near-OOD 50 0.00 0.17 0.17 (OOD) far-OOD 60 0.00 0.07 0.07 (OOD) structural-OOD 36 0.11 0.39 0.28 (OOD) ``` A model that fits its training data perfectly (`100.0%` training accuracy) and generalizes to `60%` on-distribution, `0%` far-OOD, and `11%` structural-OOD. The report names four issues. A model card that reported only the training accuracy would be technically true and completely misleading. ### Volume calculator **Formula:** | target acc | estimated N | |---|---| | 0.30 | 36 | | 0.50 | 50 | | 0.70 | 83 | | 0.80 | 124 | | 0.90 | 248 | **Empirical ladder:** | N | val_acc | |---|---| | 16 | 0.165 | | 32 | 0.320 | | 64 | 0.640 | | 128 | 1.000 | | 256 | 1.000 | | 512 | 1.000 | | 1024 | 1.000 | **Two fits:** | family | ceiling | N_half | h | MSE | recommended N for 0.80 | |---|---|---|---|---|---| | exp | 1.000 | 67.7 | — | 0.0042 | 110 | | **hill** | **1.000** | **40.8** | **2.04** | **0.0027** | **81** | The Hill fit has lower MSE (0.0027 vs 0.0042), and the fitted `h = 2.04` is the signature of a step, not a smooth curve. The ladder jumps from 0.640 at N=64 to 1.000 at N=128; the exponential fit smooths over that and recommends 110, while the Hill fit recommends 81. **Verification:** | N | held-out accuracy | target met? | |---|---|---| | 81 (hill) | 0.810 | yes, exactly | | 110 (exp) | 1.000 | yes, overshot | | 124 (formula) | 1.000 | yes, overshot | All three recommendations land above the target. The Hill recommendation is the most economical — it hits the target without overshooting — because the fit correctly identified the ladder as a step rather than a smooth curve. ## The bug class `"Our model achieves X% accuracy."` The single number hides: - which split the X% was measured on - whether any OOD split was tested - whether confidence is calibrated - whether the temperature fit degenerated - whether the output distribution collapsed - how many examples the target required in the first place Every one of these can be wrong while X% is true. ## The six defenses | level | mechanism | action | |---|---|---| | 1. split | four evaluation axes | OOD is measured, not assumed | | 2. calibrate | mean confidence + ECE on each split | confidence claims are checked | | 3. canary | output uniformity | degenerate models are flagged | | 4. bound | temperature boundary check | degenerate calibration is flagged | | 5. threshold | ECE and underconfidence thresholds | calibration issues become issues | | 6. plan | volume calculator (formula + empirical) | the target N is known before training | Any evaluation can choose its level. The choice is explicit in the report. ## When to use it - **Before publishing any accuracy number.** Run the diagnostic and quote the full report, not the top line. - **Before running an expensive training job.** The volume calculator tells you how many examples the target requires. If you do not have them, you have a data problem, not a model problem. - **When comparing two models.** Same splits, same calibration metric, same OOD axes. A model that wins on interpolation and loses on structural-OOD is a different model. - **When a model's accuracy looks too good.** The diagnostic includes a formula check: if `N * log2(C) / P` is under 0.05, the model is over-parameterized for the data. ## When *not* to use it - **When the task has no meaningful OOD axis.** Some tasks are bounded by construction (a fixed vocabulary, a closed label set). The structural-OOD split requires two different operations on the same inputs; if your task has only one, skip that split. - **When the model is already validated against a public benchmark.** Public benchmarks handle splitting. Use the diagnostic for the cases the benchmark does not cover. - **As a substitute for domain expertise.** The report flags problems; it does not tell you which ones matter for your application. ## Honest limitations - **The formula constant `k = 0.05` is calibrated on one demo.** It is a rule of thumb, not a law. Re-calibrate on your own task before trusting the number. - **The empirical ladder is expensive.** Seven training runs at two seeds each. If training is expensive, use the formula and validate with a single held-out training run. - **The Hill fit is chosen by MSE, not by hypothesis.** A Hill fit beating an exp fit is evidence that the ladder has a step, not proof. The curve families are approximations. - **The four splits do not cover every OOD axis.** Adversarial inputs, distribution shift in the feature distribution, label noise, and length-based generalization are separate axes not measured here. - **The diagnostic does not know about your loss function.** ECE is meaningful for classification. For regression, use a proper scoring rule and report residuals. - **`temperature_hit_boundary` is a heuristic.** A temperature of 19.9 on a grid with upper bound 20.0 is technically inside the grid but suspiciously close. The check uses `>= upper - 1e-6`, which catches only exact boundary hits. - **The volume calculator assumes the curve shape is fixed.** Real learning curves can have plateaus, jumps, and non-monotonic regions. The two fitted families cover two shapes; other shapes are silently mis-modelled. ## What this tool demonstrates Four findings from the demo, each visible in the report: **1. Training accuracy is not evidence.** The model reaches 100% on training data and 0% on far-OOD. **2. Calibration is not accuracy.** Mean confidence is 0.38 on interpolation where accuracy is 0.60 — the model is underconfident, not overconfident, and the temperature-scaled ECE is 0.33. **3. A single number is not a report.** Four issues, all invisible in the training accuracy figure. **4. The fit family matters.** Exponential and Hill fits produce recommendations of 110 and 81 examples for the same target. The Hill fit is more economical because it correctly identifies the ladder as a step. ## Version history | version | change | |---|---| | 0.1.0 | `LearningReport`, `diagnose()`, `print_report()`, four splits | | 0.2.0 | `estimate_volume_formula()` | | 0.3.0 | `estimate_volume_empirical()`, exponential saturation fit | | 0.4.0 | temperature-boundary check, output-uniformity canary | | 0.5.0 | ECE and underconfidence thresholds in `diagnose()` | | 0.6.0 | Hill cooperativity fit; both curve families supported | ## Reference Part of a series of small tools built in one session: | tool | reads | answers | |---|---|---| | `hv-manifold` | a corpus | the geometry of style space | | `hv-reader` | one text | how it reads | | `anomaly-or-bug` | a number and a matrix | is this a bug or a discovery? | | `frontier-check` | one claim | where does it sit relative to the frontier? | | `numerical-provenance` | a numeric pipeline | can I trust this number? | | `token-provenance` | an LLM pipeline | can I trust this token count? | | `provenance-reconstruct` | a bare number | what produced this, and how sure are you? | | `math-ood` | a small model | did it learn a rule or memorize a table? | | **`learning-report`** | **a trained model** | **can I make an honest claim about what it learned?** | The design principle — that a model claim should be as structured as the provenance of a single number — came out of a session in which the same diagnostic (four splits, calibration, OOD) was applied by hand to five different models, and then written down. ## License Apache-2.0