learning-report / README.md
zeechimp's picture
Create README.md
62b33c7 verified
|
Raw History Blame Contribute Delete
14.2 kB
---
license: apache-2.0
library_name: learning-report
tags:
- evaluation
- calibration
- out-of-distribution
- model-diagnostics
- training-data-volume
- reproducibility
- numpy
- cpu
pipeline_tag: other
language:
- en
---
# learning-report
**Two tools for making an honest claim about a model.**
A model card says `"our model achieves 97% accuracy."` That number can
be true and still hide four separate failures: no held-out evaluation,
no OOD split, a temperature that hit the edge of the search grid, and
an output distribution that has collapsed to one class. None of these
show up in the single accuracy number.
`learning-report` provides the missing pieces:
- **A diagnostic checklist** that produces a report you cannot
summarise in one number: four splits, mean confidence and ECE on
each, temperature-boundary detection, output-uniformity canary,
and calibration thresholds. The verdict counts the issues.
- **A training-data-volume calculator** that answers "how many
examples do I need?" before you run the experiment. Two
estimates: a formula and an empirical ladder with two curve
families (exponential saturation and Hill cooperativity).
Both are pure numpy, train in seconds, and ship as code — there are
no weights to download.
## The claim in one sentence
A model claim that can be summarised in a single accuracy number is
an incomplete claim; a complete claim names the splits, reports
calibration on each, and states how many examples the target
required in the first place.
## What the diagnostic reports
Every `LearningReport` contains:
- `parameters` — model size
- `training examples` — dataset size
- `training accuracy` — fit, not evidence of learning
- `temperature` — fitted value, with a `[BOUNDARY]` flag if it landed
on the edge of the search grid
- `output uniformity` — normalized entropy of the model's prediction
histogram; low values mean the model has collapsed to one class
- `mean max-prob` — average top-class probability; near-uniform
values are not calibration
- `verdict` — `[ok]`, `[?]`, `[??]`, or `[???]` by issue count
- `issues` — plain-English list of every problem detected
- `splits` — four rows, each with `n`, `accuracy`, `mean_confidence`,
`ece`, and an OOD marker
The four splits are:
| split | operands | operation | what it tests |
|---|---|---|---|
| interpolation | in training range | same | fit + generalization on distribution |
| near-OOD | one axis extended | same | partial transfer |
| far-OOD | all outside training range | same | generalization to new inputs |
| structural-OOD | in training range | different | transfer to a new operation |
The last split is the load-bearing one and it is almost never
reported. It distinguishes "learned the sum rule" from "learned the
input-output table."
## What the volume calculator reports
Two independent estimates of training-set size:
**Formula.** `N ≈ k · P / log2(C) · 1 / (1 - target_acc)` with
`k = 0.05` calibrated on the demo task. Fast, no training, useful for
the first pass.
**Empirical ladder.** Train the model at a ladder of dataset sizes,
fit a saturating curve to the resulting accuracy, and invert to
recommend an N. Two curve families:
| family | form | when it fits |
|---|---|---|
| exponential | `acc(N) = acc_max · (1 - exp(-N / N_half))` | smooth saturation |
| Hill | `acc(N) = acc_max · N^h / (N_half^h + N^h)` | step-shaped ladder |
The Hill exponent `h` is a diagnostic: `h > 1` means the transition
is cooperative, and the exponential fit is systematically
overestimating the required N.
## Install
```bash
pip install numpy
python learning_report.py
```
No downloads. No tokenizer files. No model hub. Runs in under a
minute on CPU.
## Usage
### Diagnostic checklist
```python
from learning_report import (
MLP, train, fit_temperature, diagnose, print_report,
sample_pairs, encode_pairs, labels_add, labels_mul,
)
# Train a model (or bring your own that exposes .predict_proba).
model = MLP(in_dim=40, hidden=32, out_dim=40, seed=0)
# Build four splits as lists of raw (a, b) tuples.
interp = sample_pairs(0, 9, 0, 9, 80, seed=200)
near_ood = sample_pairs(0, 9, 10, 14, 60, seed=201)
far_ood = sample_pairs(10, 19, 10, 19, 60, seed=202)
struct = sample_pairs(0, 5, 0, 5, 36, seed=203)
# diagnose() encodes internally; pass raw pairs, not arrays.
report = diagnose(
model,
train_pairs, y_train,
interp, labels_add(interp),
near_ood, labels_add(near_ood),
far_ood, labels_add(far_ood),
struct, labels_mul(struct),
)
print_report(report)
```
### Volume calculator
```python
from learning_report import (
estimate_volume_formula,
estimate_volume_empirical,
make_add_data,
make_add_model,
)
# Formula only.
n = estimate_volume_formula(n_params=2632, n_classes=40,
target_acc=0.80) # → 124
# Empirical ladder with exponential fit.
recommended_exp, diag_exp = estimate_volume_empirical(
make_data=make_add_data,
model_factory=make_add_model,
N_grid=(16, 32, 64, 128, 256, 512, 1024),
epochs=500,
seeds=2,
target_acc=0.80,
form="exp",
)
# Empirical ladder with Hill fit.
recommended_hill, diag_hill = estimate_volume_empirical(
make_data=make_add_model and make_add_data,
model_factory=make_add_model,
target_acc=0.80,
form="hill",
)
```
## Results from the demo
The demo trains a 2,632-parameter MLP on `a + b` for `a, b ∈ [0, 9]`
with 60 examples.
### Self-test
**21/21 checks pass.** The load-bearing ones:
| check | guards against |
|---|---|
| `threshold: ece fires when ece>0.20` | high ECE going unreported |
| `threshold: underconf fires when conf<0.40 and acc>0.5` | temperature over-correction going unreported |
| `curve fit hill beats exp on step data` | the fit family being chosen for the wrong shape |
| `boundary: T=20 flagged` | temperature hitting the grid edge |
### Diagnostic report
```
parameters : 2632
training examples : 60
training accuracy : 100.0%
temperature : 4.45
output uniformity : 0.658
mean max-prob : 0.244
verdict : [???]
issues:
- cannot generalise to OOD inputs: far-OOD accuracy 0.00
- cannot transfer to a different operation: structural-OOD accuracy 0.11
- poor calibration on interpolation: ECE=0.33 (threshold 0.20)
- underconfident on interpolation (acc=0.60, conf=0.38); temperature may be over-corrected
split n acc mean_conf ECE
--------------------------------------------------------
interpolation 80 0.60 0.38 0.33
near-OOD 50 0.00 0.17 0.17 (OOD)
far-OOD 60 0.00 0.07 0.07 (OOD)
structural-OOD 36 0.11 0.39 0.28 (OOD)
```
A model that fits its training data perfectly (`100.0%` training
accuracy) and generalizes to `60%` on-distribution, `0%` far-OOD,
and `11%` structural-OOD. The report names four issues. A model card
that reported only the training accuracy would be technically true
and completely misleading.
### Volume calculator
**Formula:**
| target acc | estimated N |
|---|---|
| 0.30 | 36 |
| 0.50 | 50 |
| 0.70 | 83 |
| 0.80 | 124 |
| 0.90 | 248 |
**Empirical ladder:**
| N | val_acc |
|---|---|
| 16 | 0.165 |
| 32 | 0.320 |
| 64 | 0.640 |
| 128 | 1.000 |
| 256 | 1.000 |
| 512 | 1.000 |
| 1024 | 1.000 |
**Two fits:**
| family | ceiling | N_half | h | MSE | recommended N for 0.80 |
|---|---|---|---|---|---|
| exp | 1.000 | 67.7 | — | 0.0042 | 110 |
| **hill** | **1.000** | **40.8** | **2.04** | **0.0027** | **81** |
The Hill fit has lower MSE (0.0027 vs 0.0042), and the fitted `h = 2.04`
is the signature of a step, not a smooth curve. The ladder jumps from
0.640 at N=64 to 1.000 at N=128; the exponential fit smooths over
that and recommends 110, while the Hill fit recommends 81.
**Verification:**
| N | held-out accuracy | target met? |
|---|---|---|
| 81 (hill) | 0.810 | yes, exactly |
| 110 (exp) | 1.000 | yes, overshot |
| 124 (formula) | 1.000 | yes, overshot |
All three recommendations land above the target. The Hill
recommendation is the most economical — it hits the target without
overshooting — because the fit correctly identified the ladder as a
step rather than a smooth curve.
## The bug class
`"Our model achieves X% accuracy."`
The single number hides:
- which split the X% was measured on
- whether any OOD split was tested
- whether confidence is calibrated
- whether the temperature fit degenerated
- whether the output distribution collapsed
- how many examples the target required in the first place
Every one of these can be wrong while X% is true.
## The six defenses
| level | mechanism | action |
|---|---|---|
| 1. split | four evaluation axes | OOD is measured, not assumed |
| 2. calibrate | mean confidence + ECE on each split | confidence claims are checked |
| 3. canary | output uniformity | degenerate models are flagged |
| 4. bound | temperature boundary check | degenerate calibration is flagged |
| 5. threshold | ECE and underconfidence thresholds | calibration issues become issues |
| 6. plan | volume calculator (formula + empirical) | the target N is known before training |
Any evaluation can choose its level. The choice is explicit in the
report.
## When to use it
- **Before publishing any accuracy number.** Run the diagnostic and
quote the full report, not the top line.
- **Before running an expensive training job.** The volume
calculator tells you how many examples the target requires. If
you do not have them, you have a data problem, not a model
problem.
- **When comparing two models.** Same splits, same calibration
metric, same OOD axes. A model that wins on interpolation and
loses on structural-OOD is a different model.
- **When a model's accuracy looks too good.** The diagnostic
includes a formula check: if `N * log2(C) / P` is under 0.05, the
model is over-parameterized for the data.
## When *not* to use it
- **When the task has no meaningful OOD axis.** Some tasks are
bounded by construction (a fixed vocabulary, a closed label set).
The structural-OOD split requires two different operations on the
same inputs; if your task has only one, skip that split.
- **When the model is already validated against a public benchmark.**
Public benchmarks handle splitting. Use the diagnostic for the
cases the benchmark does not cover.
- **As a substitute for domain expertise.** The report flags
problems; it does not tell you which ones matter for your
application.
## Honest limitations
- **The formula constant `k = 0.05` is calibrated on one demo.**
It is a rule of thumb, not a law. Re-calibrate on your own task
before trusting the number.
- **The empirical ladder is expensive.** Seven training runs at
two seeds each. If training is expensive, use the formula and
validate with a single held-out training run.
- **The Hill fit is chosen by MSE, not by hypothesis.** A Hill fit
beating an exp fit is evidence that the ladder has a step, not
proof. The curve families are approximations.
- **The four splits do not cover every OOD axis.** Adversarial
inputs, distribution shift in the feature distribution, label
noise, and length-based generalization are separate axes not
measured here.
- **The diagnostic does not know about your loss function.** ECE is
meaningful for classification. For regression, use a proper
scoring rule and report residuals.
- **`temperature_hit_boundary` is a heuristic.** A temperature of
19.9 on a grid with upper bound 20.0 is technically inside the
grid but suspiciously close. The check uses `>= upper - 1e-6`,
which catches only exact boundary hits.
- **The volume calculator assumes the curve shape is fixed.** Real
learning curves can have plateaus, jumps, and non-monotonic
regions. The two fitted families cover two shapes; other shapes
are silently mis-modelled.
## What this tool demonstrates
Four findings from the demo, each visible in the report:
**1. Training accuracy is not evidence.** The model reaches 100%
on training data and 0% on far-OOD.
**2. Calibration is not accuracy.** Mean confidence is 0.38 on
interpolation where accuracy is 0.60 — the model is
underconfident, not overconfident, and the temperature-scaled ECE
is 0.33.
**3. A single number is not a report.** Four issues, all invisible
in the training accuracy figure.
**4. The fit family matters.** Exponential and Hill fits produce
recommendations of 110 and 81 examples for the same target. The
Hill fit is more economical because it correctly identifies the
ladder as a step.
## Version history
| version | change |
|---|---|
| 0.1.0 | `LearningReport`, `diagnose()`, `print_report()`, four splits |
| 0.2.0 | `estimate_volume_formula()` |
| 0.3.0 | `estimate_volume_empirical()`, exponential saturation fit |
| 0.4.0 | temperature-boundary check, output-uniformity canary |
| 0.5.0 | ECE and underconfidence thresholds in `diagnose()` |
| 0.6.0 | Hill cooperativity fit; both curve families supported |
## Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
| `hv-manifold` | a corpus | the geometry of style space |
| `hv-reader` | one text | how it reads |
| `anomaly-or-bug` | a number and a matrix | is this a bug or a discovery? |
| `frontier-check` | one claim | where does it sit relative to the frontier? |
| `numerical-provenance` | a numeric pipeline | can I trust this number? |
| `token-provenance` | an LLM pipeline | can I trust this token count? |
| `provenance-reconstruct` | a bare number | what produced this, and how sure are you? |
| `math-ood` | a small model | did it learn a rule or memorize a table? |
| **`learning-report`** | **a trained model** | **can I make an honest claim about what it learned?** |
The design principle — that a model claim should be as structured
as the provenance of a single number — came out of a session in
which the same diagnostic (four splits, calibration, OOD) was
applied by hand to five different models, and then written down.
## License
Apache-2.0