|
Download README.md from zeechimp/learning-report: direct link, hf CLI and curl.
- Browser
- Download file 14.2 kB
-
https://huggingface.co/zeechimp/learning-report/resolve/main/README.md
- Command line
-
hf download hf://zeechimp/learning-report/README.md
-
curl -L -o README.md https://huggingface.co/zeechimp/learning-report/resolve/main/README.md
14.2 kB
| license: apache-2.0 | |
| library_name: learning-report | |
| tags: | |
| - evaluation | |
| - calibration | |
| - out-of-distribution | |
| - model-diagnostics | |
| - training-data-volume | |
| - reproducibility | |
| - numpy | |
| - cpu | |
| pipeline_tag: other | |
| language: | |
| - en | |
| # learning-report | |
| **Two tools for making an honest claim about a model.** | |
| A model card says `"our model achieves 97% accuracy."` That number can | |
| be true and still hide four separate failures: no held-out evaluation, | |
| no OOD split, a temperature that hit the edge of the search grid, and | |
| an output distribution that has collapsed to one class. None of these | |
| show up in the single accuracy number. | |
| `learning-report` provides the missing pieces: | |
| - **A diagnostic checklist** that produces a report you cannot | |
| summarise in one number: four splits, mean confidence and ECE on | |
| each, temperature-boundary detection, output-uniformity canary, | |
| and calibration thresholds. The verdict counts the issues. | |
| - **A training-data-volume calculator** that answers "how many | |
| examples do I need?" before you run the experiment. Two | |
| estimates: a formula and an empirical ladder with two curve | |
| families (exponential saturation and Hill cooperativity). | |
| Both are pure numpy, train in seconds, and ship as code — there are | |
| no weights to download. | |
| ## The claim in one sentence | |
| A model claim that can be summarised in a single accuracy number is | |
| an incomplete claim; a complete claim names the splits, reports | |
| calibration on each, and states how many examples the target | |
| required in the first place. | |
| ## What the diagnostic reports | |
| Every `LearningReport` contains: | |
| - `parameters` — model size | |
| - `training examples` — dataset size | |
| - `training accuracy` — fit, not evidence of learning | |
| - `temperature` — fitted value, with a `[BOUNDARY]` flag if it landed | |
| on the edge of the search grid | |
| - `output uniformity` — normalized entropy of the model's prediction | |
| histogram; low values mean the model has collapsed to one class | |
| - `mean max-prob` — average top-class probability; near-uniform | |
| values are not calibration | |
| - `verdict` — `[ok]`, `[?]`, `[??]`, or `[???]` by issue count | |
| - `issues` — plain-English list of every problem detected | |
| - `splits` — four rows, each with `n`, `accuracy`, `mean_confidence`, | |
| `ece`, and an OOD marker | |
| The four splits are: | |
| | split | operands | operation | what it tests | | |
| |---|---|---|---| | |
| | interpolation | in training range | same | fit + generalization on distribution | | |
| | near-OOD | one axis extended | same | partial transfer | | |
| | far-OOD | all outside training range | same | generalization to new inputs | | |
| | structural-OOD | in training range | different | transfer to a new operation | | |
| The last split is the load-bearing one and it is almost never | |
| reported. It distinguishes "learned the sum rule" from "learned the | |
| input-output table." | |
| ## What the volume calculator reports | |
| Two independent estimates of training-set size: | |
| **Formula.** `N ≈ k · P / log2(C) · 1 / (1 - target_acc)` with | |
| `k = 0.05` calibrated on the demo task. Fast, no training, useful for | |
| the first pass. | |
| **Empirical ladder.** Train the model at a ladder of dataset sizes, | |
| fit a saturating curve to the resulting accuracy, and invert to | |
| recommend an N. Two curve families: | |
| | family | form | when it fits | | |
| |---|---|---| | |
| | exponential | `acc(N) = acc_max · (1 - exp(-N / N_half))` | smooth saturation | | |
| | Hill | `acc(N) = acc_max · N^h / (N_half^h + N^h)` | step-shaped ladder | | |
| The Hill exponent `h` is a diagnostic: `h > 1` means the transition | |
| is cooperative, and the exponential fit is systematically | |
| overestimating the required N. | |
| ## Install | |
| ```bash | |
| pip install numpy | |
| python learning_report.py | |
| ``` | |
| No downloads. No tokenizer files. No model hub. Runs in under a | |
| minute on CPU. | |
| ## Usage | |
| ### Diagnostic checklist | |
| ```python | |
| from learning_report import ( | |
| MLP, train, fit_temperature, diagnose, print_report, | |
| sample_pairs, encode_pairs, labels_add, labels_mul, | |
| ) | |
| # Train a model (or bring your own that exposes .predict_proba). | |
| model = MLP(in_dim=40, hidden=32, out_dim=40, seed=0) | |
| # Build four splits as lists of raw (a, b) tuples. | |
| interp = sample_pairs(0, 9, 0, 9, 80, seed=200) | |
| near_ood = sample_pairs(0, 9, 10, 14, 60, seed=201) | |
| far_ood = sample_pairs(10, 19, 10, 19, 60, seed=202) | |
| struct = sample_pairs(0, 5, 0, 5, 36, seed=203) | |
| # diagnose() encodes internally; pass raw pairs, not arrays. | |
| report = diagnose( | |
| model, | |
| train_pairs, y_train, | |
| interp, labels_add(interp), | |
| near_ood, labels_add(near_ood), | |
| far_ood, labels_add(far_ood), | |
| struct, labels_mul(struct), | |
| ) | |
| print_report(report) | |
| ``` | |
| ### Volume calculator | |
| ```python | |
| from learning_report import ( | |
| estimate_volume_formula, | |
| estimate_volume_empirical, | |
| make_add_data, | |
| make_add_model, | |
| ) | |
| # Formula only. | |
| n = estimate_volume_formula(n_params=2632, n_classes=40, | |
| target_acc=0.80) # → 124 | |
| # Empirical ladder with exponential fit. | |
| recommended_exp, diag_exp = estimate_volume_empirical( | |
| make_data=make_add_data, | |
| model_factory=make_add_model, | |
| N_grid=(16, 32, 64, 128, 256, 512, 1024), | |
| epochs=500, | |
| seeds=2, | |
| target_acc=0.80, | |
| form="exp", | |
| ) | |
| # Empirical ladder with Hill fit. | |
| recommended_hill, diag_hill = estimate_volume_empirical( | |
| make_data=make_add_model and make_add_data, | |
| model_factory=make_add_model, | |
| target_acc=0.80, | |
| form="hill", | |
| ) | |
| ``` | |
| ## Results from the demo | |
| The demo trains a 2,632-parameter MLP on `a + b` for `a, b ∈ [0, 9]` | |
| with 60 examples. | |
| ### Self-test | |
| **21/21 checks pass.** The load-bearing ones: | |
| | check | guards against | | |
| |---|---| | |
| | `threshold: ece fires when ece>0.20` | high ECE going unreported | | |
| | `threshold: underconf fires when conf<0.40 and acc>0.5` | temperature over-correction going unreported | | |
| | `curve fit hill beats exp on step data` | the fit family being chosen for the wrong shape | | |
| | `boundary: T=20 flagged` | temperature hitting the grid edge | | |
| ### Diagnostic report | |
| ``` | |
| parameters : 2632 | |
| training examples : 60 | |
| training accuracy : 100.0% | |
| temperature : 4.45 | |
| output uniformity : 0.658 | |
| mean max-prob : 0.244 | |
| verdict : [???] | |
| issues: | |
| - cannot generalise to OOD inputs: far-OOD accuracy 0.00 | |
| - cannot transfer to a different operation: structural-OOD accuracy 0.11 | |
| - poor calibration on interpolation: ECE=0.33 (threshold 0.20) | |
| - underconfident on interpolation (acc=0.60, conf=0.38); temperature may be over-corrected | |
| split n acc mean_conf ECE | |
| -------------------------------------------------------- | |
| interpolation 80 0.60 0.38 0.33 | |
| near-OOD 50 0.00 0.17 0.17 (OOD) | |
| far-OOD 60 0.00 0.07 0.07 (OOD) | |
| structural-OOD 36 0.11 0.39 0.28 (OOD) | |
| ``` | |
| A model that fits its training data perfectly (`100.0%` training | |
| accuracy) and generalizes to `60%` on-distribution, `0%` far-OOD, | |
| and `11%` structural-OOD. The report names four issues. A model card | |
| that reported only the training accuracy would be technically true | |
| and completely misleading. | |
| ### Volume calculator | |
| **Formula:** | |
| | target acc | estimated N | | |
| |---|---| | |
| | 0.30 | 36 | | |
| | 0.50 | 50 | | |
| | 0.70 | 83 | | |
| | 0.80 | 124 | | |
| | 0.90 | 248 | | |
| **Empirical ladder:** | |
| | N | val_acc | | |
| |---|---| | |
| | 16 | 0.165 | | |
| | 32 | 0.320 | | |
| | 64 | 0.640 | | |
| | 128 | 1.000 | | |
| | 256 | 1.000 | | |
| | 512 | 1.000 | | |
| | 1024 | 1.000 | | |
| **Two fits:** | |
| | family | ceiling | N_half | h | MSE | recommended N for 0.80 | | |
| |---|---|---|---|---|---| | |
| | exp | 1.000 | 67.7 | — | 0.0042 | 110 | | |
| | **hill** | **1.000** | **40.8** | **2.04** | **0.0027** | **81** | | |
| The Hill fit has lower MSE (0.0027 vs 0.0042), and the fitted `h = 2.04` | |
| is the signature of a step, not a smooth curve. The ladder jumps from | |
| 0.640 at N=64 to 1.000 at N=128; the exponential fit smooths over | |
| that and recommends 110, while the Hill fit recommends 81. | |
| **Verification:** | |
| | N | held-out accuracy | target met? | | |
| |---|---|---| | |
| | 81 (hill) | 0.810 | yes, exactly | | |
| | 110 (exp) | 1.000 | yes, overshot | | |
| | 124 (formula) | 1.000 | yes, overshot | | |
| All three recommendations land above the target. The Hill | |
| recommendation is the most economical — it hits the target without | |
| overshooting — because the fit correctly identified the ladder as a | |
| step rather than a smooth curve. | |
| ## The bug class | |
| `"Our model achieves X% accuracy."` | |
| The single number hides: | |
| - which split the X% was measured on | |
| - whether any OOD split was tested | |
| - whether confidence is calibrated | |
| - whether the temperature fit degenerated | |
| - whether the output distribution collapsed | |
| - how many examples the target required in the first place | |
| Every one of these can be wrong while X% is true. | |
| ## The six defenses | |
| | level | mechanism | action | | |
| |---|---|---| | |
| | 1. split | four evaluation axes | OOD is measured, not assumed | | |
| | 2. calibrate | mean confidence + ECE on each split | confidence claims are checked | | |
| | 3. canary | output uniformity | degenerate models are flagged | | |
| | 4. bound | temperature boundary check | degenerate calibration is flagged | | |
| | 5. threshold | ECE and underconfidence thresholds | calibration issues become issues | | |
| | 6. plan | volume calculator (formula + empirical) | the target N is known before training | | |
| Any evaluation can choose its level. The choice is explicit in the | |
| report. | |
| ## When to use it | |
| - **Before publishing any accuracy number.** Run the diagnostic and | |
| quote the full report, not the top line. | |
| - **Before running an expensive training job.** The volume | |
| calculator tells you how many examples the target requires. If | |
| you do not have them, you have a data problem, not a model | |
| problem. | |
| - **When comparing two models.** Same splits, same calibration | |
| metric, same OOD axes. A model that wins on interpolation and | |
| loses on structural-OOD is a different model. | |
| - **When a model's accuracy looks too good.** The diagnostic | |
| includes a formula check: if `N * log2(C) / P` is under 0.05, the | |
| model is over-parameterized for the data. | |
| ## When *not* to use it | |
| - **When the task has no meaningful OOD axis.** Some tasks are | |
| bounded by construction (a fixed vocabulary, a closed label set). | |
| The structural-OOD split requires two different operations on the | |
| same inputs; if your task has only one, skip that split. | |
| - **When the model is already validated against a public benchmark.** | |
| Public benchmarks handle splitting. Use the diagnostic for the | |
| cases the benchmark does not cover. | |
| - **As a substitute for domain expertise.** The report flags | |
| problems; it does not tell you which ones matter for your | |
| application. | |
| ## Honest limitations | |
| - **The formula constant `k = 0.05` is calibrated on one demo.** | |
| It is a rule of thumb, not a law. Re-calibrate on your own task | |
| before trusting the number. | |
| - **The empirical ladder is expensive.** Seven training runs at | |
| two seeds each. If training is expensive, use the formula and | |
| validate with a single held-out training run. | |
| - **The Hill fit is chosen by MSE, not by hypothesis.** A Hill fit | |
| beating an exp fit is evidence that the ladder has a step, not | |
| proof. The curve families are approximations. | |
| - **The four splits do not cover every OOD axis.** Adversarial | |
| inputs, distribution shift in the feature distribution, label | |
| noise, and length-based generalization are separate axes not | |
| measured here. | |
| - **The diagnostic does not know about your loss function.** ECE is | |
| meaningful for classification. For regression, use a proper | |
| scoring rule and report residuals. | |
| - **`temperature_hit_boundary` is a heuristic.** A temperature of | |
| 19.9 on a grid with upper bound 20.0 is technically inside the | |
| grid but suspiciously close. The check uses `>= upper - 1e-6`, | |
| which catches only exact boundary hits. | |
| - **The volume calculator assumes the curve shape is fixed.** Real | |
| learning curves can have plateaus, jumps, and non-monotonic | |
| regions. The two fitted families cover two shapes; other shapes | |
| are silently mis-modelled. | |
| ## What this tool demonstrates | |
| Four findings from the demo, each visible in the report: | |
| **1. Training accuracy is not evidence.** The model reaches 100% | |
| on training data and 0% on far-OOD. | |
| **2. Calibration is not accuracy.** Mean confidence is 0.38 on | |
| interpolation where accuracy is 0.60 — the model is | |
| underconfident, not overconfident, and the temperature-scaled ECE | |
| is 0.33. | |
| **3. A single number is not a report.** Four issues, all invisible | |
| in the training accuracy figure. | |
| **4. The fit family matters.** Exponential and Hill fits produce | |
| recommendations of 110 and 81 examples for the same target. The | |
| Hill fit is more economical because it correctly identifies the | |
| ladder as a step. | |
| ## Version history | |
| | version | change | | |
| |---|---| | |
| | 0.1.0 | `LearningReport`, `diagnose()`, `print_report()`, four splits | | |
| | 0.2.0 | `estimate_volume_formula()` | | |
| | 0.3.0 | `estimate_volume_empirical()`, exponential saturation fit | | |
| | 0.4.0 | temperature-boundary check, output-uniformity canary | | |
| | 0.5.0 | ECE and underconfidence thresholds in `diagnose()` | | |
| | 0.6.0 | Hill cooperativity fit; both curve families supported | | |
| ## Reference | |
| Part of a series of small tools built in one session: | |
| | tool | reads | answers | | |
| |---|---|---| | |
| | `hv-manifold` | a corpus | the geometry of style space | | |
| | `hv-reader` | one text | how it reads | | |
| | `anomaly-or-bug` | a number and a matrix | is this a bug or a discovery? | | |
| | `frontier-check` | one claim | where does it sit relative to the frontier? | | |
| | `numerical-provenance` | a numeric pipeline | can I trust this number? | | |
| | `token-provenance` | an LLM pipeline | can I trust this token count? | | |
| | `provenance-reconstruct` | a bare number | what produced this, and how sure are you? | | |
| | `math-ood` | a small model | did it learn a rule or memorize a table? | | |
| | **`learning-report`** | **a trained model** | **can I make an honest claim about what it learned?** | | |
| The design principle — that a model claim should be as structured | |
| as the provenance of a single number — came out of a session in | |
| which the same diagnostic (four splits, calibration, OOD) was | |
| applied by hand to five different models, and then written down. | |
| ## License | |
| Apache-2.0 |