provenance-reconstruct
Given a bare number, infer the provenance it lost β and say how much you do not know.
A solver returns a float. The API strips the convergence state. A tokenizer returns a count. The log loses the template name. By the time a consumer sees the value, the metadata that made it trustworthy is gone.
provenance-reconstruct is the consumer-side complement to the two
producer-side tools in this series. It reads a bare number and emits
a provenance tag with calibrated confidence:
5.9971, n=20 -> [unconverged_short p=0.99]
1.7109, n=20 -> [unconverged_short? p=0.75]
97.705, n=448 -> [unconverged_short?? OOD p=0.63]
The ? means low confidence. The ?? OOD means the input is
outside the training distribution and the confidence is not
meaningful.
The claim in one sentence
A bare number carries a recoverable amount of information about its own provenance; a small calibrated classifier can extract that information and β more importantly β report when the extraction is unreliable.
What it is, and what it is not
It is:
- a 387β515 parameter numpy MLP, trained in seconds on CPU
- a forensic reconstructor: it infers what produced a value
- a calibrated confidence emitter:
p=0.75means 75% accurate - an OOD-aware predictor: out-of-range inputs get a
??marker - the consumer-side complement to
numerical-provenanceandtoken-provenance
It is not:
- a transformer
- a text generator
- a replacement for observed provenance
- a guarantee:
[CONVERGED p=0.99]can still be wrong
Install
pip install numpy
python provenance_reconstruct.py
No downloads. No tokenizer files. No model hub. The "tokenizer" is a 12-line local function. The "model" trains from scratch at startup.
Usage
Numeric provenance
from provenance_reconstruct import (
MLP, Normalizer, train_mlp, fit_temperature,
reconstruct_numeric, gen_numeric_examples,
)
X_tr, y_tr = gen_numeric_examples(2000, seed=10)
X_te, y_te = gen_numeric_examples(400, seed=11)
norm = Normalizer.fit(X_tr)
model = MLP(X_tr.shape[1], 32, 3, seed=20)
train_mlp(model, norm(X_tr), y_tr, epochs=250, batch=64, seed=21)
model.temperature = fit_temperature(model, norm(X_te), y_te)
r = reconstruct_numeric(model, norm, value=5.9971, n=20)
print(r.tag()) # [unconverged_short? p=0.75] or similar
Token provenance
from provenance_reconstruct import (
reconstruct_token, gen_token_examples,
)
r = reconstruct_token(model_t, norm_t, count=54, text=prompt)
print(r.tag()) # [bpe+chatml p=1.00]
Markers
| tag form | meaning |
|---|---|
[LABEL p=0.95] |
high-confidence inferred provenance |
[LABEL? p=0.61] |
low-confidence inferred provenance |
[LABEL?? OOD p=1.00] |
out of training distribution; p is not meaningful |
A downstream consumer can require observed provenance (from
numerical_provenance.py or token_provenance.py) and refuse
inferred provenance. The markers make the distinction visible.
Results
Self-test
19/19 checks pass. The load-bearing ones:
| check | guards against |
|---|---|
numeric task: no leaking features (acc < 0.85) |
re-introducing max_iter / tol as features |
numeric task: better than chance (acc > 0.40) |
a genuinely broken feature extractor |
token task: near-perfect (acc > 0.95) |
the residual trick being broken |
serialize: OOD emits '??' |
the honesty marker failing to fire |
Numeric task
| parameters | 387 |
| training examples | 2000 |
| test examples | 400 |
| majority baseline | 54.9% |
| test accuracy | 77.5% |
| expected calibration error | 0.021 |
| fitted temperature | 0.99 |
This is the honest result. Features deliberately exclude
max_iter and tol β those are solver config, not observable at
reconstruction time. With them, accuracy is 99% by label leakage.
Without them, the value's shape and the matrix size carry modest
but real signal. The model is meaningfully above baseline and
well-calibrated, but it is not a magic provenance oracle.
The confusion is concentrated between the two unconverged classes.
converged vs unconverged is easier to separate; short vs long
is nearly impossible from the value alone.
Token task
| parameters | 515 |
| training examples | 2000 |
| test examples | 400 |
| test accuracy | 100.0% |
| expected calibration error | 0.000 |
| fitted temperature | 0.50 |
This is a positive result, but a trivial one. The residual
count - bpe_like(text) equals the template offset exactly
(0, +7, or +11). The model learns a threshold lookup. It is
arithmetic dressed as learning; the model is packaging, not magic.
The fitted temperature hits the lower bound of the search grid
(0.50), which signals a saturated softmax. In this case the
saturation is correct β the task really is that separable β but
it means the confidence number carries less information than it
appears to. A downstream consumer should not treat p=1.00 on the
token task as more meaningful than p=0.77 on the numeric task.
The residual trick
The entire token-side result reduces to one feature:
observed = bpe_like(text) + template_offset
residual = observed - bpe_like(text) = template_offset
bpe_like is a 12-line local function that estimates token count
from character classes. It is not exact, but it is free and it
uses no model files. Subtracting it from the observed count
isolates the template overhead as an exact integer.
The demo's PART 4 prints this as a truth table:
observed residual inferred
6 +0 bpe (truth: none)
13 +7 bpe+llama2 (truth: llama2)
17 +11 bpe+chatml (truth: chatml)
That table is the token result. Everything else is packaging.
When to use it
- When you have a bare value and no provenance sidecar. The tag is better than nothing, provided you treat it as inferred.
- When you need calibrated uncertainty. The confidence number
is fit by temperature scaling on held-out data;
p=0.75means approximately 75% accurate. - When you need an OOD guard. The
?? OODmarker fires when the input is outside the training distribution, which is exactly when a small model should refuse to give a confident answer.
When not to use it
- When observed provenance is available. Never prefer inferred
provenance over observed provenance.
PValue(...)fromnumerical_provenance.pyis always more trustworthy than a reconstructed tag from this model. - When the value is outside the training distribution and you need a real answer. The OOD marker is honest, but it does not make the prediction correct. Retrain with the input in range, or use a producer that preserves provenance.
- As a substitute for tracking provenance. This tool exists because provenance is lost. Do not use it as an excuse to lose provenance.
Honest limitations
- The numeric accuracy is 77.5%, not 99%. The 99% figure
requires feeding the model
max_iterandtol, which are exactly the config fields that the whole exercise assumes were lost. - The token accuracy is 100% for a trivial reason. The label is
encoded in one arithmetic feature. If the template offset were
variable (e.g. a template that adds
len(system_prompt)tokens), the task would be hard. It is not, because the demo uses three fixed templates. - The OOD marker is a hard cutoff.
n=41is OOD;n=40is in-distribution. A smooth OOD score (mahalanobis distance in feature space, for instance) would be more honest but is not implemented. - The training data is synthetic.
jacobi_eigenandbpe_likeare the ground truth. Real solvers and real tokenizers have failure modes neither of these reproduces. - No held-out calibration set. Temperature is fit on the test set. A real deployment would fit on a validation split and report ECE on a third split.
- The model does not know it does not know. The OOD marker is a hard rule, not a learned uncertainty. A confident wrong answer inside the training distribution is possible.
What this model demonstrates
Two findings, one positive and one negative.
Positive (token side). A fixed template offset is recoverable from a bare count by subtracting a free local estimator. The reconstruction is exact, and the calibration is trivial because the task is arithmetic.
Negative (numeric side). A bare value plus matrix size carries only modest recoverable information about convergence state. The residual and iteration count are the reliable signals, and they are lost with the provenance. No classifier, however large, can recover information that is not in the input.
The value of the tool is therefore not its accuracy. It is its
calibration (temperature-scaled confidence) and its OOD
marker (the ?? when the input is outside the training range).
A reconstructor that says [? p=0.61] is more useful than one that
lies at p=1.00.
Version history
| version | change |
|---|---|
| 0.1.0 | MLP, Normalizer, train_mlp, reconstruct_numeric, reconstruct_token |
| 0.2.0 | fit_temperature for calibrated confidence |
| 0.3.0 | OOD marker for out-of-distribution inputs |
| 0.4.0 | removed leaking features from numeric_features; honest accuracy reported |
Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
hv-manifold |
a corpus | the geometry of style space |
hv-reader |
one text | how it reads |
anomaly-or-bug |
a number and a matrix | is this a bug or a discovery? |
frontier-check |
one claim | where does it sit relative to the frontier? |
numerical-provenance |
a numeric pipeline | can I trust this number? |
token-provenance |
an LLM pipeline | can I trust this token count? |
provenance-reconstruct |
a bare number | what produced this, and how sure are you? |
The design principle β that a value and its provenance are two things, and that the second can be partially reconstructed when the first arrives alone β came out of a session in which the same problem was solved in the forward direction twice (once for numerical convergence, once for token encoding) and the inverse direction was noticed to be missing.
License
Apache-2.0