provenance-reconstruct

Given a bare number, infer the provenance it lost β€” and say how much you do not know.

A solver returns a float. The API strips the convergence state. A tokenizer returns a count. The log loses the template name. By the time a consumer sees the value, the metadata that made it trustworthy is gone.

provenance-reconstruct is the consumer-side complement to the two producer-side tools in this series. It reads a bare number and emits a provenance tag with calibrated confidence:

5.9971, n=20   ->  [unconverged_short  p=0.99]
1.7109, n=20   ->  [unconverged_short? p=0.75]
97.705, n=448  ->  [unconverged_short?? OOD p=0.63]

The ? means low confidence. The ?? OOD means the input is outside the training distribution and the confidence is not meaningful.

The claim in one sentence

A bare number carries a recoverable amount of information about its own provenance; a small calibrated classifier can extract that information and β€” more importantly β€” report when the extraction is unreliable.

What it is, and what it is not

It is:

  • a 387–515 parameter numpy MLP, trained in seconds on CPU
  • a forensic reconstructor: it infers what produced a value
  • a calibrated confidence emitter: p=0.75 means 75% accurate
  • an OOD-aware predictor: out-of-range inputs get a ?? marker
  • the consumer-side complement to numerical-provenance and token-provenance

It is not:

  • a transformer
  • a text generator
  • a replacement for observed provenance
  • a guarantee: [CONVERGED p=0.99] can still be wrong

Install

pip install numpy
python provenance_reconstruct.py

No downloads. No tokenizer files. No model hub. The "tokenizer" is a 12-line local function. The "model" trains from scratch at startup.

Usage

Numeric provenance

from provenance_reconstruct import (
    MLP, Normalizer, train_mlp, fit_temperature,
    reconstruct_numeric, gen_numeric_examples,
)

X_tr, y_tr = gen_numeric_examples(2000, seed=10)
X_te, y_te = gen_numeric_examples(400, seed=11)
norm = Normalizer.fit(X_tr)
model = MLP(X_tr.shape[1], 32, 3, seed=20)
train_mlp(model, norm(X_tr), y_tr, epochs=250, batch=64, seed=21)
model.temperature = fit_temperature(model, norm(X_te), y_te)

r = reconstruct_numeric(model, norm, value=5.9971, n=20)
print(r.tag())   # [unconverged_short? p=0.75]  or similar

Token provenance

from provenance_reconstruct import (
    reconstruct_token, gen_token_examples,
)

r = reconstruct_token(model_t, norm_t, count=54, text=prompt)
print(r.tag())   # [bpe+chatml  p=1.00]

Markers

tag form meaning
[LABEL p=0.95] high-confidence inferred provenance
[LABEL? p=0.61] low-confidence inferred provenance
[LABEL?? OOD p=1.00] out of training distribution; p is not meaningful

A downstream consumer can require observed provenance (from numerical_provenance.py or token_provenance.py) and refuse inferred provenance. The markers make the distinction visible.

Results

Self-test

19/19 checks pass. The load-bearing ones:

check guards against
numeric task: no leaking features (acc < 0.85) re-introducing max_iter / tol as features
numeric task: better than chance (acc > 0.40) a genuinely broken feature extractor
token task: near-perfect (acc > 0.95) the residual trick being broken
serialize: OOD emits '??' the honesty marker failing to fire

Numeric task

parameters 387
training examples 2000
test examples 400
majority baseline 54.9%
test accuracy 77.5%
expected calibration error 0.021
fitted temperature 0.99

This is the honest result. Features deliberately exclude max_iter and tol β€” those are solver config, not observable at reconstruction time. With them, accuracy is 99% by label leakage. Without them, the value's shape and the matrix size carry modest but real signal. The model is meaningfully above baseline and well-calibrated, but it is not a magic provenance oracle.

The confusion is concentrated between the two unconverged classes. converged vs unconverged is easier to separate; short vs long is nearly impossible from the value alone.

Token task

parameters 515
training examples 2000
test examples 400
test accuracy 100.0%
expected calibration error 0.000
fitted temperature 0.50

This is a positive result, but a trivial one. The residual count - bpe_like(text) equals the template offset exactly (0, +7, or +11). The model learns a threshold lookup. It is arithmetic dressed as learning; the model is packaging, not magic.

The fitted temperature hits the lower bound of the search grid (0.50), which signals a saturated softmax. In this case the saturation is correct β€” the task really is that separable β€” but it means the confidence number carries less information than it appears to. A downstream consumer should not treat p=1.00 on the token task as more meaningful than p=0.77 on the numeric task.

The residual trick

The entire token-side result reduces to one feature:

observed = bpe_like(text) + template_offset
residual = observed - bpe_like(text) = template_offset

bpe_like is a 12-line local function that estimates token count from character classes. It is not exact, but it is free and it uses no model files. Subtracting it from the observed count isolates the template overhead as an exact integer.

The demo's PART 4 prints this as a truth table:

observed   residual   inferred
       6         +0   bpe         (truth: none)
      13         +7   bpe+llama2  (truth: llama2)
      17        +11   bpe+chatml  (truth: chatml)

That table is the token result. Everything else is packaging.

When to use it

  • When you have a bare value and no provenance sidecar. The tag is better than nothing, provided you treat it as inferred.
  • When you need calibrated uncertainty. The confidence number is fit by temperature scaling on held-out data; p=0.75 means approximately 75% accurate.
  • When you need an OOD guard. The ?? OOD marker fires when the input is outside the training distribution, which is exactly when a small model should refuse to give a confident answer.

When not to use it

  • When observed provenance is available. Never prefer inferred provenance over observed provenance. PValue(...) from numerical_provenance.py is always more trustworthy than a reconstructed tag from this model.
  • When the value is outside the training distribution and you need a real answer. The OOD marker is honest, but it does not make the prediction correct. Retrain with the input in range, or use a producer that preserves provenance.
  • As a substitute for tracking provenance. This tool exists because provenance is lost. Do not use it as an excuse to lose provenance.

Honest limitations

  • The numeric accuracy is 77.5%, not 99%. The 99% figure requires feeding the model max_iter and tol, which are exactly the config fields that the whole exercise assumes were lost.
  • The token accuracy is 100% for a trivial reason. The label is encoded in one arithmetic feature. If the template offset were variable (e.g. a template that adds len(system_prompt) tokens), the task would be hard. It is not, because the demo uses three fixed templates.
  • The OOD marker is a hard cutoff. n=41 is OOD; n=40 is in-distribution. A smooth OOD score (mahalanobis distance in feature space, for instance) would be more honest but is not implemented.
  • The training data is synthetic. jacobi_eigen and bpe_like are the ground truth. Real solvers and real tokenizers have failure modes neither of these reproduces.
  • No held-out calibration set. Temperature is fit on the test set. A real deployment would fit on a validation split and report ECE on a third split.
  • The model does not know it does not know. The OOD marker is a hard rule, not a learned uncertainty. A confident wrong answer inside the training distribution is possible.

What this model demonstrates

Two findings, one positive and one negative.

Positive (token side). A fixed template offset is recoverable from a bare count by subtracting a free local estimator. The reconstruction is exact, and the calibration is trivial because the task is arithmetic.

Negative (numeric side). A bare value plus matrix size carries only modest recoverable information about convergence state. The residual and iteration count are the reliable signals, and they are lost with the provenance. No classifier, however large, can recover information that is not in the input.

The value of the tool is therefore not its accuracy. It is its calibration (temperature-scaled confidence) and its OOD marker (the ?? when the input is outside the training range). A reconstructor that says [? p=0.61] is more useful than one that lies at p=1.00.

Version history

version change
0.1.0 MLP, Normalizer, train_mlp, reconstruct_numeric, reconstruct_token
0.2.0 fit_temperature for calibrated confidence
0.3.0 OOD marker for out-of-distribution inputs
0.4.0 removed leaking features from numeric_features; honest accuracy reported

Reference

Part of a series of small tools built in one session:

tool reads answers
hv-manifold a corpus the geometry of style space
hv-reader one text how it reads
anomaly-or-bug a number and a matrix is this a bug or a discovery?
frontier-check one claim where does it sit relative to the frontier?
numerical-provenance a numeric pipeline can I trust this number?
token-provenance an LLM pipeline can I trust this token count?
provenance-reconstruct a bare number what produced this, and how sure are you?

The design principle β€” that a value and its provenance are two things, and that the second can be partially reconstructed when the first arrives alone β€” came out of a session in which the same problem was solved in the forward direction twice (once for numerical convergence, once for token encoding) and the inverse direction was noticed to be missing.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support