holo-invariance
Measure the invariance group of any model.
A mode is characterized by what it can ignore, not by what it can do. This tool measures the invariance group directly: the set of input transformations under which a model's operation is unchanged.
What it measures
For each transformation applied to the input, the tool reports three quantities:
| Quantity | Meaning |
|---|---|
mean_similarity |
Continuous. Average closeness of outputs. |
invariance |
Fraction of trials with similarity above a threshold. Defines group membership. |
decision_change |
Fraction of trials where the discrete output changed. The "accuracy drop" analog. |
The three can move independently. A model can be invariant in operation but not in decisions (same path, different argmax). Or invariant in decisions but not operation (different path, same answer). The signature of a model is its invariance vector across the transformation set.
Why this matters
Standard robustness tests ask does accuracy drop? Invariance asks does the model's internal operation change? Two models with identical robustness can have entirely different invariance groups. The invariance group is what tells you what the model treats as the same. That is the mode, not the accuracy.
The framework comes from the paper series on mode-interface mismatch. A mode is characterized by its invariance group — the class of transformations under which its primary operation is unaffected. The tool makes that definition operational.
Installation
pip install numpy matplotlib
No other dependencies. Single file, approximately 500 lines.
Usage
CLI
python holo_invariance.py
python holo_invariance.py --output results/ --threshold 0.85
Python API
import numpy as np
from holo_invariance import InvarianceProbe, TEXT_TRANSFORMS
def my_model(text):
# Return a vector, scalar, or string.
scores = np.array([...])
return scores / scores.sum()
probe = InvarianceProbe(
my_model,
TEXT_TRANSFORMS,
sim_threshold=0.90,
name="my_model",
)
report = probe.run(my_texts, n_repeats=3)
report.print()
report.save_json("report.json")
report.plot("signature.png")
For LLMs
Wrap them with a feature extractor. Raw text comparison is too strict; compare embeddings or token distributions instead.
from holo_invariance import InvarianceProbe, TEXT_TRANSFORMS, wrap_llm
def my_llm_call(text):
# Returns raw model output.
return llm.generate(text)
def my_embedding(output):
# Returns a vector representation of the output.
return embed(output)
model = wrap_llm(my_llm_call, feature_fn=my_embedding)
probe = InvarianceProbe(model, TEXT_TRANSFORMS)
report = probe.run(my_inputs, n_repeats=3)
report.print()
Multi-model comparison
from holo_invariance import compare_models
reports = {}
for name, model in my_models.items():
probe = InvarianceProbe(model, TEXT_TRANSFORMS, name=name)
reports[name] = probe.run(my_texts)
compare_models(reports)
What the demo shows
The included demo runs four deliberately different text models through the same transformation set:
| Model | Description |
|---|---|
ExactMatch |
Matches literal keywords only. Not synonym-invariant. |
Concept |
Broad synonym sets. Synonym-invariant by construction. |
LengthOnly |
Uses text length as its only feature. |
Random |
Control. Ignores input entirely. |
And one numeric model: f(x) = sum(x^2).
Results
Text models (10 transformations, 15 inputs, 3 repeats each):
| Transformation | ExactMatch | Concept | LengthOnly | Random |
|---|---|---|---|---|
| case_flip | 1.00 | 1.00 | 1.00 | 0.04 |
| case_flip_letters | 1.00 | 1.00 | 1.00 | 0.27 |
| whitespace | 1.00 | 1.00 | 1.00 | 0.22 |
| stopword_drop | 1.00 | 1.00 | 1.00 | 0.33 |
| punct_strip | 1.00 | 1.00 | 1.00 | 0.13 |
| punct_add | 1.00 | 1.00 | 1.00 | 0.20 |
| synonym_swap | 0.78 | 0.93 | 1.00 | 0.31 |
| word_shuffle | 1.00 | 1.00 | 1.00 | 0.27 |
| char_typo_5pct | 0.87 | 0.82 | 1.00 | 0.47 |
| length_double | 1.00 | 1.00 | 1.00 | 0.20 |
Centered signature similarity:
| ExactMatch | Concept | LengthOnly | Random | |
|---|---|---|---|---|
| ExactMatch | 1.00 | 0.96 | 0.92 | -0.98 |
| Concept | 0.96 | 1.00 | 0.96 | -0.99 |
| LengthOnly | 0.92 | 0.96 | 1.00 | -0.98 |
| Random | -0.98 | -0.99 | -0.98 | 1.00 |
The two classifiers (ExactMatch, Concept) cluster together. LengthOnly sits nearby. Random is orthogonal to everything, including itself in the off-diagonal sense. The tool separates the models that differ and groups the models that are similar.
Numeric model f(x) = sum(x^2):
| Transformation | Invariance | Δdecision | In group |
|---|---|---|---|
| permute | 1.00 | 0.00 | YES |
| sign_flip | 1.00 | 0.00 | YES |
| zero_pad | 1.00 | 0.00 | YES |
| noise_0.1 | 1.00 | 0.48 | YES |
| shift_+0.5 | 0.30 | 0.90 | no |
| scale_1.5 | 0.00 | 1.00 | no |
The function is invariant to permutation, sign flip, and zero padding.
It is not invariant to additive shift or scale. The noise_0.1 row is
the interesting one: the operation is stable (similarity 0.97), but the
discrete readout flips almost half the time. The two are separable.
Built-in transformations
Text (10)
case_flip, case_flip_letters, whitespace, stopword_drop,
punct_strip, punct_add, synonym_swap, word_shuffle,
char_typo_5pct, length_double
Numeric (6)
permute, scale_1.5, shift_+0.5, sign_flip, noise_0.1, zero_pad
Custom transformations are simple functions:
def my_transform(x, rng):
# Returns a transformed input.
return modified_x
MY_TRANSFORMS = {"my_transform": my_transform}
probe = InvarianceProbe(model, MY_TRANSFORMS)
Output format
Reports are JSON-serializable:
{
"model_name": "SumOfSquares",
"sim_threshold": 0.9,
"signature": [1.0, 0.0, 0.3, 1.0, 1.0, 1.0],
"transform_names": ["permute", "scale_1.5", "shift_+0.5",
"sign_flip", "noise_0.1", "zero_pad"],
"results": {
"permute": {
"name": "permute",
"n_trials": 60,
"n_effective": 60,
"mean_similarity": 1.0,
"invariance": 1.0,
"decision_change": 0.0,
"in_group": true
}
}
}
Design notes
Effective trial counting
Each transformation reports how often it actually changed the input
(n_effective). A transformation that did nothing on the test set
reports in_group: n/a and is excluded from the signature. This
prevents no-op transformations from appearing invariant by default.
Centered signature similarity
Cosine similarity on uncentered invariance vectors is dominated by the shared ones. The tool centers signatures before comparison, so the difference lives in the dimensions where models actually differ.
Relative-tolerance scalar readout
Scalars are compared with a relative tolerance bin, not exact match. Small perturbations of a scalar are not decision changes.
Limitations
Readout collapses mechanism. A model can be mechanism-sensitive but readout-invariant if the readout collapses the information. The tool measures output-level invariance, not mechanism-level invariance.
Transformation set determines the group. The invariance group is measured relative to the transformations tested. Transformations not in the set are not measured.
Cosine misses ratio changes. For normalized output vectors, cosine is insensitive to changes in the ratio of components. Use
relative_l1as the comparator when ratios matter.Single-threshold. The threshold
0.90defines group membership. Different thresholds give different groups. The tool reportsmean_similarityfor users who want a continuous measure.
Citation
@misc{holo-invariance2026,
title = {holo-invariance: Measuring the invariance group of any model},
author = {zeechimp},
year = {2026},
note = {Tool for measuring model invariances.
Companion to the mode-interface mismatch paper series.}
}
References
- Plate, T. A. "Holographic Reduced Representations." IEEE Transactions on Neural Networks 6:3 (1995), 623–641.
- Kanerva, P. "Hyperdimensional Computing." Cognitive Computation 1:2 (2009), 139–159.
- Paper series on mode-interface mismatch (2026).
License
Apache 2.0