holo-invariance

Measure the invariance group of any model.

A mode is characterized by what it can ignore, not by what it can do. This tool measures the invariance group directly: the set of input transformations under which a model's operation is unchanged.

What it measures

For each transformation applied to the input, the tool reports three quantities:

Quantity Meaning
mean_similarity Continuous. Average closeness of outputs.
invariance Fraction of trials with similarity above a threshold. Defines group membership.
decision_change Fraction of trials where the discrete output changed. The "accuracy drop" analog.

The three can move independently. A model can be invariant in operation but not in decisions (same path, different argmax). Or invariant in decisions but not operation (different path, same answer). The signature of a model is its invariance vector across the transformation set.

Why this matters

Standard robustness tests ask does accuracy drop? Invariance asks does the model's internal operation change? Two models with identical robustness can have entirely different invariance groups. The invariance group is what tells you what the model treats as the same. That is the mode, not the accuracy.

The framework comes from the paper series on mode-interface mismatch. A mode is characterized by its invariance group — the class of transformations under which its primary operation is unaffected. The tool makes that definition operational.

Installation

pip install numpy matplotlib

No other dependencies. Single file, approximately 500 lines.

Usage

CLI

python holo_invariance.py
python holo_invariance.py --output results/ --threshold 0.85

Python API

import numpy as np
from holo_invariance import InvarianceProbe, TEXT_TRANSFORMS

def my_model(text):
    # Return a vector, scalar, or string.
    scores = np.array([...])
    return scores / scores.sum()

probe = InvarianceProbe(
    my_model,
    TEXT_TRANSFORMS,
    sim_threshold=0.90,
    name="my_model",
)

report = probe.run(my_texts, n_repeats=3)
report.print()
report.save_json("report.json")
report.plot("signature.png")

For LLMs

Wrap them with a feature extractor. Raw text comparison is too strict; compare embeddings or token distributions instead.

from holo_invariance import InvarianceProbe, TEXT_TRANSFORMS, wrap_llm

def my_llm_call(text):
    # Returns raw model output.
    return llm.generate(text)

def my_embedding(output):
    # Returns a vector representation of the output.
    return embed(output)

model = wrap_llm(my_llm_call, feature_fn=my_embedding)

probe = InvarianceProbe(model, TEXT_TRANSFORMS)
report = probe.run(my_inputs, n_repeats=3)
report.print()

Multi-model comparison

from holo_invariance import compare_models

reports = {}
for name, model in my_models.items():
    probe = InvarianceProbe(model, TEXT_TRANSFORMS, name=name)
    reports[name] = probe.run(my_texts)

compare_models(reports)

What the demo shows

The included demo runs four deliberately different text models through the same transformation set:

Model Description
ExactMatch Matches literal keywords only. Not synonym-invariant.
Concept Broad synonym sets. Synonym-invariant by construction.
LengthOnly Uses text length as its only feature.
Random Control. Ignores input entirely.

And one numeric model: f(x) = sum(x^2).

Results

Text models (10 transformations, 15 inputs, 3 repeats each):

Transformation ExactMatch Concept LengthOnly Random
case_flip 1.00 1.00 1.00 0.04
case_flip_letters 1.00 1.00 1.00 0.27
whitespace 1.00 1.00 1.00 0.22
stopword_drop 1.00 1.00 1.00 0.33
punct_strip 1.00 1.00 1.00 0.13
punct_add 1.00 1.00 1.00 0.20
synonym_swap 0.78 0.93 1.00 0.31
word_shuffle 1.00 1.00 1.00 0.27
char_typo_5pct 0.87 0.82 1.00 0.47
length_double 1.00 1.00 1.00 0.20

Centered signature similarity:

ExactMatch Concept LengthOnly Random
ExactMatch 1.00 0.96 0.92 -0.98
Concept 0.96 1.00 0.96 -0.99
LengthOnly 0.92 0.96 1.00 -0.98
Random -0.98 -0.99 -0.98 1.00

The two classifiers (ExactMatch, Concept) cluster together. LengthOnly sits nearby. Random is orthogonal to everything, including itself in the off-diagonal sense. The tool separates the models that differ and groups the models that are similar.

Numeric model f(x) = sum(x^2):

Transformation Invariance Δdecision In group
permute 1.00 0.00 YES
sign_flip 1.00 0.00 YES
zero_pad 1.00 0.00 YES
noise_0.1 1.00 0.48 YES
shift_+0.5 0.30 0.90 no
scale_1.5 0.00 1.00 no

The function is invariant to permutation, sign flip, and zero padding. It is not invariant to additive shift or scale. The noise_0.1 row is the interesting one: the operation is stable (similarity 0.97), but the discrete readout flips almost half the time. The two are separable.

Built-in transformations

Text (10)

case_flip, case_flip_letters, whitespace, stopword_drop, punct_strip, punct_add, synonym_swap, word_shuffle, char_typo_5pct, length_double

Numeric (6)

permute, scale_1.5, shift_+0.5, sign_flip, noise_0.1, zero_pad

Custom transformations are simple functions:

def my_transform(x, rng):
    # Returns a transformed input.
    return modified_x

MY_TRANSFORMS = {"my_transform": my_transform}
probe = InvarianceProbe(model, MY_TRANSFORMS)

Output format

Reports are JSON-serializable:

{
  "model_name": "SumOfSquares",
  "sim_threshold": 0.9,
  "signature": [1.0, 0.0, 0.3, 1.0, 1.0, 1.0],
  "transform_names": ["permute", "scale_1.5", "shift_+0.5",
                      "sign_flip", "noise_0.1", "zero_pad"],
  "results": {
    "permute": {
      "name": "permute",
      "n_trials": 60,
      "n_effective": 60,
      "mean_similarity": 1.0,
      "invariance": 1.0,
      "decision_change": 0.0,
      "in_group": true
    }
  }
}

Design notes

Effective trial counting

Each transformation reports how often it actually changed the input (n_effective). A transformation that did nothing on the test set reports in_group: n/a and is excluded from the signature. This prevents no-op transformations from appearing invariant by default.

Centered signature similarity

Cosine similarity on uncentered invariance vectors is dominated by the shared ones. The tool centers signatures before comparison, so the difference lives in the dimensions where models actually differ.

Relative-tolerance scalar readout

Scalars are compared with a relative tolerance bin, not exact match. Small perturbations of a scalar are not decision changes.

Limitations

  1. Readout collapses mechanism. A model can be mechanism-sensitive but readout-invariant if the readout collapses the information. The tool measures output-level invariance, not mechanism-level invariance.

  2. Transformation set determines the group. The invariance group is measured relative to the transformations tested. Transformations not in the set are not measured.

  3. Cosine misses ratio changes. For normalized output vectors, cosine is insensitive to changes in the ratio of components. Use relative_l1 as the comparator when ratios matter.

  4. Single-threshold. The threshold 0.90 defines group membership. Different thresholds give different groups. The tool reports mean_similarity for users who want a continuous measure.

Citation

@misc{holo-invariance2026,
  title  = {holo-invariance: Measuring the invariance group of any model},
  author = {zeechimp},
  year   = {2026},
  note   = {Tool for measuring model invariances.
            Companion to the mode-interface mismatch paper series.}
}

References

  • Plate, T. A. "Holographic Reduced Representations." IEEE Transactions on Neural Networks 6:3 (1995), 623–641.
  • Kanerva, P. "Hyperdimensional Computing." Cognitive Computation 1:2 (2009), 139–159.
  • Paper series on mode-interface mismatch (2026).

License

Apache 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support