hv-locality
Locality diagnostic for binary feature maps. Feed it a function. Get back a verdict on whether the function preserves locality — or destroys it beyond the point where any classifier can recover signal.
The claim in one sentence
No model on Hugging Face answers "will these features work for my classifier?" Every model either produces features or consumes them. This one diagnoses them.
What it measures
Given a callable f: {0,1}^n -> {0,1}^m:
| quantity | meaning |
|---|---|
| avalanche ε | L(1) — correlation of output bits under single-bit flips |
| locality curve L(k) | ` |
| fitted ε | least-squares fit of L(k) ≈ ε^k on the decreasing region |
| saturation k* | smallest k where L(k) < plateau_threshold |
| plateau L_∞ | residual correlation at large k (noise + structure) |
| verdict | usable/marginal/unusable for a task with known k_intra |
The theorem behind it
From XuanJi-ISA issue #113 (the incompatibility theorem):
For any ε-avalanche function
fand any linear classifier trained onf(x)with N samples, expected accuracy is bounded by1/2 + ε^(d_intra) + C·sqrt(m/N).
Saturation: with ε small, L(k) falls to the usable-signal threshold
at k* ≈ log₂(m). Beyond that, inputs differing by more than k* bits
map to statistically uncorrelated features. No amount of training
fixes that.
The MNIST confirmation: a cryptographic hash (MEP sponge, ε ≈ 0.001, m = 512) produces 512-bit features that stay at 9.79% accuracy across every output dimension from 16 to 4096 — statistically identical to random bits (9.91%) and chance (10%). Raw pixels reach 86.52%.
Install
pip install numpy
## Usage
### Diagnose a function
```python
from hv_locality import HVLocality
m = HVLocality()
def my_encoder(x):
# x is a uint8 array of {0,1} of length n_bits_in
# return a uint8 array of {0,1} of length n_bits_out
...
rpt = m.diagnose(my_encoder, n_bits_in=64, n_bits_out=512)
print(rpt.avalanche_epsilon) # e.g. 0.035
print(rpt.saturation_k) # e.g. 1
print(rpt.plateau_value) # e.g. 0.036
Verdict for a specific task
v = m.verdict(rpt, k_intra=30) # typical within-class Hamming distance
print(v["tier"]) # "unusable"
print(v["recommendation"]) # "features carry no usable class signal..."
Full report
print(m.render(rpt, k_intra=30))
Profile a task's intra-class distance
from hv_locality import task_profile
profile = task_profile(X, y, n_bits=64)
print(profile["k_intra"]["median"]) # typical within-class Hamming distance
print(profile["k_inter"]["median"]) # typical between-class Hamming distance
CLI
# run demos
python hv_locality.py
# diagnose a built-in function
python hv_locality.py --function random_hash --n-bits-in 64 --n-bits-out 512
# with a task verdict
python hv_locality.py --function weak_avalanche --k-intra 30
Benchmarks
Three functions, same I/O dimensions (64 → 512 bits)
| function | ε | L(1) | L(10) | k* | L_∞ |
|---|---|---|---|---|---|
| identity | 0.996 | 0.996 | 0.961 | 51 | 0.828 |
| weak_avalanche (10% mixing) | 0.828 | 0.827 | 0.103 | 11 | 0.037 |
| random_hash (splitmix) | 0.035 | 0.035 | 0.039 | 1 | 0.036 |
Reading the table:
- Identity preserves locality perfectly.
L(k)stays above 0.82 for every k through 50. A classifier always works. - Weak avalanche mixes ~10% of the input into each output bit.
L(k)decays geometrically withε ≈ 0.79and crosses the useful-signal threshold (0.10) at k = 11. - Random hash destroys locality in a single bit flip.
L(1)is already at the plateau (0.035). Same-class inputs at any distance above 1 map to uncorrelated features.
Verdict matrix
k_intra |
identity | weak_avalanche | random_hash |
|---|---|---|---|
| 1 | usable | usable | unusable |
| 5 | usable | usable | unusable |
| 10 | usable | marginal | unusable |
| 20 | usable | unusable | unusable |
| 50 | usable | unusable | unusable |
Random hashes and cryptographic primitives are unusable for every
realistic vision task (MNIST k_intra ≈ 30, CIFAR-10 ≈ 200).
Real task profile
| dataset | k_intra (median) |
k_inter (median) |
|---|---|---|
| synthetic 10-class 64-dim | 14 | 34 |
| MNIST-like 8×8 bits | 8 | 16 |
| MNIST full (784 bits) | ~30 | ~100 |
| CIFAR-10 (3072 bits) | ~200 | ~1000 |
When to use it
- Before training: check whether a new feature extractor preserves
locality. If
L(k_intra)is below the useful-signal threshold, stop — the features are information-theoretically incapable of separating the classes. - During debugging: when a hidden layer underperforms a downstream task, measure its locality curve. Low saturation k* explains the failure directly.
- When evaluating cryptographic primitives as features: the theorem predicts universal failure. This is the empirical confirmation tool.
- When comparing feature extractors: rank them by saturation k*. Higher is better for classification.
When not to use it
- Continuous inputs. The math is for bit-Hamming metrics. Continuous
inputs must be binarized first (see
task_profile). - Nonlinear classifiers on the features. The theorem bounds linear classifiers. Nonlinear classifiers may in principle invert the map, but for avalanche functions this requires exponential time.
- Feature maps with very large m. If m > 2^30, the saturation threshold exceeds realistic within-class distances and the theorem's bound is vacuous. Not a practical concern at typical m.
Honest limitations
- The
plateau_thresholdis functional, not statistical. It is a chosen level (0.10) below whichL(k)is too weak for a linear classifier. The statistical noise floor is much lower (1/sqrt(N·m) ≈ 0.002here) and is reported separately. - The theorem bounds linear classifiers. It's an upper bound, not an exact prediction. Nonlinear classifiers can do better on non-avalanche functions.
- The fitted ε assumes the chain lemma applies. For functions that aren't strictly ε-avalanche, the fit may over- or under-estimate ε. The fit is restricted to the decreasing region, not the plateau.
- Task profiles assume bit-quantization is meaningful. For dense
continuous features, the median-threshold binarization is a lossy
projection. Different quantizations give different
k_intra. - The k* analytic bound is
log₂(m)only when ε = 0. For ε > 0 the effective threshold islog(N/m) / (2·log(1/ε)), which depends on the training size N.
Why this is a new category
Every model on Hugging Face produces or consumes features. None
diagnoses them. hv-locality packages a proven theorem as a usable
tool. It answers a question every ML practitioner asks — "why don't my
features work?" — that nobody has a first-class tool for.
It extends the reader-model family: hv-tail classifies the user,
hv-mode shapes the answer, hv-split bundles readings, hv-ttu
predicts comprehension cost. hv-locality classifies whether a
feature map is usable. Same philosophical move — the model
classifies something nobody else classifies.
Reference
From XuanJi-ISA exploratory track, issue #113 ("Cryptographic Primitives Are Provably Unsuitable as Feature Maps"). The chain lemma proof, the saturation threshold, and the empirical MNIST confirmation are all in that issue.
License
Apache-2.0