license: mit pipeline_tag: feature-extraction
hv-failure-catalogue
A catalogue of 27 curated failure modes across 11 model families from the XuanJi corpus. Given a new input or configuration, returns the top-3 closest known failures with diagnosed causes and remediation.
Keywords: failure-catalogue, negative-knowledge, retrieval, diagnosis, numpy-only, no-training, model-forensics.
NumPy only. No training. Runtime < 0.1 s.
Size: ~22 KB source, no weights. Dependencies: NumPy only.
The primitive
Every shipped model has specific failure modes. Each is encoded as a signature vector (16 features): length, alphabet size, entropy, skew, autocorrelation, variance, mean, zero fraction, repeat fraction, max run length, specialty density, plus 5 config-specific slots.
Given a new input:
- Extract its signature.
- Compute Euclidean distance to every catalogue entry.
- Apply family hint weighting if a family is specified.
- Return top-k with distances, diagnoses, and remediations.
Headline numbers
Catalogue composition:
| family | entries |
|---|---|
| acausal-memory | 3 |
| chemical-bubble | 3 |
| chemical-life | 3 |
| csi-detector | 3 |
| drift-rl | 3 |
| chemical-classifier | 2 |
| chemical-compute | 2 |
| metacognition | 2 |
| paleo-llm | 2 |
| stigmergic-router | 2 |
| writhe-stories | 2 |
| total | 27 |
Ground-truth matching on 15 labeled test cases:
| metric | value |
|---|---|
| Top-1 accuracy | 1.000 |
| Top-3 accuracy | 1.000 |
16/16 consistency checks pass.
What the matching actually does
Ground-truth matching is perfect, but the test cases are the same ones used to construct the catalogue β the test is somewhat circular. The more informative experiment is the "unseen input" section:
| input | top-1 match | correct? |
|---|---|---|
| long English text | csi_collapse_on_skew |
β |
| sparse binary stream | memory_capacity_exceeded |
β |
| random walk | csi_dc_peak_artifact |
β |
small chemical-life config |
hypercycle_overflow |
β (with family hint) |
Family hint is required. The signature vector does not carry family-discriminative information on its own. With a hint, top-1 accuracy on unseen inputs is high; without a hint, it collapses to chance-level agreement.
How to use
from hv_failure_catalogue import build_catalogue, extract_signature
cat = build_catalogue()
# Match with a family hint
results = cat.match(
data=my_sequence,
config={'pop_cap': 200, 'point_mut_rate': 0.001},
family_hint='csi-detector',
top_k=3,
)
for r in results:
print(f"[{r['distance']:.3f}] {r['name']}")
print(f" {r['diagnosis']}")
print(f" Fix: {r['remediation']}")
Command line:
python hv_failure_catalogue.py # full benchmark
python hv_failure_catalogue.py summary # catalogue stats
python hv_failure_catalogue.py list # full listing
python hv_failure_catalogue.py match # ground-truth test
python hv_failure_catalogue.py match-with --text "hello world" --family csi-detector
The 27 failure modes
csi-detector (3): specialty class absent β NaN; heavy skew collapses csi_comp; random-walk DC peak artifact.
chemical-life (3): hypercycle overflow at large N; collapse above critical N*; zombie regime misclassification.
chemical-bubble (3): linear oscillator at low drive; single-threshold ceiling at 14/16; LLE = 0 from fixed point.
chemical-classifier (2): global saturation produces coexistence; correlation blindness creates degenerate spike.
chemical-compute (2): register overflow beyond n_max; out-of-range target β undeclared species.
acausal-memory (3): capacity exceeded beyond D/(2 log D); full decay collapse; resurrection threshold too high.
drift-rl (3): reward clipping destroys terminal signal; anchoring worse than nothing; EMA on terminal destroys learning.
writhe-stories (2): myth reads tragic; bittersweet reads balanced.
paleo-llm (2): no attention core β chance; depth doesn't help.
stigmergic-router (2): sticky reset; coarse granularity.
metacognition (2): plateau forecast; vanishing gradient forecast.
Intended use
- Debugging a new run. When training fails, extract the signature of the failing run and query the catalogue. You may be reproducing a known failure.
- Design review. Before shipping a new model, check whether its expected configuration matches any known failure mode.
- Cross-domain transfer. If an input matches a failure from a different family, that's evidence the two families share a structural issue (e.g., overflow, capacity, correlation).
- Reference. A catalogue of 27 diagnosed failure modes is useful documentation for the shipped corpus.
Limitations
- Hand-crafted signatures. The 16 features are chosen for interpretability. A learned embedding might match better.
- No family-discriminative signal. Without a family hint, top-1 accuracy drops sharply. The signature vector does not separate families.
- Small catalogue. 27 entries. Real FMEA catalogs have hundreds.
- Ground-truth benchmark is circular. Test cases are the same ones used to construct the trigger signatures.
- Family hint weighting is a hyperparameter. 0.3Γ boost / 2.0Γ penalty was chosen empirically. Other values would shift the trade-off.
- No novel failure detection. If you feed a new failure mode, the catalogue returns the closest known failure β which may be uninformative.
- No explanation generation. Diagnoses are pre-written, not generated from the input.
Comparison to related tools
| tool | scope | deps | programmatic |
|---|---|---|---|
| FMEA | engineering process | manual | no |
| HAZOP | process hazard analysis | manual | no |
| CHESS (bug-finding) | specific to C/OpenMP | C compiler | semi |
| hv-failure-catalogue | ML model failure modes | numpy | yes |
The catalogue's contribution is programmatic: any new input gets matched against a corpus of known failures via a fixed feature signature. No manual classification required.
Files
hv_failure_catalogue.pyβ full source (NumPy only)config.jsonβ catalogue + benchmark resultsexample.pyβ usage demonstrationsREADME.mdβ this card
Citation
If you use this, cite it as hv-failure-catalogue from the zeechimp
HuggingFace organization.
- Downloads last month
- 17