license: mit pipeline_tag: feature-extraction

hv-failure-catalogue

A catalogue of 27 curated failure modes across 11 model families from the XuanJi corpus. Given a new input or configuration, returns the top-3 closest known failures with diagnosed causes and remediation.

Keywords: failure-catalogue, negative-knowledge, retrieval, diagnosis, numpy-only, no-training, model-forensics.

NumPy only. No training. Runtime < 0.1 s.

Size: ~22 KB source, no weights. Dependencies: NumPy only.

The primitive

Every shipped model has specific failure modes. Each is encoded as a signature vector (16 features): length, alphabet size, entropy, skew, autocorrelation, variance, mean, zero fraction, repeat fraction, max run length, specialty density, plus 5 config-specific slots.

Given a new input:

  1. Extract its signature.
  2. Compute Euclidean distance to every catalogue entry.
  3. Apply family hint weighting if a family is specified.
  4. Return top-k with distances, diagnoses, and remediations.

Headline numbers

Catalogue composition:

family entries
acausal-memory 3
chemical-bubble 3
chemical-life 3
csi-detector 3
drift-rl 3
chemical-classifier 2
chemical-compute 2
metacognition 2
paleo-llm 2
stigmergic-router 2
writhe-stories 2
total 27

Ground-truth matching on 15 labeled test cases:

metric value
Top-1 accuracy 1.000
Top-3 accuracy 1.000

16/16 consistency checks pass.

What the matching actually does

Ground-truth matching is perfect, but the test cases are the same ones used to construct the catalogue β€” the test is somewhat circular. The more informative experiment is the "unseen input" section:

input top-1 match correct?
long English text csi_collapse_on_skew βœ“
sparse binary stream memory_capacity_exceeded βœ“
random walk csi_dc_peak_artifact βœ“
small chemical-life config hypercycle_overflow βœ“ (with family hint)

Family hint is required. The signature vector does not carry family-discriminative information on its own. With a hint, top-1 accuracy on unseen inputs is high; without a hint, it collapses to chance-level agreement.

How to use

from hv_failure_catalogue import build_catalogue, extract_signature

cat = build_catalogue()

# Match with a family hint
results = cat.match(
    data=my_sequence,
    config={'pop_cap': 200, 'point_mut_rate': 0.001},
    family_hint='csi-detector',
    top_k=3,
)

for r in results:
    print(f"[{r['distance']:.3f}] {r['name']}")
    print(f"  {r['diagnosis']}")
    print(f"  Fix: {r['remediation']}")

Command line:

python hv_failure_catalogue.py                                    # full benchmark
python hv_failure_catalogue.py summary                            # catalogue stats
python hv_failure_catalogue.py list                               # full listing
python hv_failure_catalogue.py match                              # ground-truth test
python hv_failure_catalogue.py match-with --text "hello world" --family csi-detector

The 27 failure modes

csi-detector (3): specialty class absent β†’ NaN; heavy skew collapses csi_comp; random-walk DC peak artifact.

chemical-life (3): hypercycle overflow at large N; collapse above critical N*; zombie regime misclassification.

chemical-bubble (3): linear oscillator at low drive; single-threshold ceiling at 14/16; LLE = 0 from fixed point.

chemical-classifier (2): global saturation produces coexistence; correlation blindness creates degenerate spike.

chemical-compute (2): register overflow beyond n_max; out-of-range target β†’ undeclared species.

acausal-memory (3): capacity exceeded beyond D/(2 log D); full decay collapse; resurrection threshold too high.

drift-rl (3): reward clipping destroys terminal signal; anchoring worse than nothing; EMA on terminal destroys learning.

writhe-stories (2): myth reads tragic; bittersweet reads balanced.

paleo-llm (2): no attention core β†’ chance; depth doesn't help.

stigmergic-router (2): sticky reset; coarse granularity.

metacognition (2): plateau forecast; vanishing gradient forecast.

Intended use

  • Debugging a new run. When training fails, extract the signature of the failing run and query the catalogue. You may be reproducing a known failure.
  • Design review. Before shipping a new model, check whether its expected configuration matches any known failure mode.
  • Cross-domain transfer. If an input matches a failure from a different family, that's evidence the two families share a structural issue (e.g., overflow, capacity, correlation).
  • Reference. A catalogue of 27 diagnosed failure modes is useful documentation for the shipped corpus.

Limitations

  • Hand-crafted signatures. The 16 features are chosen for interpretability. A learned embedding might match better.
  • No family-discriminative signal. Without a family hint, top-1 accuracy drops sharply. The signature vector does not separate families.
  • Small catalogue. 27 entries. Real FMEA catalogs have hundreds.
  • Ground-truth benchmark is circular. Test cases are the same ones used to construct the trigger signatures.
  • Family hint weighting is a hyperparameter. 0.3Γ— boost / 2.0Γ— penalty was chosen empirically. Other values would shift the trade-off.
  • No novel failure detection. If you feed a new failure mode, the catalogue returns the closest known failure β€” which may be uninformative.
  • No explanation generation. Diagnoses are pre-written, not generated from the input.

Comparison to related tools

tool scope deps programmatic
FMEA engineering process manual no
HAZOP process hazard analysis manual no
CHESS (bug-finding) specific to C/OpenMP C compiler semi
hv-failure-catalogue ML model failure modes numpy yes

The catalogue's contribution is programmatic: any new input gets matched against a corpus of known failures via a fixed feature signature. No manual classification required.

Files

  • hv_failure_catalogue.py β€” full source (NumPy only)
  • config.json β€” catalogue + benchmark results
  • example.py β€” usage demonstrations
  • README.md β€” this card

Citation

If you use this, cite it as hv-failure-catalogue from the zeechimp HuggingFace organization.

Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support