hv-chemical-classifier

Hypotheses as species. Posterior as steady state. A NumPy-only toolkit that compiles Bayesian inference into a competitive Lotka-Volterra system and reads the classification from the equilibrium.

Size: ~24 KB source, no weights. Runtime: ~2 s for the full benchmark. Dependencies: NumPy only.

The primitive

Each hypothesis is a species. Its growth rate is its posterior weight (prior × product of likelihoods). Its dynamics are competitive Lotka-Volterra. The steady state is the classification.

dH_i/dt = a_i · H_i · (1 − H_i/K) − c · H_i · Σ_{j≠i} H_j

Four mechanisms turn this into a classifier:

mechanism what it does
M1 Winner-take-all per-species cross-inhibition, not global saturation
M2 Pairwise messenger fires only when two hypotheses coexist
M3 Adaptive tempering α = 0.5 + 0.5·sharpness resolves prior vs likelihood
M4 Correlation discount N correlated channels → √N effective evidence

Plus three extensions:

  • Multi-seed bistability — refuses (returns AMBIGUOUS) when genuinely ambiguous
  • True VOI — which channel should I run next?
  • Entropy stopping — stop on max posterior, not on entropy

Headline numbers

Winner-take-all (4 hypotheses, 4 observations):

hypothesis prior growth rate final conc posterior
theropod 0.450 0.0066 0.0000 0.0000
sauropod 0.200 1.0000 0.9980 1.0000
ornithopod 0.200 0.0525 0.0000 0.0000
ceratopsian 0.150 0.0175 0.0000 0.0000

Global vs per-species saturation — the falsifier:

mode max concentration
global dH_i = a_i·H_i·(1 − Σ H/K) 0.820
per-species dH_i = a_i·H_i·(1 − H_i/K) − c·H_i·Σ_{j≠i}H_j 0.998

Global saturation keeps concentrations near uniform. Per-species cross-inhibition produces true winner-take-all.

Pairwise messenger — fires only when coexistence exists:

scenario peak M final M
single hypothesis 0.0000 0.0000
two close hypotheses 0.4751 0.1013
clear winner 0.1545 0.0112

Adaptive tempering — α responds to likelihood sharpness:

specimen α raw log-likelihoods
sharp 0.886 [−7.42, −1.58, −4.53, −5.34]
weak 0.554 [−3.73, −4.42, −2.81, −3.04]

Correlation discount — 5 channels in one group give √5 evidence:

config raw log-odds contribution
independent 5 × log(0.6/0.4) = +2.027
correlated √5 × log(0.6/0.4) = +0.907

Multi-seed bistability — the honest refusal:

scenario verdict winner counts
clear signal sauropod {sauropod: 10}
weak signal AMBIGUOUS {theropod: 2, ornithopod: 8}

13/13 consistency checks pass.

Why per-species cross-inhibition matters

The single most important finding from the source archive:

Global saturation (1 − Σ_j H_j / K) produces coexistence, not winner-take-all. It cannot produce bistability. Any classifier built on global saturation will return near-uniform posterior when the evidence is sparse.

Per-species inhibition − c · H_i · Σ_{j≠i} H_j produces true WTA (winner >0.99) and genuine bistability (multiple stable fixed points). This is a structural property of the dynamics, not a tuning parameter.

The pairwise messenger

A total-activity messenger dM/dt = k · Σ H_i − k_out · M fires on any activity. The pairwise messenger fires only at saddles:

dM/dt = k_pair · Σ_{i<j} H_i · H_j − k_m_out · M

One hypothesis → Σ pairs = 0 → M decays to zero. Two coexist → M rises. Clear winner → M rises during transient, decays to steady state.

The messenger magnitude is a confidence signal orthogonal to the posterior. It fires when the classifier is uncertain between two hypotheses, not when one is winning.

How to use

from hv_chemical_classifier import (
    Hypothesis, EvidenceChannel, ChemicalClassifier,
    adaptive_tempering, true_voi, correlation_discount,
)

# Define hypotheses with priors
hypotheses = [
    Hypothesis('theropod', prior=0.45),
    Hypothesis('sauropod', prior=0.20),
    Hypothesis('ornithopod', prior=0.20),
    Hypothesis('ceratopsian', prior=0.15),
]

# Build the classifier
clf = ChemicalClassifier(hypotheses, saturation_mode='per_species')

# Add evidence channels (likelihoods[hypothesis_idx, outcome])
clf.add_channel(EvidenceChannel(
    'femur_morphology',
    likelihoods=np.array([
        [0.6, 0.3, 0.1],   # theropod
        [0.1, 0.3, 0.6],   # sauropod
        [0.3, 0.4, 0.3],   # ornithopod
        [0.4, 0.4, 0.2],   # ceratopsian
    ]),
    cost=1.0,
))

# Observe evidence
clf.observe('femur_morphology', outcome=2)  # "strong"

# Classify
result = clf.classify()
# result['winner'] = 'sauropod'
# result['posterior'] = [0.0, 1.0, 0.0, 0.0]

# Multi-seed bistability check
ms = clf.classify_multiseed(n_seeds=10)
# ms['ambiguous'] = False, ms['verdict'] = 'sauropod'

# Adaptive tempering: use with likelihood-uncertain specimens
r = clf.classify(apply_tempering=True)

# Correlation discount: use when channels are correlated
r = clf.classify(apply_discount=True)

# True VOI: which channel to run next?
voi = true_voi(clf, current_posterior, ['channel_a', 'channel_b'])
# voi['channel_a'] = {'voi': ..., 'cost': ..., 'voi_per_cost': ...}
Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support