hv-horizon

Boundary map of a signaling corpus. Six horizons, one capacity.

The claim in one sentence

hv-horizon maps the six boundaries of a signaling corpus β€” semantic, spatial, contextual, temporal, individual, combinatorial β€” and reports the range over which the corpus is doing something vs merely making sound.

The unified object

H = (S, C, G, βˆ‚)

S = symbol streams per source
C = context per source (sender, filename, position)
G = transition graph over the alphabet
βˆ‚ = the six-dimensional boundary

hv-sign provides S and G. Sender context comes from filenames. βˆ‚ is the output.

What it produces

Six horizons, each in [0, 1] unless noted:

horizon reads high value means
semantic structure without meaning ritualized, repetitive
spatial tolerance to noise long-range, robust
contextual uncertainty of expectations large shadow, unpredictable
temporal passes required (fold) needs re-decoding
individual sender distinguishability many identifiable sources
combinatorial unvisited bigram fraction sparse transition graph

Six emergent scalars:

scalar formula reads
semantic range spatial Γ— (1 βˆ’ semantic) meters of meaning
fold-adjusted range spatial / max(1, fold) single-pass decoding limit
ghost density = combinatorial fraction of unrealized signals
sender entropy = individual diversity of sources
shadow coherence 1 βˆ’ contextual how predictable
signaling capacity geometric mean of six normalized signals one scalar: 0 to 1

Install

pip install numpy

Actually β€” no dependencies beyond `hv_sign.py` itself, which is pure stdlib.

## Usage

### Demo

```bash
python hv_horizon.py

### Demo

```bash
python hv_horizon.py

Synthesizes a nine-file corpus across three fictional senders, runs the full horizon analysis, saves and reloads the model.

Real corpus

python hv_horizon.py --dir ./fsdd/recordings --limit 100 --sender-field 1 --decay-sample 20
python hv_horizon.py --dir ./esc50/audio --limit 100 --decay-sample 20
python hv_horizon.py --dir ./corpus --json > horizons.json

Python

from hv_horizon import HVHorizon

m = HVHorizon()
report = m.analyze("./fsdd/recordings", max_files=100)

print(report.horizons.semantic)             # structure without meaning
print(report.horizons.spatial)              # noise tolerance
print(report.horizons.temporal)             # passes required
print(report.emergent.signaling_capacity)   # the scalar
print(report.sender_counts)                 # {sender: n_files}

The six horizons

1. Semantic β€” hv-gap

Mean per-file gap from hv_sign's void analysis. Structure without meaning. A repetitive file (all same symbol) has structure = 1, diversity = 0, meaning = 0, gap = 1. A patterned file with variation has structure and meaning both high, gap low.

semantic = mean(gap) over files with β‰₯ 2 events.

2. Spatial β€” hv-decay

Tolerance to noise. For each sampled file, run the decay analysis, read symbolic_snr β€” the lowest SNR at which the score still recovers.

symbolic_snr meaning spatial
None fails at clean 0.0
30 (top of range) fails almost immediately 0.0
15 recovers to 15 dB 0.5
5 recovers to 5 dB 0.833
0 (bottom of range) recovers to 0 dB or lower 1.0

spatial = mean(1 βˆ’ symbolic_snr / max_snr) over the sample. A signal that recovers at 5 dB is more robust than one that only recovers at 30 dB, so low SNR threshold = high spatial horizon.

3. Contextual β€” hv-shadow

Conditional entropy of the next symbol given the previous, divided by the unconditional entropy of the next symbol.

contextual = H(next | prev) / H(next)

High value = the previous symbol doesn't predict the next = large shadow. Low value = strong sequencing structure = small shadow.

4. Temporal β€” hv-fold

Passes required. From the same entropies:

predictability = (H(next) βˆ’ H(next | prev)) / H(next)

fold = 1 / max(predictability, 0.1), capped at 10.

A corpus where the next symbol is fully predictable has fold = 1. A corpus where the next symbol is completely independent has predictability β‰ˆ 0.1, fold β‰ˆ 10.

5. Individual β€” hv-mirror

Sender distinguishability. Between-sender variance divided by total variance, over the 4-D feature space.

individual = Var_between / (Var_between + Var_within)

Sender extraction is by filename field. Config sender_field selects which underscore/dash-separated field holds the sender id.

  • FSDD: 0_jackson_0 β†’ sender_field = 1 β†’ jackson
  • ESC-50: 1-100032-A-0 β†’ sender_field = 1 β†’ 100032 (clip id, not sender β€” see limitations)

individual = 0.0 if the corpus has fewer than two senders.

6. Combinatorial β€” hv-ghost

Unvisited bigram fraction.

visited = |{(a, b) : a β†’ b observed}| max_bigrams = kΒ² (including self-loops) combinatorial = 1 βˆ’ visited / max_bigrams

High value = large empty space in the transition graph. Low value = saturated system.

The emergent scalars

Semantic range

spatial Γ— (1 βˆ’ semantic)

The product of "how robust" and "how meaningful." A signal that is robust but repetitive has low semantic range β€” it survives the distance but says nothing. A signal that is meaningful but fragile also has low semantic range. Maximum when both terms are high.

Fold-adjusted range

spatial / max(1, fold)

A signal that requires two passes to decode effectively has half the range, because at the edge of audibility there is no second pass. This is a physical constraint, not a preference.

Ghost density, sender entropy, shadow coherence

Direct copies of combinatorial, individual, 1 βˆ’ contextual. Kept separate for naming clarity: they are the consequences of the horizons, not the horizons themselves.

Signaling capacity

Geometric mean of six normalized signals:

1 - semantic      (low gap is good)
spatial           (high tolerance is good)
1 - contextual    (low shadow is good)
1 / max(1, fold)  (low fold is good)
individual        (high sender diversity is good)
1 - combinatorial (low ghost is good)

capacity = (product of six) ^ (1/6), bounded to [0, 1].

A designed communication system should score high. Random environmental sound should score low.

Benchmarks

Synthesized demo corpus

Nine files, three senders (beeps, chirps, claps), three files each:

horizon value reading
semantic 0.621 mixed β€” beeps are highly ritualized, claps vary
spatial 0.767 high β€” pure tones survive low SNR
contextual 0.328 moderate β€” some sequencing structure
temporal 1.489 low β€” modest decoding effort
individual 0.614 high β€” three distinct senders
combinatorial 0.680 high β€” sparse transition graph
signaling capacity 0.543 moderate

The demo corpus is honest mid-range. It has senders and events but the per-file behaviour is repetitive, so semantic stays elevated.

Expected on real corpora

corpus prediction reasoning
FSDD (spoken digits) capacity 0.6–0.8 designed communication system
ESC-50 (environmental sound) capacity 0.2–0.4 not signaling
Zebra finch call corpus capacity 0.5–0.7 categorical, semantically organized
Sperm whale codas capacity 0.4–0.6 structured but partially opaque

The prediction is falsifiable. If hv-horizon cleanly separates a communication corpus (FSDD) from a non-communication corpus (ESC-50), the model has recovered the fundamental distinction between signaling and sound.

When to use it

  • Corpus triage. Given a large unlabeled audio collection, rank subsets by signaling capacity.
  • Cross-species comparison. Compare two communication systems on the same six axes.
  • Bioacoustics hypothesis testing. A species said to have individual-specific calls should score high on individual. A species said to have a ritualized display should score high on semantic.
  • Annotation planning. Low combinatorial = the alphabet is saturated, annotation is worthwhile. High combinatorial = many unvisited paths, annotation may not cover the space.
  • Data quality control. A corpus with anomalously high semantic or low spatial may be a processing artifact, not real.

When not to use it

  • For speech recognition. No words, no phonemes.
  • For labeling. The output is a boundary, not a category.
  • For single files. hv-horizon reads corpora. Single files should go to hv-sign.
  • When filenames don't encode sender. Without a sender structure, individual = 0 and the model loses a sixth of its signal.
  • For continuous sound without onsets. Drones and hisses produce zero events and every horizon collapses.

Honest limitations

  • Sender extraction is filename-based. hv-horizon does not infer sender identity from acoustic features. If the sender is not encoded in the filename, the individual horizon is zero. FSDD works because every file is digit_speaker_take. ESC-50 does not β€” the middle field is a clip id, not a speaker.
  • Spatial is a proxy, not meters. The decay horizon measures SNR tolerance, not physical distance. Converting to meters requires an acoustic propagation model and is out of scope for v0.1.
  • The six horizons are not independent. A repetitive corpus scores high on semantic and low on individual because both derive from the same low symbol diversity. The geometric mean partly corrects for this by requiring all six to be favorable, but the individual axes still correlate.
  • Capacity weights are implicit, not fit. The geometric mean treats all six horizons equally. Whether speech should weight semantic more than spatial is an empirical question the model does not answer.
  • Alpha dominates fold for small k. With k = 5, fold is bounded by 1 / (1 βˆ’ log2(5) / log2(5)) = 1 / (1 βˆ’ 1) β€” but the min_predictability = 0.1 floor caps this. On a corpus with k > 2²⁰ the fold would saturate. On small-k corpora the fold rarely exceeds 3.
  • The demo corpus is small. Nine files is not a corpus. The numbers in the benchmark table are illustrative, not measurements.
  • No calibration against labeled communication corpora. The predictions in the "Expected on real corpora" table are extrapolated from the theory, not fit to data. They are stated to be falsified.

Reference

Part of the reader-model series, which has grown beyond reading.

Companion to hv-sign (symbolic audio signature), hv-note (fixed alphabet), hv-core (argument robustness), hv-fold (reading passes), hv-tempo (pace variation).

hv-sign reads a corpus. hv-horizon reads the boundary of a corpus. Same data, one level up.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support