hv-horizon
Boundary map of a signaling corpus. Six horizons, one capacity.
The claim in one sentence
hv-horizon maps the six boundaries of a signaling corpus β
semantic, spatial, contextual, temporal, individual, combinatorial β
and reports the range over which the corpus is doing something vs
merely making sound.
The unified object
H = (S, C, G, β)
S = symbol streams per source
C = context per source (sender, filename, position)
G = transition graph over the alphabet
β = the six-dimensional boundary
hv-sign provides S and G. Sender context comes from filenames.
β is the output.
What it produces
Six horizons, each in [0, 1] unless noted:
| horizon | reads | high value means |
|---|---|---|
| semantic | structure without meaning | ritualized, repetitive |
| spatial | tolerance to noise | long-range, robust |
| contextual | uncertainty of expectations | large shadow, unpredictable |
| temporal | passes required (fold) | needs re-decoding |
| individual | sender distinguishability | many identifiable sources |
| combinatorial | unvisited bigram fraction | sparse transition graph |
Six emergent scalars:
| scalar | formula | reads |
|---|---|---|
| semantic range | spatial Γ (1 β semantic) |
meters of meaning |
| fold-adjusted range | spatial / max(1, fold) |
single-pass decoding limit |
| ghost density | = combinatorial |
fraction of unrealized signals |
| sender entropy | = individual |
diversity of sources |
| shadow coherence | 1 β contextual |
how predictable |
| signaling capacity | geometric mean of six normalized signals | one scalar: 0 to 1 |
Install
pip install numpy
Actually β no dependencies beyond `hv_sign.py` itself, which is pure stdlib.
## Usage
### Demo
```bash
python hv_horizon.py
### Demo
```bash
python hv_horizon.py
Synthesizes a nine-file corpus across three fictional senders, runs the full horizon analysis, saves and reloads the model.
Real corpus
python hv_horizon.py --dir ./fsdd/recordings --limit 100 --sender-field 1 --decay-sample 20
python hv_horizon.py --dir ./esc50/audio --limit 100 --decay-sample 20
python hv_horizon.py --dir ./corpus --json > horizons.json
Python
from hv_horizon import HVHorizon
m = HVHorizon()
report = m.analyze("./fsdd/recordings", max_files=100)
print(report.horizons.semantic) # structure without meaning
print(report.horizons.spatial) # noise tolerance
print(report.horizons.temporal) # passes required
print(report.emergent.signaling_capacity) # the scalar
print(report.sender_counts) # {sender: n_files}
The six horizons
1. Semantic β hv-gap
Mean per-file gap from hv_sign's void analysis. Structure without
meaning. A repetitive file (all same symbol) has structure = 1,
diversity = 0, meaning = 0, gap = 1. A patterned file with variation
has structure and meaning both high, gap low.
semantic = mean(gap) over files with β₯ 2 events.
2. Spatial β hv-decay
Tolerance to noise. For each sampled file, run the decay analysis,
read symbolic_snr β the lowest SNR at which the score still recovers.
symbolic_snr |
meaning | spatial |
|---|---|---|
None |
fails at clean | 0.0 |
| 30 (top of range) | fails almost immediately | 0.0 |
| 15 | recovers to 15 dB | 0.5 |
| 5 | recovers to 5 dB | 0.833 |
| 0 (bottom of range) | recovers to 0 dB or lower | 1.0 |
spatial = mean(1 β symbolic_snr / max_snr) over the sample. A signal
that recovers at 5 dB is more robust than one that only recovers at
30 dB, so low SNR threshold = high spatial horizon.
3. Contextual β hv-shadow
Conditional entropy of the next symbol given the previous, divided by the unconditional entropy of the next symbol.
contextual = H(next | prev) / H(next)
High value = the previous symbol doesn't predict the next = large shadow. Low value = strong sequencing structure = small shadow.
4. Temporal β hv-fold
Passes required. From the same entropies:
predictability = (H(next) β H(next | prev)) / H(next)
fold = 1 / max(predictability, 0.1), capped at 10.
A corpus where the next symbol is fully predictable has fold = 1. A corpus where the next symbol is completely independent has predictability β 0.1, fold β 10.
5. Individual β hv-mirror
Sender distinguishability. Between-sender variance divided by total variance, over the 4-D feature space.
individual = Var_between / (Var_between + Var_within)
Sender extraction is by filename field. Config sender_field selects
which underscore/dash-separated field holds the sender id.
- FSDD:
0_jackson_0βsender_field = 1βjackson - ESC-50:
1-100032-A-0βsender_field = 1β100032(clip id, not sender β see limitations)
individual = 0.0 if the corpus has fewer than two senders.
6. Combinatorial β hv-ghost
Unvisited bigram fraction.
visited = |{(a, b) : a β b observed}|
max_bigrams = kΒ² (including self-loops)
combinatorial = 1 β visited / max_bigrams
High value = large empty space in the transition graph. Low value = saturated system.
The emergent scalars
Semantic range
spatial Γ (1 β semantic)
The product of "how robust" and "how meaningful." A signal that is robust but repetitive has low semantic range β it survives the distance but says nothing. A signal that is meaningful but fragile also has low semantic range. Maximum when both terms are high.
Fold-adjusted range
spatial / max(1, fold)
A signal that requires two passes to decode effectively has half the range, because at the edge of audibility there is no second pass. This is a physical constraint, not a preference.
Ghost density, sender entropy, shadow coherence
Direct copies of combinatorial, individual, 1 β contextual.
Kept separate for naming clarity: they are the consequences of the
horizons, not the horizons themselves.
Signaling capacity
Geometric mean of six normalized signals:
1 - semantic (low gap is good)
spatial (high tolerance is good)
1 - contextual (low shadow is good)
1 / max(1, fold) (low fold is good)
individual (high sender diversity is good)
1 - combinatorial (low ghost is good)
capacity = (product of six) ^ (1/6), bounded to [0, 1].
A designed communication system should score high. Random environmental sound should score low.
Benchmarks
Synthesized demo corpus
Nine files, three senders (beeps, chirps, claps), three files each:
| horizon | value | reading |
|---|---|---|
| semantic | 0.621 | mixed β beeps are highly ritualized, claps vary |
| spatial | 0.767 | high β pure tones survive low SNR |
| contextual | 0.328 | moderate β some sequencing structure |
| temporal | 1.489 | low β modest decoding effort |
| individual | 0.614 | high β three distinct senders |
| combinatorial | 0.680 | high β sparse transition graph |
| signaling capacity | 0.543 | moderate |
The demo corpus is honest mid-range. It has senders and events but the per-file behaviour is repetitive, so semantic stays elevated.
Expected on real corpora
| corpus | prediction | reasoning |
|---|---|---|
| FSDD (spoken digits) | capacity 0.6β0.8 | designed communication system |
| ESC-50 (environmental sound) | capacity 0.2β0.4 | not signaling |
| Zebra finch call corpus | capacity 0.5β0.7 | categorical, semantically organized |
| Sperm whale codas | capacity 0.4β0.6 | structured but partially opaque |
The prediction is falsifiable. If hv-horizon cleanly separates a
communication corpus (FSDD) from a non-communication corpus (ESC-50),
the model has recovered the fundamental distinction between signaling
and sound.
When to use it
- Corpus triage. Given a large unlabeled audio collection, rank subsets by signaling capacity.
- Cross-species comparison. Compare two communication systems on the same six axes.
- Bioacoustics hypothesis testing. A species said to have
individual-specific calls should score high on
individual. A species said to have a ritualized display should score high onsemantic. - Annotation planning. Low
combinatorial= the alphabet is saturated, annotation is worthwhile. Highcombinatorial= many unvisited paths, annotation may not cover the space. - Data quality control. A corpus with anomalously high
semanticor lowspatialmay be a processing artifact, not real.
When not to use it
- For speech recognition. No words, no phonemes.
- For labeling. The output is a boundary, not a category.
- For single files.
hv-horizonreads corpora. Single files should go tohv-sign. - When filenames don't encode sender. Without a sender structure,
individual = 0and the model loses a sixth of its signal. - For continuous sound without onsets. Drones and hisses produce zero events and every horizon collapses.
Honest limitations
- Sender extraction is filename-based.
hv-horizondoes not infer sender identity from acoustic features. If the sender is not encoded in the filename, theindividualhorizon is zero. FSDD works because every file isdigit_speaker_take. ESC-50 does not β the middle field is a clip id, not a speaker. - Spatial is a proxy, not meters. The decay horizon measures SNR tolerance, not physical distance. Converting to meters requires an acoustic propagation model and is out of scope for v0.1.
- The six horizons are not independent. A repetitive corpus scores
high on
semanticand low onindividualbecause both derive from the same low symbol diversity. The geometric mean partly corrects for this by requiring all six to be favorable, but the individual axes still correlate. - Capacity weights are implicit, not fit. The geometric mean treats
all six horizons equally. Whether speech should weight
semanticmore thanspatialis an empirical question the model does not answer. - Alpha dominates fold for small k. With k = 5,
foldis bounded by1 / (1 β log2(5) / log2(5)) = 1 / (1 β 1)β but themin_predictability = 0.1floor caps this. On a corpus with k > 2Β²β° the fold would saturate. On small-k corpora the fold rarely exceeds 3. - The demo corpus is small. Nine files is not a corpus. The numbers in the benchmark table are illustrative, not measurements.
- No calibration against labeled communication corpora. The predictions in the "Expected on real corpora" table are extrapolated from the theory, not fit to data. They are stated to be falsified.
Reference
Part of the reader-model series, which has grown beyond reading.
Companion to hv-sign (symbolic audio signature), hv-note (fixed
alphabet), hv-core (argument robustness), hv-fold (reading passes),
hv-tempo (pace variation).
hv-sign reads a corpus. hv-horizon reads the boundary of a
corpus. Same data, one level up.
License
Apache-2.0