hv-contour
Contour-based symbolic encoding of frequency-modulated animal vocalizations.
Instead of decoding audio with a large model, reduce each event to a contour shape, encode the sequence as a single hypervector, and decode by similarity.
The claim in one sentence
Dolphins, whales, birds, and tonal languages all carry categorical
information in frequency modulation over time. hv-contour extracts the
contour, encodes it as a hypervector, and compares by similarity β no
transformer, no 400M parameters.
What it produces
For a WAV file:
- contours β the sequence of contour tokens, each with start/end time, start/end Hz, and inflection count
- hypervector β a single binary vector (50,000 dim by default) encoding the entire sequence
For a corpus:
- pairwise similarity β Hamming similarity between every pair of encoded hypervectors (0.5 = orthogonal)
- subsequence query β given a contour pattern, rank corpus entries by similarity of the best matching window
Install
pip install numpy
Actually β no dependencies. Pure stdlib.
## Usage
### Demo
```bash
python hv_contour.py
### Demo
```bash
python hv_contour.py
Encodes a sample corpus of seven synthetic phrases, prints the pairwise similarity matrix, runs two queries, and does an audio round-trip.
Single file
python hv_contour.py --file whistle.wav
Python
from hv_contour import HVContour, extract_contour, _read_wav
m = HVContour(dim=50000, seed=42)
contours, hv = m.encode_file("whistle.wav")
print([c.token for c in contours]) # ['R', 'W', 'F']
print(hv.similarity(other_hv)) # 0.0 - 1.0
# Query a corpus of contour sequences
corpus = {
"rising_call": ["R", "R", "R"],
"falling_call": ["F", "F", "F"],
"contact_arc": ["R", "W", "F"],
}
results = m.query_subsequence(corpus, ["R", "R", "R"])
# [("rising_call", 1.0), ("contact_arc", 0.56), ("falling_call", 0.63)]
The contour alphabet
| token | shape | tonal analogue |
|---|---|---|
R |
rising, no inflection | Mandarin 2nd tone |
F |
falling, no inflection | Mandarin 4th tone |
W |
rise-then-fall | Mandarin 3rd tone |
U |
fall-then-rise | Thai rising |
S |
steady / flat | Mandarin 1st tone |
M |
multi-inflection | complex contour |
Six tokens. Fixed. Unnamed in the sense that they do not carry semantic
meaning β R is "rising", not "alarm call".
The encoding
Each contour token gets a random binary base hypervector. Each position gets a random binary position vector. The encoding of a contour sequence is the majority vote over two families of terms:
unigram_t = token_t XOR position_i
bigram_t = token_i XOR token_{i+1} XOR position_i
H = majority_vote({unigram_t} βͺ {bigram_t} βͺ {continuation})
Binding is XOR. Superposition is majority vote. The result is a single fixed-length binary hypervector that encodes the entire sequence.
Two sequences with the same contour tokens in the same order produce nearly identical hypervectors. Two sequences with different contours produce less similar ones. Orthogonality baseline is 0.5.
Why this works
This is hyperdimensional computing (also called vector-symbolic architecture). The specific choice of encoding β token-bound-to-position plus bigram-bound-to-position, superposed by majority vote β is the classic Kanerva / Plate scheme.
The choice to apply it to contour rather than to raw audio is the modelling claim: contour is the signal. Frequency-modulated vocalizations are perceptually categorized by their shape, not their absolute frequency. The six-token alphabet reflects the canonical contour classes documented in the dolphin and bird literature.
Benchmarks
Synthetic corpus (7 phrases)
| phrase | tokens | ones (of 50000) |
|---|---|---|
| rising_call | R R R | 24916 |
| falling_call | F F F | 24999 |
| contact_arc | R W F | 24905 |
| alarm_double | W U | 25017 |
| multi_phrase | R F W U S | 24825 |
| flat_long | S S S S | 25075 |
| mixed | R S F M | 24862 |
Pairwise similarity (selected)
| pair | similarity |
|---|---|
| rising_call vs falling_call | 0.625 |
| contact_arc vs mixed | 0.637 |
| contact_arc vs multi_phrase | 0.627 |
| alarm_double vs flat_long | 0.435 |
| rising_call vs contact_arc | 0.562 |
Orthogonal baseline: 0.5. The matrix is meaningfully structured: phrases that share a token or a bigram score higher than phrases that don't.
Audio round-trip
| synthesized | extracted |
|---|---|
| R | R |
| F | F |
| S | S |
The zero-crossing frequency estimator is accurate enough on clean whistles to recover the intended contour.
When to use it
- Symbolic search. Given a corpus of vocalizations, find entries whose contour pattern matches a query.
- Clustering by contour. Group signals by shape, not by frequency.
- Compression. A 5-token contour sequence becomes a single 50 Kbit vector, comparable by Hamming distance in O(d) time.
- Cross-species comparison. The six-token alphabet is language-agnostic and species-agnostic.
When not to use it
- For pitch tracking. The zero-crossing estimator is crude. Harmonics, overlapping calls, and broadband noise break it.
- For polyphonic audio. The model assumes one dominant frequency at a time.
- For very long recordings.
max_len = 64events per phrase. - For matching without reference data. The hypervector is only meaningful compared to another hypervector. There is no absolute semantic readout.
Honest limitations
- The contour alphabet is hand-designed. The six tokens come from
the dolphin literature. A corpus-adapted alphabet would discover its
own tokens (see the
hv-lexidea from the same series). - The zero-crossing frequency estimator is crude. It works on clean whistles and fails on anything with harmonics.
- The similarity range is narrow. At
dim=50000, values lie roughly between 0.44 and 0.64. Biggerdimimproves resolution at linear memory cost. - The bigram binding introduces a small systematic offset. Two
sequences with the same number of tokens at the same positions score
higher than orthogonal even if the tokens differ entirely. This is
visible in the
rising_callvsfalling_callsimilarity of 0.625. - The demo corpus is small. Seven synthetic phrases. On a real corpus of hundreds of whistles, the similarity landscape will be noisier.
Reference
Part of the reader-model series, which has grown beyond reading and into symbolic audio.
Companion to hv-sign (symbolic audio signature), hv-note (fixed
alphabet), hv-core (argument robustness), and hv-tempo (pace
variation).
hv-contour is the piece that reduces frequency modulation to a single
hypervector in a fixed-dimensional space where similarity is meaningful.
License
Apache-2.0