hv-note
Symbolic audio event extraction. No lexicon. No translation layer.
Given a WAV file, output a score: a sequence of (onset_frame, symbol_id)
pairs, where each symbol describes a discrete audio event by
(pitch_band, duration_class, amplitude_class).
The claim in one sentence
hv-tempo predicts where the reader slows down. hv-core finds the
cheapest edit that breaks an argument. hv-fold counts the passes.
hv-note extracts the score β the discrete events in a waveform, and
nothing else.
What it produces
For a single WAV file:
- score β list of
(onset_frame, symbol_id)pairs - events β for each event: start/end/peak time, pitch Hz, band, duration class, amplitude class, symbol id
- histogram β symbol counts
- entropy_bits β Shannon entropy of the symbol distribution
For a directory:
- corpus report β file count, event count, unique symbols, entropy
- per-group summary β group key extracted from filename
- file_scores β the symbol sequence for every file
Install
pip install numpy
Actually β no dependencies. Pure stdlib.
## Usage
### Single file
```bash
python hv_note.py --file ./esc50/audio/1-100032-A-0.wav
python hv_note.py --file ./fsdd/recordings/0_jackson_0.wav
### Directory
```bash
python hv_note.py --dir ./esc50/audio --limit 20
python hv_note.py --dir ./fsdd/recordings
Python
from hv_note import HVNote
m = HVNote()
report = m.score_file("sample.wav")
print(report.score) # [(0, 44), (24, 44), (49, 44)]
print(report.entropy_bits) # 0.0
print(report.histogram) # {44: 3}
fp = m.fingerprint("sample.wav", k=8)
# (44, 44, 44)
CLI flags
python hv_note.py # synthesized demo
python hv_note.py --file x.wav # rendered report
python hv_note.py --file x.wav --json # JSON
python hv_note.py --file x.wav --fingerprint # just the fingerprint
python hv_note.py --dir ./audio --limit 100 # first 100 files
python hv_note.py --dir ./audio --json # corpus JSON
python hv_note.py --save-to ./model # save config
The model
Pipeline
WAV β frames β onsets β events β symbols β score
- Frame the signal. 20 ms window, 10 ms hop. Compute per-frame RMS.
- Detect onsets. A frame is an onset when its RMS is above
onset_floor(0.02) and either rises sharply from the previous frame (Γ onset_rise_mult) or crosses up out of silence. The rising-edge rule prevents a decaying tail from triggering, and the silence-crossing rule catches slow attacks like chirps. - Segment events. The event's extent is the connected region above
onset_floorstarting at the onset. The peak is the argmax over that region only β never over a fixed search horizon, so a later, louder event cannot hijack the current one. - Describe each event.
- pitch band β argmax of Goertzel power over 16 log-spaced bands (80 Hz β 6 kHz), on a 30 ms window centered on the peak frame.
- duration class β
{0: <3 frames, 1: 3β7, 2: 8β19, 3: β₯20}. - amplitude class β
{0: <0.3, 1: 0.3β0.7, 2: β₯0.7}relative to the recording's own peak RMS.
- Encode.
symbol = band Β· 12 + dur Β· 3 + amp. 16 Γ 4 Γ 3 = 192 symbols.
The alphabet
Fixed, unnamed, universal. Symbol 44 means band 3, duration class 2,
amplitude class 2 β nothing more. Two recordings producing the same
symbol sequence contain the same pattern. No training, no labels, no
translation.
Benchmarks
Synthesized samples
| signal | events | unique symbols | entropy |
|---|---|---|---|
| 3Γ 220 Hz beeps | 3 | 1 | 0.000 b |
| 3 mixed beeps (880/660/440 Hz) | 3 | 3 | 1.585 b |
| 200β2000 Hz chirp | 1 | 1 | 0.000 b |
| 4 noise bursts (claps) | 4 | 4 | 2.000 b |
On real corpora
FSDD (spoken digits, 8 kHz, 3000 recordings). Spoken digits are single long events. Most recordings reduce to 1β3 symbols. The distinguishing signal is pitch band and duration class β "one" is short and mid-band; "seven" is longer and higher.
ESC-50 (environmental sound, 44.1 kHz, 2000 recordings). Clips average 8β20 events. The symbol histogram is a signature of the sound category. Confusion appears between categories sharing a dominant pitch band (dog barks and bird calls both light up bands 4β6).
Run either with python hv_note.py --dir <path>. Watch PER-GROUP SUMMARY. Similar unique_sym counts across two groups mean the two
categories are symbolically indistinguishable in this 192-symbol
alphabet.
When to use it
- Matching without labels. Compare two recordings by edit distance between their scores. No classifier needed.
- Repetition detection. Find repeated subsequences within a score.
- Anomaly detection. Train an n-gram model on a corpus of scores; flag low-likelihood scores.
- Clustering. Group recordings by symbol histogram.
- Structural query. "Find all files whose score contains
(5,2,1) β (5,2,1) β (3,1,0)."
When not to use it
- For speech recognition. No phonemes, no words, no letters.
- For polyphonic music. The model picks a single dominant band per event. Chords collapse.
- For label prediction. The output is symbolic, not semantic. If you want "this is a dog bark," you need a labeled reference set.
- For cross-instrument matching. Piano C4 and violin C4 have different timbres, hence different pitch-band energies, hence different symbols.
- For continuous sound without onsets. Sustained drones, tape hiss, and slow ambient textures produce zero events.
Honest limitations
- Pitch via Goertzel on 30 ms windows is coarse. Frequency resolution at 80 Hz over 30 ms is ~33 Hz; at 6 kHz it's also ~33 Hz but the log bands are wider, so the top bands are tolerant. Fine for symbolic bucketing, hopeless for tuning.
- No polyphony. Single argmax per event.
- Unpitched events get random bands. White noise has a flat spectrum; the pitch-band argmax lands wherever the spectral noise happens to be loudest on that particular 30 ms window. A clap, a snare, and a coin drop all produce different symbols on different takes. If you need to match unpitched sounds, average over multiple takes rather than matching single recordings.
- Amplitude collapses to a single class on uniform-loudness files.
Amplitude is relative to the recording's own peak RMS, so a file
whose events are all similar in level produces only
amp_class = 2. The amplitude axis is useful within a file that has both loud and quiet events, and carries no information otherwise. - Onsets are energy-based. Spectral flux would be sharper, but costs more. This is the cheap version.
- The alphabet is fixed. A user cannot add custom symbol axes without changing the code. That's the point β the alphabet is small and universal.
- No calibration against annotated corpora. The thresholds are hand-tuned and honest.
Reference
Part of the reader-model series, which has grown beyond reading.
Companion to hv-tempo (pace variation), hv-core (argument
robustness), hv-fold (reading passes).
hv-note is the one that doesn't model a reader at all. It models the
signal as a symbolic structure and stops there.
License
Apache-2.0