hFT-Transformer β GGUF
Toyama et al.'s hierarchical frequency-time transformer for piano
transcription (ISMIR 2023, sony/hFT-Transformer, MIT),
converted to GGUF for the ggml runtime in
CrispASR:
crispasr --piano -m hft-transformer-q4_0.gguf -f input.wav
Files
| file | size | note F1 | F1 with offsets | solo piano |
|---|---|---|---|---|
hft-transformer-f32.gguf |
21.8 MiB | 52.21% | 18.49% | 70.52% |
hft-transformer-q8_0.gguf |
7.0 MiB | 52.23% | 18.50% | 70.51% |
hft-transformer-q4_0.gguf |
4.5 MiB | 52.55% | 18.77% | 70.71% |
Quantisation is free here, q4_0 included. That is not the usual result
and the reason is specific: the head q4_0 perturbs most is velocity
(cosine 0.870), and in this model velocity feeds the decoder's ignore_zero
gate rather than a note duration β so its errors change which notes
survive, not how long they last. Contrast the Onsets & Frames GGUF, where
q4_0 damages the frame head and costs 0.5 points of F1-with-offsets,
because that head sets durations.
4.5 MiB reaching 70.7% on solo piano makes this the smallest model in this family that reaches that accuracy.
Accuracy and provenance
Converted from the pruned ONNX export of the sony/hFT-Transformer
checkpoint. The converter fuses the (1, 5) convolution into
tok_embedding_freq β both are linear in the 65-tap window with nothing
between them, so they collapse exactly into one Linear(65, 256)
(verified at 7.1e-07 relative).
The f32 build is the model, not an approximation of it: head-identical to the ONNX export (onset max abs 2.5e-06, cosine 1.00000000, 100.0000% of decisions identical including the velocity-gate argmax), and F1-identical on all ten pieces of MusicNet's test split.
Figures above are all ten MusicNet test pieces, 13,589 reference notes,
scored with mir_eval.transcription (50 ms onset tolerance, 50 cents,
maximum bipartite matching), every arm through the same decoder. Method and
full tables: docs/music-transcription/HFT_TRANSFORMER.md
and Β§35β36 of the
CrispTuner report.
Cost β read this before choosing it
On a 4-vCPU Skylake-SP VPS, f32 is the fastest arm at 2.14Γ real time on four threads (4.84 CPU-s per audio-second single-threaded). Quantisation costs 29% of throughput while buying the size above, because ggml's repacked int8 GEMM is reached only through the CPU device's extra buffer types and this build does not request them β and that host has AVX-512F but no VNNI, so there is no int8 dot-product instruction to exploit either. The int8 case is untested here rather than disproved.
Native ONNX Runtime is 1.28Γ real time on the same box, so ggml is 1.56Γ its CPU cost β but at 237 MiB peak RSS against ORT's ~1001 MiB.
For throughput, onsets-and-frames is the better piano arm at 0.44Γ real
time. Choose hFT for accuracy per megabyte, or where 4.5 MiB matters.
- Downloads last month
- 721
4-bit
8-bit
32-bit