hFT-Transformer β€” GGUF

Toyama et al.'s hierarchical frequency-time transformer for piano transcription (ISMIR 2023, sony/hFT-Transformer, MIT), converted to GGUF for the ggml runtime in CrispASR:

crispasr --piano -m hft-transformer-q4_0.gguf -f input.wav

Files

file size note F1 F1 with offsets solo piano
hft-transformer-f32.gguf 21.8 MiB 52.21% 18.49% 70.52%
hft-transformer-q8_0.gguf 7.0 MiB 52.23% 18.50% 70.51%
hft-transformer-q4_0.gguf 4.5 MiB 52.55% 18.77% 70.71%

Quantisation is free here, q4_0 included. That is not the usual result and the reason is specific: the head q4_0 perturbs most is velocity (cosine 0.870), and in this model velocity feeds the decoder's ignore_zero gate rather than a note duration β€” so its errors change which notes survive, not how long they last. Contrast the Onsets & Frames GGUF, where q4_0 damages the frame head and costs 0.5 points of F1-with-offsets, because that head sets durations.

4.5 MiB reaching 70.7% on solo piano makes this the smallest model in this family that reaches that accuracy.

Accuracy and provenance

Converted from the pruned ONNX export of the sony/hFT-Transformer checkpoint. The converter fuses the (1, 5) convolution into tok_embedding_freq β€” both are linear in the 65-tap window with nothing between them, so they collapse exactly into one Linear(65, 256) (verified at 7.1e-07 relative).

The f32 build is the model, not an approximation of it: head-identical to the ONNX export (onset max abs 2.5e-06, cosine 1.00000000, 100.0000% of decisions identical including the velocity-gate argmax), and F1-identical on all ten pieces of MusicNet's test split.

Figures above are all ten MusicNet test pieces, 13,589 reference notes, scored with mir_eval.transcription (50 ms onset tolerance, 50 cents, maximum bipartite matching), every arm through the same decoder. Method and full tables: docs/music-transcription/HFT_TRANSFORMER.md and Β§35–36 of the CrispTuner report.

Cost β€” read this before choosing it

On a 4-vCPU Skylake-SP VPS, f32 is the fastest arm at 2.14Γ— real time on four threads (4.84 CPU-s per audio-second single-threaded). Quantisation costs 29% of throughput while buying the size above, because ggml's repacked int8 GEMM is reached only through the CPU device's extra buffer types and this build does not request them β€” and that host has AVX-512F but no VNNI, so there is no int8 dot-product instruction to exploit either. The int8 case is untested here rather than disproved.

Native ONNX Runtime is 1.28Γ— real time on the same box, so ggml is 1.56Γ— its CPU cost β€” but at 237 MiB peak RSS against ORT's ~1001 MiB.

For throughput, onsets-and-frames is the better piano arm at 0.44Γ— real time. Choose hFT for accuracy per megabyte, or where 4.5 MiB matters.

Downloads last month
721
GGUF
Model size
5.7M params
Architecture
hft-transformer
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support