audio-cpp commited on
Commit
4cd7a27
·
verified ·
1 Parent(s): ef2ff9d

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +40 -23
README.md CHANGED
@@ -50,33 +50,50 @@ or diarization results.
50
  For server use, configure family `audio_flamingo`, task `asr`, and mode `offline`,
51
  then submit audio to `/v1/audio/transcriptions` with the instruction in `prompt`.
52
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
53
  ## CUDA Performance and Quantization
54
 
55
- The following results use an RTX 5090, the CUDA backend, 8 CPU threads, and the
56
- same decoded 10-second, 16-kHz mono speech input. Wall time and RTF are from the
57
- second request in an already-loaded session. Peak VRAM covers the complete
58
- three-request session and was sampled every 100 ms.
59
-
60
- | Model | Weight type | Warm wall time | RTF | Peak VRAM | Transcript compared with BF16 |
61
- | --- | --- | ---: | ---: | ---: | --- |
62
- | Audio Flamingo 3 | BF16 | 283.3 ms | 0.0283 | 16,854 MiB | Reference |
63
- | Audio Flamingo 3 | Q8_0 | 184.9 ms | 0.0185 | 9,956 MiB | Exact |
64
- | Audio Flamingo 3 | Q4_K | 151.7 ms | 0.0152 | 6,280 MiB | Same words; wrapper and punctuation changed |
65
- | Audio Flamingo Next | BF16 | 246.3 ms | 0.0246 | 16,890 MiB | Reference |
66
- | Audio Flamingo Next | Q8_0 | 160.9 ms | 0.0161 | 9,992 MiB | Same words; one comma omitted |
67
- | Audio Flamingo Next | Q4_K | 129.9 ms | 0.0130 | 6,314 MiB | Exact |
68
-
69
- The BF16 Audio Flamingo 3 transcription and music-answer outputs were
70
- byte-identical to the official Transformers implementation. Audio Flamingo Next
71
- was also byte-identical on transcription. On a separate three-minute timestamped
72
- request, its words and line boundaries matched Python after timestamps were
73
- removed, while generated timestamps drifted.
74
 
75
  Quantized outputs remained coherent on the transcription and music-understanding
76
- requests. Q8_0 is the balanced default for lower VRAM and faster inference. Q4_K
77
- offers the lowest tested VRAM, but quantization can change wording, punctuation,
78
- or factual details in open-ended answers and should not be treated as exact
79
- parity with BF16.
80
 
81
  ## Audio Preprocessing
82
 
 
50
  For server use, configure family `audio_flamingo`, task `asr`, and mode `offline`,
51
  then submit audio to `/v1/audio/transcriptions` with the instruction in `prompt`.
52
 
53
+ ## Correctness Reference
54
+
55
+ Official Python BF16 output is the reference. Correctness comparisons and
56
+ performance measurements are separate tests.
57
+
58
+ | Model | Runtime / weights | Result compared with official Python |
59
+ | --- | --- | --- |
60
+ | Audio Flamingo 3 | Official Python BF16 | Reference |
61
+ | Audio Flamingo 3 | audio.cpp BF16 | Exact transcription and music answer |
62
+ | Audio Flamingo 3 | audio.cpp Q8_0 | Exact transcription; coherent but different music answer |
63
+ | Audio Flamingo 3 | audio.cpp Q4_K | Same transcription words with different wrapper and punctuation; coherent but different music answer |
64
+ | Audio Flamingo Next | Official Python BF16 | Reference |
65
+ | Audio Flamingo Next | audio.cpp BF16 | Exact short transcription; matching words and line boundaries on the three-minute test after timestamps were removed |
66
+ | Audio Flamingo Next | audio.cpp Q8_0 | Same short transcription words with one comma omitted |
67
+ | Audio Flamingo Next | audio.cpp Q4_K | Exact short transcription |
68
+
69
+ Audio Flamingo Next generates timestamps as text rather than structured
70
+ alignment. Its BF16 three-minute response used the same words and line boundaries
71
+ as Python after timestamps were removed, but the generated timestamps drifted.
72
+
73
  ## CUDA Performance and Quantization
74
 
75
+ The following results use an RTX 5090 and the same decoded 10-second, 16-kHz
76
+ mono speech input. Python uses the official Transformers BF16 model with SDPA.
77
+ audio.cpp uses the CUDA backend and 8 CPU threads. Wall time and RTF are from
78
+ the second request in an already-loaded session. Peak VRAM covers the complete
79
+ process and was sampled every 100 ms.
80
+
81
+ | Model | Runtime / weights | Warm wall time | RTF | Peak VRAM |
82
+ | --- | --- | ---: | ---: | ---: |
83
+ | Audio Flamingo 3 | Official Python BF16 | 292.4 ms | 0.0292 | 16,766 MiB |
84
+ | Audio Flamingo 3 | audio.cpp BF16 | 283.3 ms | 0.0283 | 16,854 MiB |
85
+ | Audio Flamingo 3 | audio.cpp Q8_0 | 184.9 ms | 0.0185 | 9,956 MiB |
86
+ | Audio Flamingo 3 | audio.cpp Q4_K | 151.7 ms | 0.0152 | 6,280 MiB |
87
+ | Audio Flamingo Next | Official Python BF16 | 249.0 ms | 0.0249 | 16,784 MiB |
88
+ | Audio Flamingo Next | audio.cpp BF16 | 246.3 ms | 0.0246 | 16,890 MiB |
89
+ | Audio Flamingo Next | audio.cpp Q8_0 | 160.9 ms | 0.0161 | 9,992 MiB |
90
+ | Audio Flamingo Next | audio.cpp Q4_K | 129.9 ms | 0.0130 | 6,314 MiB |
 
 
 
91
 
92
  Quantized outputs remained coherent on the transcription and music-understanding
93
+ requests. Q4_K is the default package and offers the lowest tested VRAM. Q8_0 is
94
+ the higher-precision quantized alternative. Quantization can change wording,
95
+ punctuation, or factual details in open-ended answers and should not be treated
96
+ as exact parity with BF16.
97
 
98
  ## Audio Preprocessing
99