Update README.md
Browse files
README.md
CHANGED
|
@@ -50,33 +50,50 @@ or diarization results.
|
|
| 50 |
For server use, configure family `audio_flamingo`, task `asr`, and mode `offline`,
|
| 51 |
then submit audio to `/v1/audio/transcriptions` with the instruction in `prompt`.
|
| 52 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 53 |
## CUDA Performance and Quantization
|
| 54 |
|
| 55 |
-
The following results use an RTX 5090
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
|
| 62 |
-
|
|
| 63 |
-
| Audio Flamingo 3 |
|
| 64 |
-
| Audio Flamingo 3 |
|
| 65 |
-
| Audio Flamingo
|
| 66 |
-
| Audio Flamingo
|
| 67 |
-
| Audio Flamingo Next |
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
was also byte-identical on transcription. On a separate three-minute timestamped
|
| 72 |
-
request, its words and line boundaries matched Python after timestamps were
|
| 73 |
-
removed, while generated timestamps drifted.
|
| 74 |
|
| 75 |
Quantized outputs remained coherent on the transcription and music-understanding
|
| 76 |
-
requests.
|
| 77 |
-
|
| 78 |
-
or factual details in open-ended answers and should not be treated
|
| 79 |
-
parity with BF16.
|
| 80 |
|
| 81 |
## Audio Preprocessing
|
| 82 |
|
|
|
|
| 50 |
For server use, configure family `audio_flamingo`, task `asr`, and mode `offline`,
|
| 51 |
then submit audio to `/v1/audio/transcriptions` with the instruction in `prompt`.
|
| 52 |
|
| 53 |
+
## Correctness Reference
|
| 54 |
+
|
| 55 |
+
Official Python BF16 output is the reference. Correctness comparisons and
|
| 56 |
+
performance measurements are separate tests.
|
| 57 |
+
|
| 58 |
+
| Model | Runtime / weights | Result compared with official Python |
|
| 59 |
+
| --- | --- | --- |
|
| 60 |
+
| Audio Flamingo 3 | Official Python BF16 | Reference |
|
| 61 |
+
| Audio Flamingo 3 | audio.cpp BF16 | Exact transcription and music answer |
|
| 62 |
+
| Audio Flamingo 3 | audio.cpp Q8_0 | Exact transcription; coherent but different music answer |
|
| 63 |
+
| Audio Flamingo 3 | audio.cpp Q4_K | Same transcription words with different wrapper and punctuation; coherent but different music answer |
|
| 64 |
+
| Audio Flamingo Next | Official Python BF16 | Reference |
|
| 65 |
+
| Audio Flamingo Next | audio.cpp BF16 | Exact short transcription; matching words and line boundaries on the three-minute test after timestamps were removed |
|
| 66 |
+
| Audio Flamingo Next | audio.cpp Q8_0 | Same short transcription words with one comma omitted |
|
| 67 |
+
| Audio Flamingo Next | audio.cpp Q4_K | Exact short transcription |
|
| 68 |
+
|
| 69 |
+
Audio Flamingo Next generates timestamps as text rather than structured
|
| 70 |
+
alignment. Its BF16 three-minute response used the same words and line boundaries
|
| 71 |
+
as Python after timestamps were removed, but the generated timestamps drifted.
|
| 72 |
+
|
| 73 |
## CUDA Performance and Quantization
|
| 74 |
|
| 75 |
+
The following results use an RTX 5090 and the same decoded 10-second, 16-kHz
|
| 76 |
+
mono speech input. Python uses the official Transformers BF16 model with SDPA.
|
| 77 |
+
audio.cpp uses the CUDA backend and 8 CPU threads. Wall time and RTF are from
|
| 78 |
+
the second request in an already-loaded session. Peak VRAM covers the complete
|
| 79 |
+
process and was sampled every 100 ms.
|
| 80 |
+
|
| 81 |
+
| Model | Runtime / weights | Warm wall time | RTF | Peak VRAM |
|
| 82 |
+
| --- | --- | ---: | ---: | ---: |
|
| 83 |
+
| Audio Flamingo 3 | Official Python BF16 | 292.4 ms | 0.0292 | 16,766 MiB |
|
| 84 |
+
| Audio Flamingo 3 | audio.cpp BF16 | 283.3 ms | 0.0283 | 16,854 MiB |
|
| 85 |
+
| Audio Flamingo 3 | audio.cpp Q8_0 | 184.9 ms | 0.0185 | 9,956 MiB |
|
| 86 |
+
| Audio Flamingo 3 | audio.cpp Q4_K | 151.7 ms | 0.0152 | 6,280 MiB |
|
| 87 |
+
| Audio Flamingo Next | Official Python BF16 | 249.0 ms | 0.0249 | 16,784 MiB |
|
| 88 |
+
| Audio Flamingo Next | audio.cpp BF16 | 246.3 ms | 0.0246 | 16,890 MiB |
|
| 89 |
+
| Audio Flamingo Next | audio.cpp Q8_0 | 160.9 ms | 0.0161 | 9,992 MiB |
|
| 90 |
+
| Audio Flamingo Next | audio.cpp Q4_K | 129.9 ms | 0.0130 | 6,314 MiB |
|
|
|
|
|
|
|
|
|
|
| 91 |
|
| 92 |
Quantized outputs remained coherent on the transcription and music-understanding
|
| 93 |
+
requests. Q4_K is the default package and offers the lowest tested VRAM. Q8_0 is
|
| 94 |
+
the higher-precision quantized alternative. Quantization can change wording,
|
| 95 |
+
punctuation, or factual details in open-ended answers and should not be treated
|
| 96 |
+
as exact parity with BF16.
|
| 97 |
|
| 98 |
## Audio Preprocessing
|
| 99 |
|