SAM-Audio-GGUF / README.md
audio-cpp's picture
Update README.md
d429aa7 verified
|
Raw
History Blame Contribute Delete
9.96 kB
metadata
license: other
license_name: sam-license
license_link: LICENSE
base_model: facebook/sam-audio-large
library_name: audio.cpp
pipeline_tag: audio-to-audio
tags:
  - audio.cpp
  - gguf
  - audio-separation

SAM Audio GGUF

GGUF weights for native inference with audio.cpp. SAM Audio separates a sound described by text, a masked image or video, or positive and negative time spans from an input recording.

Upstream

These packages use the original Meta checkpoints, not the separate -tv variants.

Variant GGUF Pinned checkpoint
Small sam-audio-small-f32.gguf facebook/sam-audio-small
Base sam-audio-base-f32.gguf facebook/sam-audio-base
Large sam-audio-large-f32.gguf facebook/sam-audio-large

The official SAM Audio implementation is pinned to bb4c699.

Each GGUF contains SAM weights, T5 text encoder, tokenizer, configuration, and audio.cpp model spec. No separate text encoder is needed. F32 packages retain the original weights, checked byte-for-byte against the source tensors. BF16 and Q8_0 packages are also available for each variant; both were converted directly from the source safetensors. Biases, normalization vectors, and scale/shift tables remain F32. Q8_0 denotes mixed quantized storage, not that every tensor is quantized.

Usage

audiocpp_cli --task s2s --family sam_audio \
  --model /path/to/SAM-Audio-GGUF/sam-audio-small-f32.gguf --backend cuda \
  --audio input.wav --text "man speaking" --seed -1 \
  --out-dir outputs/separation --log

Outputs are target.wav and residual.wav, mono 48 kHz. Optional temporal prompts use --request-option 'anchors=[["+",0.5,2.0],["-",3.0,4.0]]'. The runtime also accepts a masked static image with --request-option reference_image_path=masked.png, or a masked video with --request-option reference_video_path=masked.mp4. Video decoding requires FFmpeg/libav runtime libraries. Use one visual option at a time and keep the video and input audio time origins aligned. Automatic span prediction and candidate ranking are not included.

See the model usage guide for all request options. Code tested on CUDA/Vulkan/Metal.

All parity, performance, and VRAM results below were measured with watermarking enabled to match the official Python implementation, which adds a watermark to its output. The final audio.cpp version will return audio WITHOUT the added watermark, so these results do not describe the final implementation.

Controlled Parity With Python

Parity was tested separately from performance on CUDA with F32 weights and TF32 disabled in both implementations (NVIDIA_TF32_OVERRIDE=0). Python used seed 42; C++ replayed its exact diffusion noise and watermark bits. Both used the same 14.975-second office recording, man speaking prompt, and 16 midpoint steps. C++ computed its own encoder features, diffusion, and decoded audio.

Each row compares the final C++ waveforms directly against official Python, including bounded-memory mode. Higher cosine similarity is better; 1 is exact directional agreement, not a claim of byte-identical waveforms.

Variant C++ mode Target cosine Residual cosine
Small F32 Normal 0.99999945 0.99999864
Small F32 Bounded memory 0.99999968 0.99999896
Base F32 Normal 0.99999956 0.99999558
Base F32 Bounded memory 0.99999950 0.99999753
Large F32 Normal 0.99999980 0.99999888
Large F32 Bounded memory 0.99999984 0.99999936

All rows exceeded 0.9999 for both outputs. Repeating each controlled C++ run within the same session produced byte-identical outputs.

Ordinary CUDA Performance With Python

Measured on an NVIDIA RTX 5090 (32 GB), using a 14.975-second recording from the official office example and the text prompt man speaking.

  • Both implementations use the original weights, 16 midpoint steps, one candidate, and no visual or span prompts. Optional ranking and span-prediction models are not loaded in Python.
  • Ordinary inference: no fixed or replayed random noise, no TF32 override, and no forced precision/autocast changes. C++ uses --seed -1; Python does not set a seed. Python uses PyTorch 2.11.0+cu128; C++ uses a debug build and 8 threads.
  • Warm inference time is the median of the last three of five sequential requests in one session. It excludes model loading and file I/O. RTF is inference time divided by input duration; lower is faster.
  • Peak VRAM is process memory sampled with nvidia-smi every 50 ms across loading and all five requests, using the same method for Python and C++.
Variant C++ warm time C++ RTF C++ peak VRAM Python warm time Python RTF Python peak VRAM
Small F32 684 ms 0.0456 8,500 MiB 746 ms 0.0498 10,972 MiB
Base F32 918 ms 0.0613 11,010 MiB 1,105 ms 0.0738 13,414 MiB
Large F32 1,542 ms 0.1030 17,826 MiB 2,103 ms 0.1405 20,568 MiB

These are text-conditioned, single-recording measurements, not a benchmark of visual prompting, automatic span prediction, or candidate ranking.

Outputs from these independent-randomness performance runs are not used for parity comparisons.

BF16 and Q8 Compared With C++ F32

The following are a separate C++-only comparison on the same RTX 5090 and 14.975-second input. Ordinary performance uses the five-request methodology above, --seed -1, and no TF32 or precision overrides. F32 was remeasured for this comparison; no Python model was run.

Variant Storage Warm time RTF Peak VRAM
Small F32 700 ms 0.04673 8,500 MiB
Small BF16 674 ms 0.04501 7,366 MiB
Small Q8_0 673 ms 0.04493 6,834 MiB
Base F32 960 ms 0.06408 11,010 MiB
Base BF16 785 ms 0.05241 8,644 MiB
Base Q8_0 794 ms 0.05302 7,524 MiB
Large F32 1,566 ms 0.10457 17,826 MiB
Large BF16 1,146 ms 0.07652 12,092 MiB
Large Q8_0 1,051 ms 0.07018 9,378 MiB

Output drift was measured in separate controlled C++ runs, using seed 42 and NVIDIA_TF32_OVERRIDE=0 for every dtype. Each output is compared against that variant's C++ F32 waveform, not Python. These runs supply no performance or VRAM figures.

Variant Storage Target cosine vs F32 Residual cosine vs F32
Small BF16 0.999940 0.999763
Small Q8_0 0.999476 0.999095
Base BF16 0.999918 0.999788
Base Q8_0 0.999623 0.999069
Large BF16 0.999743 0.999712
Large Q8_0 0.999590 0.999273

BF16 reduced peak VRAM by 13-32%; Q8_0 reduced it by 20-47% in this workload. Drift is expected. These short text-conditioned checks do not establish equal quality for every recording or validate reduced-precision visual prompting and bounded-memory long-form execution. F32 remains available as the reference.

Bounded-Memory Mode

Enable the experimental session option to reduce long-recording GPU workspace:

--session-option sam_audio.memory_bounded=true

The option defaults to false, and its name is provisional. The existing Small GGUF predates this option; add --model-spec-override model_specs/sam_audio.json from a checkout that includes bounded-memory support. The Base and Large GGUFs include the option in their embedded specs.

The mode retains full-recording diffusion and attention. Codec convolutions use tiles with receptive-field context, and watermark LSTM state carries between tiles. It does not separate independent audio chunks and crossfade them. Intermediate sequences reside in host RAM, which still grows with input length. Transfers and overlap computation can make it slower on short inputs; small numerical differences are possible. The configured model context limit remains 400 seconds.

The following CUDA tests used a 180-second input made by repeating the same example to 180 seconds, with the same prompt and inference settings as above. Each C++ result is the first inference in a fresh session, including graph setup but excluding model loading and file I/O. Peak VRAM includes loading and inference. These are not warm medians.

Variant C++ bounded time C++ RTF C++ peak VRAM Official Python, same input
Small F32 8.768 s 0.0487 6,550 MiB CUDA out of memory
Base F32 11.109 s 0.0617 9,212 MiB CUDA out of memory
Large F32 17.038 s 0.0947 16,244 MiB CUDA out of memory

All three C++ runs produced the complete target and residual recordings. Python ran out of memory in its audio encoder on the same 32 GB GPU, using its default precision and allocator settings. This demonstrates completion and memory usage for this workload, not 180-second parity against Python: Python did not produce a full-length reference. Other hardware or Python memory-saving configurations may behave differently.

License

SAM Audio weights are distributed under the SAM License, dated November 19, 2025. The included T5 encoder is from Google T5 Base and retains its Apache-2.0 license. Conversion does not replace the upstream license terms.