Voxtral Mini Standalone Audio Feature Extractor

A lightweight standalone audio feature extractor derived from Voxtral-Mini-3B-2507.

This repository isolates the Whisper-based audio encoder and multi-modal projector from the original 3B language model. The extracted module is intended for offline audio preprocessing, dataset preparation, feature caching, and downstream multimodal training pipelines.

The language-model decoder is not included.

Voxtral Audio Tower Architecture

Raw audio
    โ”‚
    โ–ผ
HF Audio Feature Extractor
    โ”‚
    โ–ผ
Log-Mel features
    โ”‚
    โ–ผ
Voxtral Whisper-encoder
    โ”‚
    โ–ผ
[B, T, 1280]
    โ”‚
    โ–ผ
Feature packing ร—4
    โ”‚
    โ–ผ
[B, T/4, 5120]
    โ”‚
    โ–ผ
Multi-Modal Projector
    โ”‚
    โ–ผ
[B, T/4, 3072]

The extracted audio branch contains:

  • Voxtral / Whisper-based audio encoder
  • Voxtral multi-modal projector
  • Voxtral feature packing operation

It does not contain the 3B LLaMA language-model decoder.

Why?

For offline preprocessing, loading the complete multimodal language model is unnecessary when the only required output is the projected audio representation.

This standalone checkpoint can therefore be used as a dedicated audio feature extraction stage:

audio
  โ†“
audio processor
  โ†“
precomputed audio embeddings
  โ†“
dataset cache
  โ†“
LLM / adapter / multimodal training

This is particularly useful for large datasets where audio features can be computed once and reused across multiple training runs.

Output

For an input Mel tensor with shape:

[B, 128, T]

the extractor produces:

[B, T/4, 3072]

For example:

Input:
[1, 128, 1500]

Output:
[1, 375, 3072]

The current reference implementation uses bfloat16.

Quickstart

import torch
import soundfile as sf
import torchaudio.functional as F

from transformers import AutoModel, AutoFeatureExtractor

MODEL_ID = "vxltxr/voxtral-mini-audio-extractor"

device = "cuda" if torch.cuda.is_available() else "cpu"

# Standalone audio encoder + projector.
# No 3B LLM decoder is loaded.
model = AutoModel.from_pretrained(
    MODEL_ID,
    trust_remote_code=True,
    dtype=torch.bfloat16,
).to(device).eval()

# Use the original Voxtral feature extractor for audio -> log-Mel preprocessing.
feature_extractor = AutoFeatureExtractor.from_pretrained(
    "mistralai/Voxtral-Mini-3B-2507"
)

# Load audio.
audio_data, sampling_rate = sf.read(
    "sample.wav",
    dtype="float32",
)

# Convert stereo -> mono if necessary.
if audio_data.ndim > 1:
    audio_data = audio_data.mean(axis=1)

# Resample to 16 kHz when necessary.
if sampling_rate != 16000:
    waveform = torch.from_numpy(audio_data)
    waveform = F.resample(
        waveform,
        orig_freq=sampling_rate,
        new_freq=16000,
    )
    audio_data = waveform.numpy()
    sampling_rate = 16000

# Audio -> log-Mel features.
inputs = feature_extractor(
    audio_data,
    sampling_rate=sampling_rate,
    return_tensors="pt",
)

mel = inputs["input_features"].to(
    device=device,
    dtype=torch.bfloat16,
)

# Audio encoder -> packing -> multi-modal projector.
with torch.inference_mode():
    audio_embeds = model.extract_features(mel)

print("Mel shape:       ", mel.shape)
print("Embeddings shape:", audio_embeds.shape)
print("Embeddings dtype:", audio_embeds.dtype)

Offline Dataset Preprocessing

The intended use case is to precompute audio embeddings before training.

Conceptually:

def preprocess_batch(batch):
    inputs = feature_extractor(
        batch["audio"],
        sampling_rate=16000,
        return_tensors="pt",
    )

    mel = inputs["input_features"].to(
        device="cuda",
        dtype=torch.bfloat16,
    )

    with torch.inference_mode():
        features = model.extract_features(mel)

    return {
        "audio_features": features.cpu().numpy(),
    }

This allows the expensive audio encoder to run once during dataset preparation instead of during every training step.

Standalone Checkpoint

The extracted artifact contains the weights of:

audio_tower.*
multi_modal_projector.*

The extraction currently contains 489 tensors.

The standalone branch was validated against the corresponding reference audio graph with:

Output shape:       [1, 375, 3072]
Max absolute diff:  0.0
Mean absolute diff: 0.0
Cosine similarity:  0.999999881

This verifies the extracted BF16 audio branch against the reference implementation for the tested input.

Relationship to Voxtral

This repository is derived from:

mistralai/Voxtral-Mini-3B-2507

The upstream Voxtral architecture combines an audio encoder and multi-modal projector with a language-model decoder. This repository extracts only the audio feature path for standalone preprocessing.

The Hugging Face Voxtral implementation describes get_audio_features() as the path that takes log-Mel audio features through the audio encoder and multi-modal projector to obtain audio embeddings.

Intended Use

Good fits include:

  • offline audio feature extraction
  • multimodal dataset preprocessing
  • cached audio embeddings
  • adapter / projector experiments
  • multimodal LLM training pipelines
  • large-scale dataset preparation
  • debugging and analysis of the Voxtral audio branch

Not Intended For

This checkpoint is not a speech-to-text model.

It does not contain:

  • the 3B LLaMA decoder
  • text generation weights
  • a tokenizer for generation
  • the full Voxtral conditional-generation pipeline

Its output is an intermediate audio representation intended to be consumed by a downstream model.

Notes

The current checkpoint expects the Voxtral-compatible audio preprocessing pipeline to produce the appropriate log-Mel input_features.

For production dataset preprocessing, keep the audio preprocessing configuration aligned with the original Voxtral model.

License

Apache-2.0.

This repository contains extracted components derived from the upstream Voxtral model. Please also review the upstream model's license and terms before redistribution or deployment.

Downloads last month
84
Safetensors
Model size
0.7B params
Tensor type
F32
ยท
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support