Automatic Speech Recognition
Transformers
Safetensors
English
Chinese
voxtral_realtime
asr
speech-recognition
streaming-asr
speech
audio
multimodal
voxtral
chinese
english
code-switching
Instructions to use x-square-robot/X2Streaming-ASR-4B-1009 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use x-square-robot/X2Streaming-ASR-4B-1009 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="x-square-robot/X2Streaming-ASR-4B-1009")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("x-square-robot/X2Streaming-ASR-4B-1009") model = AutoModelForMultimodalLM.from_pretrained("x-square-robot/X2Streaming-ASR-4B-1009", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download processor_config.json from x-square-robot/X2Streaming-ASR-4B-1009: direct link, hf CLI and curl.
- Browser
- Download file 384 Bytes
-
https://huggingface.co/x-square-robot/X2Streaming-ASR-4B-1009/resolve/main/processor_config.json
- Command line
-
hf download hf://x-square-robot/X2Streaming-ASR-4B-1009/processor_config.json
-
curl -L -o processor_config.json https://huggingface.co/x-square-robot/X2Streaming-ASR-4B-1009/resolve/main/processor_config.json
384 Bytes
| { | |
| "feature_extractor": { | |
| "feature_extractor_type": "VoxtralRealtimeFeatureExtractor", | |
| "feature_size": 128, | |
| "global_log_mel_max": 1.5, | |
| "hop_length": 160, | |
| "n_fft": 400, | |
| "padding_side": "right", | |
| "padding_value": 0.0, | |
| "return_attention_mask": true, | |
| "sampling_rate": 16000, | |
| "win_length": 400 | |
| }, | |
| "processor_class": "VoxtralRealtimeProcessor" | |
| } | |