Voice Activity Detection
NeMo
Safetensors
GGUF
Transformers
nemotron3_diarization
audio-frame-classification
speaker-diarization
streaming-sortformer
speaker-tagging
Instructions to use nvidia/Nemotron-3-Diarization with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use nvidia/Nemotron-3-Diarization with NeMo:
# tag did not correspond to a valid NeMo domain.
- Transformers
How to use nvidia/Nemotron-3-Diarization with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForAudioFrameClassification processor = AutoProcessor.from_pretrained("nvidia/Nemotron-3-Diarization") model = AutoModelForAudioFrameClassification.from_pretrained("nvidia/Nemotron-3-Diarization", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download processor_config.json from nvidia/Nemotron-3-Diarization: direct link, hf CLI and curl.
- Browser
- Download file 623 Bytes
-
https://huggingface.co/nvidia/Nemotron-3-Diarization/resolve/main/processor_config.json
- Command line
-
hf download hf://nvidia/Nemotron-3-Diarization/processor_config.json
-
curl -L -o processor_config.json https://huggingface.co/nvidia/Nemotron-3-Diarization/resolve/main/processor_config.json
623 Bytes
| { | |
| "feature_extractor": { | |
| "feature_extractor_type": "NemotronAsrStreamingFeatureExtractor", | |
| "feature_size": 128, | |
| "hop_length": 160, | |
| "n_fft": 512, | |
| "padding_side": "right", | |
| "padding_value": 0.0, | |
| "preemphasis": 0.97, | |
| "return_attention_mask": true, | |
| "sampling_rate": 16000, | |
| "win_length": 400 | |
| }, | |
| "processor_class": "Nemotron3DiarizationProcessor", | |
| "streaming_mode": "low_latency", | |
| "streaming_modes": { | |
| "low_latency": [ | |
| 9, | |
| 4 | |
| ], | |
| "ultra_low_latency": [ | |
| 3, | |
| 1 | |
| ], | |
| "very_low_latency": [ | |
| 6, | |
| 2 | |
| ] | |
| }, | |
| "subsampling_factor": 8 | |
| } | |