Instructions to use nvidia/Nemotron-3-Diarization with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use nvidia/Nemotron-3-Diarization with NeMo:
# tag did not correspond to a valid NeMo domain.
- Transformers
How to use nvidia/Nemotron-3-Diarization with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForAudioFrameClassification processor = AutoProcessor.from_pretrained("nvidia/Nemotron-3-Diarization") model = AutoModelForAudioFrameClassification.from_pretrained("nvidia/Nemotron-3-Diarization", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CPU-only on Windows + Japanese audio: test results
Sharing results in case they're useful to others:
NeMo-Speech.cpp, CPU build, v3-offline preset, no GPU
193 s Japanese two-speaker conversation: 7.6 s on Ryzen 7 5800X (RTF 0.039), 9.9 s on Core i5-8500 (RTF 0.051), ~170 MB peak
Japanese isn't in the listed dataset languages, but a listen-through showed speakers mostly matched; boundaries bleed slightly at quick turn changes
Output was byte-identical across runs and matched between the two machines
Called from C# (.NET Framework 4.8) via the C API with P/Invoke; same result as the CLI
On Windows, the build needed vcomp140.dll bundled next to the EXE.
Full write-up with code: https://kuboya.dev/articles/nemotron-3-diarization-on-windows-cpu/
For overlap speech look at : https://huggingface.co/nvidia/multitalker-parakeet-streaming-0.6b-v1 There are open source models that supports jp. The model card explains how to train a paired asr_diar
Thanks! That's exactly the problem I ran into at overlaps. If I understand correctly, the released checkpoint is English-only for now, and Japanese would mean fine-tuning a paired ASR + diarization setup myself β I'll look into it. For now my tool flags overlapping lines for manual review instead of assigning them to one speaker.
The Diarization model is language agnostic.
Yes, you are correct that the MultiTalker is En only for now as an example. You can check : https://huggingface.co/blog/nvidia/fine-tuning-nemotron-35-asr which is fine tuning for RNNT-Cache-aware. similar architecture.
The skill can also help :https://github.com/NVIDIA/skills/tree/main/skills/nemotron-asr-finetune
Thanks, that's really helpful β good to have it confirmed that the diarization model is language-agnostic, since that's what I saw with Japanese.
I'll take a look at the fine-tuning guide and the skill. Appreciate the pointers!