CPU-only on Windows + Japanese audio: test results

#5
by kuboya - opened

Sharing results in case they're useful to others:

NeMo-Speech.cpp, CPU build, v3-offline preset, no GPU
193 s Japanese two-speaker conversation: 7.6 s on Ryzen 7 5800X (RTF 0.039), 9.9 s on Core i5-8500 (RTF 0.051), ~170 MB peak
Japanese isn't in the listed dataset languages, but a listen-through showed speakers mostly matched; boundaries bleed slightly at quick turn changes
Output was byte-identical across runs and matched between the two machines
Called from C# (.NET Framework 4.8) via the C API with P/Invoke; same result as the CLI

On Windows, the build needed vcomp140.dll bundled next to the EXE.

Full write-up with code: https://kuboya.dev/articles/nemotron-3-diarization-on-windows-cpu/

NVIDIA org

For overlap speech look at : https://huggingface.co/nvidia/multitalker-parakeet-streaming-0.6b-v1 There are open source models that supports jp. The model card explains how to train a paired asr_diar

Thanks! That's exactly the problem I ran into at overlaps. If I understand correctly, the released checkpoint is English-only for now, and Japanese would mean fine-tuning a paired ASR + diarization setup myself β€” I'll look into it. For now my tool flags overlapping lines for manual review instead of assigning them to one speaker.

NVIDIA org

The Diarization model is language agnostic.
Yes, you are correct that the MultiTalker is En only for now as an example. You can check : https://huggingface.co/blog/nvidia/fine-tuning-nemotron-35-asr which is fine tuning for RNNT-Cache-aware. similar architecture.
The skill can also help :https://github.com/NVIDIA/skills/tree/main/skills/nemotron-asr-finetune

Thanks, that's really helpful β€” good to have it confirmed that the diarization model is language-agnostic, since that's what I saw with Japanese.
I'll take a look at the fine-tuning guide and the skill. Appreciate the pointers!

Sign up or log in to comment