NVIDIA Nemotron 3 Diarization Explained
Share

Post Contenthttps://www.youtube.com/shorts/D7Gn9IraysI 

 [[{“value”:”NVIDIA just dropped Nemotron 3 Diarization, a 100M-parameter open model that tells apart up to 8 speakers in real-time streaming audio. Pair it with a speech-to-text model like Parakeet and every line of your meeting notes knows who said it. It runs entirely locally, so your audio never leaves your machine.

How it works:
– Every fraction of a second, it decides which of 8 speakers is talking, even when two people talk at once
– Speakers are numbered in the order they first show up
– A small speaker memory keeps the labels stable when someone comes back later

Links:
Model: https://huggingface.co/nvidia/Nemotron-3-Diarization
NVIDIA blog: https://huggingface.co/blog/nvidia/nemotron-diarization
Live demo: https://huggingface.co/spaces/nvidia/nemotron-diarization
Parakeet (speech-to-text): https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
NeMo Speech toolkit: https://github.com/NVIDIA-NeMo/Speech
Streaming Sortformer paper: https://arxiv.org/abs/2507.18446
Sortformer paper: https://arxiv.org/abs/2409.06656

#shorts #nvidia #ai #speechrecognition #opensource”}]] Read More Prompt Engineering 

#Promptengineering #AI

By ali

Leave a Reply