Oossa

Transformers 5.18.0 adds streaming speaker diarization model

Hugging Face’s new release includes Nemotron 3 Diarization, a streaming model that can label up to eight speakers in live audio.

NoteOossaPublished by Oossa: 1 min read

Hugging Face released Transformers v5.18.0. The update adds Nemotron 3 Diarization, an open‑weight model that can tell “who spoke when” in real‑time or offline audio. It works with up to eight speakers and lets developers choose latency from 80 ms to 30.4 seconds.

The release also brings multimodal models NemotronH Omni and HyperCLOVAX Vision V2, plus a range of bug fixes and performance tweaks.

Why it matters

Developers can now add live speaker‑labelled transcription to apps, making meetings and podcasts easier to follow.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.