Hugging Face announced an Open TTS Leaderboard that rates open‑source text‑to‑speech models with automatic metrics. It measures intelligibility via word error rate, speed with real‑time factor and time‑to‑first‑audio, and speaker similarity using cosine similarity of embeddings. The system can evaluate a model in a few hours, compared with the weeks needed for human‑preference arenas. The leaderboard also lets users listen to outputs and give feedback.
Why it matters
It gives developers a quick, reproducible way to compare open‑source speech models, helping them choose faster or more accurate options for apps.