Krisp Voice Isolation Benchmark on HuggingFace
by Nate: Hugging Face
A benchmark for voice isolation just landed on HuggingFace, and the headline number is wild: word error rate went from 23.3% down to 6.2% once Krisp Voice Isolation was in the loop. That's roughly a 73% reduction, and it's the kind of jump that makes you wonder why every STT stack isn't running this by default. The setup: 265 real recordings, not synthetic mixes, run through eleven different STT configurations. The gains weren't uniform though. Shared offices and call-center floors saw the biggest lift, which tracks when you think about how much competing speech is floating around in those rooms. Phone calls barely budged, and that also tracks, because there's way less background chatter to strip out in the first place. Krisp Voice Isolation already supports 200+ voice platforms and has chewed through 10B+ minutes of voice-agent audio, so this isn't some lab experiment. The full dataset is on HuggingFace for reproducibility, and the tested setups include ElevenLabs Scribe v2 streaming, AssemblyAI Universal-3 Pro batch, Soniox STT v5 streaming, Deepgram Nova-3 streaming, Cartesia Ink-Whisper batch, and Nvidia Nemotron 0.6B streaming. Now the numbers are open for anyone to poke at
