Research
Deepgram Nova 3 with Eigen: WER Results Across Four Datasets
TL;DR: In the July 2026 Eigen benchmark, Deepgram Nova 3 WER decreased after Eigen processing on all four evaluated datasets. On the three-hour mixed English/Hindi partner-call set, WER fell from 27.86% to 20.33%, a 7.53-point reduction. The same gain or a latency effect has not been established for another deployment.
When Deepgram Nova 3 makes errors on noisy calls, the audio entering ASR is one variable worth testing. The benchmark measured the same samples before and after Eigen enhancement across four datasets, with Deepgram transcribing in streaming mode.
Results across all four datasets
Dataset | Original WER | With Eigen | Absolute reduction | Relative improvement |
|---|---|---|---|---|
LibriSpeech Clean | 3.20% | 2.61% | 0.59 pts | 18.4% |
VoxPopuli | 9.62% | 7.60% | 2.02 pts | 21.0% |
Kathbath (Hindi) | 17.70% | 12.56% | 5.14 pts | 29.0% |
Real World Audio | 27.86% | 20.33% | 7.53 pts | 27.0% |
Deepgram Nova 3 improved with Eigen in each reported dataset. The largest absolute reduction, 7.53 points, occurred on Real World Audio, an internally sourced three-hour English/Hindi set intended to reflect partner calls. Its sample selection and language mix are not described in the report, so it should not be treated as representative of all production calls.
How Deepgram Nova 3 compared to competitor enhancement on real-world audio
System | WER on Real World Audio |
|---|---|
No enhancement | 27.86% |
Krisp VIVA 2.0 (BVC) | 26.90% |
ai-coustics Quail VF 2.2L | 28.64% |
Eigen | 20.33% |
With Deepgram Nova 3, ai-coustics raised WER by 0.78 points relative to unenhanced audio and Krisp lowered it by 0.96 points. Eigen lowered WER by 7.53 points relative to unenhanced audio, a 6.57-point larger reduction than Krisp on this set.
What this means for a Deepgram-based voice agent
Run the same comparison on consented calls from your own deployment: preserve unprocessed audio, pass a copy through Eigen, and transcribe both paths with identical Deepgram settings. Compare WER and important fields such as names, numbers, and codes by noise cohort. Separately measure added processing and end-to-end response time on your deployment path; the benchmark report provides WER, not latency or conversation-delay measurements.
Related benchmark pages
Eigen, Krisp, and ai-coustics on the same Real World Audio set