Research

Deepgram Nova 3 with Eigen: WER Results Across Four Datasets

TL;DR: In the July 2026 Eigen benchmark, Deepgram Nova 3 WER decreased after Eigen processing on all four evaluated datasets. On the three-hour mixed English/Hindi partner-call set, WER fell from 27.86% to 20.33%, a 7.53-point reduction. The same gain or a latency effect has not been established for another deployment.

When Deepgram Nova 3 makes errors on noisy calls, the audio entering ASR is one variable worth testing. The benchmark measured the same samples before and after Eigen enhancement across four datasets, with Deepgram transcribing in streaming mode.

Results across all four datasets

Dataset

Original WER

With Eigen

Absolute reduction

Relative improvement

LibriSpeech Clean

3.20%

2.61%

0.59 pts

18.4%

VoxPopuli

9.62%

7.60%

2.02 pts

21.0%

Kathbath (Hindi)

17.70%

12.56%

5.14 pts

29.0%

Real World Audio

27.86%

20.33%

7.53 pts

27.0%

Deepgram Nova 3 improved with Eigen in each reported dataset. The largest absolute reduction, 7.53 points, occurred on Real World Audio, an internally sourced three-hour English/Hindi set intended to reflect partner calls. Its sample selection and language mix are not described in the report, so it should not be treated as representative of all production calls.

How Deepgram Nova 3 compared to competitor enhancement on real-world audio

System

WER on Real World Audio

No enhancement

27.86%

Krisp VIVA 2.0 (BVC)

26.90%

ai-coustics Quail VF 2.2L

28.64%

Eigen

20.33%

With Deepgram Nova 3, ai-coustics raised WER by 0.78 points relative to unenhanced audio and Krisp lowered it by 0.96 points. Eigen lowered WER by 7.53 points relative to unenhanced audio, a 6.57-point larger reduction than Krisp on this set.

What this means for a Deepgram-based voice agent

Run the same comparison on consented calls from your own deployment: preserve unprocessed audio, pass a copy through Eigen, and transcribe both paths with identical Deepgram settings. Compare WER and important fields such as names, numbers, and codes by noise cohort. Separately measure added processing and end-to-end response time on your deployment path; the benchmark report provides WER, not latency or conversation-delay measurements.

Related benchmark pages