Research
Eigen Benchmark: WER Results for Four ASRs Across Selected Datasets
TL;DR: In the July 2026 Eigen benchmark, word error rate (WER) decreased for every reported ASR/dataset pairing after Eigen processing. The largest relative reduction was 44.3% for Soniox v5 on a three-hour English/Hindi call set sourced from Arctan design partners. The evaluation measured transcription accuracy, not agent task success or response time.
A degraded transcript can contribute to a voice agent failure before the language model receives the caller's words. Background noise, overlapping speech, reverberation, and weak capture are possible causes. This benchmark measures how Eigen, an audio enhancement layer from Arctan, changes WER for four streaming ASR systems on selected public and partner-sourced datasets. It does not identify the cause of every voice agent error.
What was measured
The benchmark treats Eigen as a pre-ASR layer: input audio is enhanced first, and the resulting audio is streamed into speech recognition systems for transcription. Two questions were tested:
Does Eigen reduce WER for four selected ASR models, compared with unenhanced audio, across clean, structured, noisy, and partner-call conditions?
How does Eigen compare with two other enhancement systems on the internally sourced Real World Audio set?
Methodology
The evaluation design was adapted from the Open ASR Leaderboard paper, including its approach to dataset handling, transcript normalization, and WER scoring. The paper's WER values were not reused. The published method does not provide enough run detail to establish that every step of the paper's protocol was reproduced.
For each evaluated sample, the original audio was retained as a baseline and compared with Eigen-enhanced audio. Transcription used streaming ASR, with transcripts normalized before WER scoring. Paired results are available for all four ASRs on LibriSpeech Clean, VoxPopuli, and Real World Audio; only Deepgram and Soniox have reported Kathbath results. The published competitor comparison is limited to Real World Audio. Exact normalization rules, complete invocation settings, a per-language breakdown of the mixed call set, and uncertainty estimates are not published here. The results therefore should not be read as independently reproducible or statistically significant.
WER is calculated as (Substitutions + Deletions + Insertions) / Reference Words. Lower is better.
Scope: Deepgram Nova 3, AssemblyAI Universal 3 Pro, Soniox v5, and Cartesia ink-2 in streaming transcription mode. The English/Hindi coverage differs by dataset and ASR, as shown below.
Datasets:
LibriSpeech Clean (test-clean split): read audiobook speech under clean recording conditions
VoxPopuli (5-hour subset): formal European Parliament speech, structured and professionally captured
Kathbath (Hindi) (1,000 curated noisy clips): challenging acoustic conditions in Hindi
Real World Audio (3-hour English/Hindi subset): sourced from Arctan design partner companies, reflecting conversational pacing, background noise, interruptions, and natural speech variation. Ground truth for this dataset was established through manual audit of AssemblyAI Universal 3 Pro transcripts.
Results: original audio vs. Eigen-enhanced audio
Dataset | ASR Model | Original WER | With Eigen | Absolute reduction | Relative improvement |
|---|---|---|---|---|---|
LibriSpeech Clean | Deepgram Nova 3 | 3.20% | 2.61% | 0.59 pts | 18.4% |
LibriSpeech Clean | AssemblyAI Universal 3 Pro | 1.60% | 1.40% | 0.20 pts | 12.5% |
LibriSpeech Clean | Soniox v5 | 2.16% | 1.56% | 0.60 pts | 27.8% |
LibriSpeech Clean | Cartesia ink-2 | 2.00% | 1.54% | 0.46 pts | 23.0% |
VoxPopuli | Deepgram Nova 3 | 9.62% | 7.60% | 2.02 pts | 21.0% |
VoxPopuli | AssemblyAI Universal 3 Pro | 7.40% | 6.50% | 0.90 pts | 12.2% |
VoxPopuli | Soniox v5 | 7.21% | 6.62% | 0.59 pts | 8.2% |
VoxPopuli | Cartesia ink-2 | 7.90% | 6.84% | 1.06 pts | 13.4% |
Kathbath (Hindi) | Deepgram Nova 3 | 17.70% | 12.56% | 5.14 pts | 29.0% |
Kathbath (Hindi) | Soniox v5 | 22.90% | 18.10% | 4.80 pts | 21.0% |
Real World Audio | Deepgram Nova 3 | 27.86% | 20.33% | 7.53 pts | 27.0% |
Real World Audio | AssemblyAI Universal 3 Pro | 22.68% | 16.91% | 5.77 pts | 25.4% |
Real World Audio | Soniox v5 | 12.81% | 7.14% | 5.67 pts | 44.3% |
Real World Audio | Cartesia ink-2 | 11.05% | 9.38% | 1.67 pts | 15.1% |
AssemblyAI Universal 3 Pro and Cartesia ink-2 were not evaluated for Hindi in this report, so no Kathbath results are reported for those pairings.
Across every dataset and ASR model where a paired result exists, Eigen reduced WER. The gains are modest on already-clean speech (LibriSpeech Clean) and become substantially larger on noisy and real-world audio, with the single largest relative improvement, 44.3 percent, occurring on Soniox v5 with Real World Audio.
For model-specific interpretation, see the results for Deepgram Nova 3, AssemblyAI Universal 3 Pro, Soniox v5, and Cartesia ink-2. The Hindi Kathbath analysis covers the two ASRs evaluated on that dataset.
Results: Eigen vs. named competitors
On the Real World Audio dataset, Eigen was also compared directly against two competitor enhancement systems, Krisp VIVA 2.0 (BVC) and ai-coustics Quail VF 2.2L, using the same ASR models and audio.
ASR Model | Original WER | Krisp VIVA 2.0 (BVC) | ai-coustics Quail VF 2.2L | Eigen |
|---|---|---|---|---|
Deepgram Nova 3 | 27.86% | 26.90% | 28.64% | 20.33% |
AssemblyAI Universal 3 Pro | 22.68% | 27.10% | 27.86% | 16.91% |
Soniox v5 | 12.81% | 12.02% | 15.64% | 7.14% |
Cartesia ink-2 | 11.05% | 11.63% | 16.72% | 9.38% |
Average | 18.6% | 19.4% | 22.2% | 13.4% |
The average row is an unweighted mean of the four ASR WERs, not a pooled WER over all reference words. On this call set, the means were 13.4% for Eigen, 18.6% for unenhanced audio, 19.4% for Krisp VIVA 2.0 (BVC), and 22.2% for ai-coustics Quail VF 2.2L. Krisp improved two individual ASR pairings and worsened two, a distinction the average hides.
The full named-system comparison breaks down where each enhancement condition improved or worsened the unenhanced result.
How should these results be interpreted?
The evaluation covers English and Hindi only. The three-hour Real World Audio subset comes from Arctan design partners; its sample-selection criteria and English/Hindi distribution are not published. Its reference transcripts were manually audited from AssemblyAI Universal 3 Pro output, so reference errors or model-related bias may remain. The ASR and competitor model versions are named, but complete run settings, per-sample scores, confidence intervals, and an independent audit are not included. The benchmark measures WER, not perceived audio quality, added latency, false interruptions, agent task completion, or saved LLM calls. Teams should test the same comparisons on consented recordings from their own deployment before generalizing these numbers.
Conclusion
Across the reported pairings, Eigen reduced WER. The largest absolute reductions appeared on the partner-sourced call set, where Eigen had lower WER than the unenhanced and two named competitor conditions for each of four ASRs. These results support testing Eigen on similar calls; they do not establish the size of the benefit for other deployments or downstream agent behavior.
Planned future work includes evaluation on additional public datasets (AMI, Earnings22, GigaSpeech), more languages, real-time factor (RTFx) analysis, and coverage of additional STT providers, including open-source options.
Related guidance
Why background noise breaks speech-to-text accuracy explains how to diagnose the input before changing the ASR stack.
How to add real-time noise cancellation to a Python voice agent turns the evaluation pattern into an integration checklist.