Research
Eigen, Krisp, and ai-coustics: WER on the Real World Audio Set
TL;DR: On the benchmark’s three-hour English/Hindi call set sourced from Arctan design partners, Eigen had the lowest WER with each of four streaming ASRs. Krisp improved two ASR pairings and worsened two; ai-coustics worsened all four relative to unenhanced audio. The comparison applies to the named versions and test conditions, not every version or deployment of those products.
The useful question is whether an enhancement system improves recognition on the calls and ASR model a team actually uses. The July 2026 Eigen benchmark compared Eigen with Krisp VIVA 2.0 (BVC) and ai-coustics Quail VF 2.2L on the same three-hour Real World Audio set. The audio comes from Arctan design partners and contains English and Hindi; a per-language breakdown is not available for this set.
The setup
The report describes paired processing: the same input samples were evaluated without enhancement and after each enhancement system, then transcribed in streaming mode by Deepgram Nova 3, AssemblyAI Universal 3 Pro, Soniox v5, and Cartesia ink-2. Word error rate (WER) is the count of substitutions, deletions, and insertions divided by reference words; lower is better. The real-world reference transcripts came from manually audited AssemblyAI Universal 3 Pro output. The report does not publish all processing settings or uncertainty estimates, so the table supports a comparison within this run, not a universal ranking.
The results
ASR Model | Original (no enhancement) | Krisp VIVA 2.0 (BVC) | ai-coustics Quail VF 2.2L | Eigen |
|---|---|---|---|---|
Deepgram Nova 3 | 27.86% | 26.90% | 28.64% | 20.33% |
AssemblyAI Universal 3 Pro | 22.68% | 27.10% | 27.86% | 16.91% |
Soniox v5 | 12.81% | 12.02% | 15.64% | 7.14% |
Cartesia ink-2 | 11.05% | 11.63% | 16.72% | 9.38% |
Average across all four ASR models | 18.6% | 19.4% | 22.2% | 13.4% |
What stands out
The unweighted mean across the four ASR WERs was 18.6% without enhancement, 19.4% with Krisp, 22.2% with ai-coustics, and 13.4% with Eigen. Krisp improved Deepgram (27.86% to 26.90%) and Soniox (12.81% to 12.02%) but worsened AssemblyAI (22.68% to 27.10%) and Cartesia (11.05% to 11.63%). ai-coustics had higher WER than the unenhanced baseline in all four rows. The mean is not a pooled error rate across all words.
Eigen reduced WER relative to the unenhanced baseline in all four rows. Its 13.4% unweighted mean is about 28% lower than the 18.6% unenhanced mean. An average alone hides model-specific differences, so a team should begin with the row for its own ASR.
What can WER tell us about the cause?
The table measures recognition error, not why an error changed. Enhancement can sometimes remove useful speech cues or introduce artifacts, but the benchmark does not include audio-level error analysis that would attribute these results to a specific mechanism. Under the evaluated conditions, Eigen had lower WER than the other three input conditions for each ASR tested.
Why this matters for evaluation, not just marketing
For a deployment decision, compare each candidate with an unenhanced baseline on the same consented calls and ASR configuration. Include quiet and noisy cohorts, and inspect high-value words as well as WER. The benchmark evidence page describes the dataset mix, scoring approach, reported results, and what is publicly available. Raw partner recordings and the complete run configuration are not published.
Continue with the relevant ASR
Start with the benchmark evidence page for the complete dataset and methodology scope.
Read the measured row in context for Deepgram Nova 3, AssemblyAI Universal 3 Pro, Soniox v5, or Cartesia ink-2.