Research
Eigen on Noisy Hindi Audio: Kathbath WER with Deepgram Nova 3 and Soniox v5
TL;DR: On 1,000 curated noisy Hindi Kathbath clips, WER decreased after Eigen processing from 17.70% to 12.56% with Deepgram Nova 3 and from 22.90% to 18.10% with Soniox v5. These results cover two Hindi ASR pairings, not other languages or complete multilingual voice agents.
The benchmark included a Hindi evaluation using 1,000 curated noisy Kathbath clips. It provides a separate non-English check alongside the English public datasets and mixed English/Hindi partner-call set.
Why this matters
If an enhancement layer is intended for more than one language, evaluate it on each target language with the actual recognizer and audio conditions. A gain on English read speech does not answer whether the same processing helps noisy Hindi clips. Kathbath supplies one such test, but it does not establish behavior across all languages, accents, or code-switching calls.
Results
ASR Model | Original WER | With Eigen | Absolute reduction | Relative improvement |
|---|---|---|---|---|
Deepgram Nova 3 | 17.70% | 12.56% | 5.14 pts | 29.0% |
Soniox v5 | 22.90% | 18.10% | 4.80 pts | 21.0% |
AssemblyAI Universal 3 Pro and Cartesia ink-2 were not evaluated for Hindi in this report, so the table has no Kathbath results for those pairings.
For the two evaluated ASR models, the relative WER reductions were 29.0% with Deepgram and 21.0% with Soniox. These Hindi-only figures should not be compared as though the Real World Audio results came from an English-only set: that three-hour partner-call subset contains both English and Hindi, and the report does not separate its scores by language.
Figure 1. WER on 1,000 curated noisy Hindi Kathbath clips; lower is better. Source: Arctan Eigen Benchmark Report, July 2026.
What this suggests for multilingual deployments
Both evaluated Hindi pairings improved WER with Eigen on this curated set, so the report's gains are not limited to its English public datasets. The WER scores do not reveal the acoustic or linguistic mechanism behind the improvement. A multilingual deployment should test each language, language mix, and ASR path it actually uses.
Limitations to note
The report covers English and Hindi only, with Kathbath results for two ASRs. It does not provide confidence intervals, a code-switching breakdown, or outcomes for other languages. The Real World Audio references were manually audited from AssemblyAI output; Kathbath is a separate public dataset. For another language or recognizer, treat this as a reason to run a local test, not a predicted WER reduction.