Research

Eigen on Noisy Hindi Audio: Kathbath WER with Deepgram Nova 3 and Soniox v5

TL;DR: On 1,000 curated noisy Hindi Kathbath clips, WER decreased after Eigen processing from 17.70% to 12.56% with Deepgram Nova 3 and from 22.90% to 18.10% with Soniox v5. These results cover two Hindi ASR pairings, not other languages or complete multilingual voice agents.

The benchmark included a Hindi evaluation using 1,000 curated noisy Kathbath clips. It provides a separate non-English check alongside the English public datasets and mixed English/Hindi partner-call set.

Why this matters

If an enhancement layer is intended for more than one language, evaluate it on each target language with the actual recognizer and audio conditions. A gain on English read speech does not answer whether the same processing helps noisy Hindi clips. Kathbath supplies one such test, but it does not establish behavior across all languages, accents, or code-switching calls.

Results

ASR Model

Original WER

With Eigen

Absolute reduction

Relative improvement

Deepgram Nova 3

17.70%

12.56%

5.14 pts

29.0%

Soniox v5

22.90%

18.10%

4.80 pts

21.0%

AssemblyAI Universal 3 Pro and Cartesia ink-2 were not evaluated for Hindi in this report, so the table has no Kathbath results for those pairings.

For the two evaluated ASR models, the relative WER reductions were 29.0% with Deepgram and 21.0% with Soniox. These Hindi-only figures should not be compared as though the Real World Audio results came from an English-only set: that three-hour partner-call subset contains both English and Hindi, and the report does not separate its scores by language.

Horizontal comparison of Hindi Kathbath word error rate before and after Eigen for Deepgram Nova 3 and Soniox v5; lower is better.

Figure 1. WER on 1,000 curated noisy Hindi Kathbath clips; lower is better. Source: Arctan Eigen Benchmark Report, July 2026.

What this suggests for multilingual deployments

Both evaluated Hindi pairings improved WER with Eigen on this curated set, so the report's gains are not limited to its English public datasets. The WER scores do not reveal the acoustic or linguistic mechanism behind the improvement. A multilingual deployment should test each language, language mix, and ASR path it actually uses.

Limitations to note

The report covers English and Hindi only, with Kathbath results for two ASRs. It does not provide confidence intervals, a code-switching breakdown, or outcomes for other languages. The Real World Audio references were manually audited from AssemblyAI output; Kathbath is a separate public dataset. For another language or recognizer, treat this as a reason to run a local test, not a predicted WER reduction.

Related benchmark pages