Research

AssemblyAI Universal 3 Pro with Eigen: WER Results Across Three Datasets

TL;DR: In the benchmark's AssemblyAI Universal 3 Pro pairings, WER decreased after Eigen processing on LibriSpeech Clean, VoxPopuli, and a three-hour partner-call set. On the mixed English/Hindi call set, WER fell from 22.68% to 16.91%; the two named competitor enhancement conditions had higher WER than unenhanced audio. Hindi-only Kathbath results were not reported for AssemblyAI.

The benchmark tested whether upstream enhancement changed AssemblyAI Universal 3 Pro transcription accuracy on the three datasets for which this pairing has published results.

Results across datasets

Dataset

Original WER

With Eigen

Absolute reduction

Relative improvement

LibriSpeech Clean

1.60%

1.40%

0.20 pts

12.5%

VoxPopuli

7.40%

6.50%

0.90 pts

12.2%

Real World Audio

22.68%

16.91%

5.77 pts

25.4%

AssemblyAI Universal 3 Pro was not evaluated on the Hindi Kathbath set in this report. Across the three reported datasets, WER improved with Eigen, with the largest absolute gain, 5.77 points, on Real World Audio.

How AssemblyAI compared with competitor enhancement on real-world audio

System

WER on Real World Audio

No enhancement

22.68%

Krisp VIVA 2.0 (BVC)

27.10%

ai-coustics Quail VF 2.2L

27.86%

Eigen

16.91%

On this Real World Audio set, Krisp raised AssemblyAI WER by 4.42 points and ai-coustics by 5.18 points compared with the 22.68% unenhanced baseline. Eigen lowered WER by 5.77 points. These are results for the named product versions in one internally run comparison; they do not describe every available version or audio condition.

What this means for an AssemblyAI-based voice agent

This pairing illustrates why adding a processor should be tested against unenhanced audio rather than assumed to help. If AssemblyAI is your recognizer, replay consented calls through each candidate with the same ASR settings and examine the error types, not only the aggregate WER. The report's references for the partner-call set were manually audited from AssemblyAI output, a potential source of reference bias to consider when interpreting this particular comparison.

Related benchmark pages