Research

Cartesia ink-2 with Eigen: WER Results Across Three Datasets

TL;DR: Cartesia ink-2 had the lowest unenhanced WER of the four ASRs on the benchmark's mixed English/Hindi partner-call set. WER decreased from 11.05% to 9.38% after Eigen processing, a 1.67-point or 15.1% relative reduction. The evaluation measures recognition accuracy, not whether that gain changes Cartesia's end-to-end speed.

The benchmark tested whether upstream enhancement changed Cartesia ink-2 WER on the three datasets for which this pairing has published results.

Results across datasets

Dataset

Original WER

With Eigen

Absolute reduction

Relative improvement

LibriSpeech Clean

2.00%

1.54%

0.46 pts

23.0%

VoxPopuli

7.90%

6.84%

1.06 pts

13.4%

Real World Audio

11.05%

9.38%

1.67 pts

15.1%

Cartesia ink-2 was not evaluated on the Hindi Kathbath set in this report. Across the three reported datasets, Eigen improved WER in every case.

How Cartesia compared with competitor enhancement on real-world audio

System

WER on Real World Audio

No enhancement

11.05%

Krisp VIVA 2.0 (BVC)

11.63%

ai-coustics Quail VF 2.2L

16.72%

Eigen

9.38%

Cartesia ink-2 had the lowest unenhanced WER of the four ASRs on Real World Audio (11.05%). Krisp raised WER to 11.63% and ai-coustics to 16.72%, while Eigen reduced it to 9.38%. The smaller absolute Eigen gain here than for the other ASRs is an observation, not evidence that Cartesia has a fixed accuracy ceiling.

What this means for a Cartesia-based voice agent

If Cartesia is part of a latency-sensitive pipeline, compare the 1.67-point WER gain with separately measured processing time, time to first transcript, and end-to-end response time. The benchmark does not report those timing measures. Also test on your own calls: the Real World Audio set is an internally sourced three-hour English/Hindi subset, and its reference transcripts were manually audited from another ASR's output.

Related benchmark pages