Research
Cartesia ink-2 with Eigen: WER Results Across Three Datasets
TL;DR: Cartesia ink-2 had the lowest unenhanced WER of the four ASRs on the benchmark's mixed English/Hindi partner-call set. WER decreased from 11.05% to 9.38% after Eigen processing, a 1.67-point or 15.1% relative reduction. The evaluation measures recognition accuracy, not whether that gain changes Cartesia's end-to-end speed.
The benchmark tested whether upstream enhancement changed Cartesia ink-2 WER on the three datasets for which this pairing has published results.
Results across datasets
Dataset | Original WER | With Eigen | Absolute reduction | Relative improvement |
|---|---|---|---|---|
LibriSpeech Clean | 2.00% | 1.54% | 0.46 pts | 23.0% |
VoxPopuli | 7.90% | 6.84% | 1.06 pts | 13.4% |
Real World Audio | 11.05% | 9.38% | 1.67 pts | 15.1% |
Cartesia ink-2 was not evaluated on the Hindi Kathbath set in this report. Across the three reported datasets, Eigen improved WER in every case.
How Cartesia compared with competitor enhancement on real-world audio
System | WER on Real World Audio |
|---|---|
No enhancement | 11.05% |
Krisp VIVA 2.0 (BVC) | 11.63% |
ai-coustics Quail VF 2.2L | 16.72% |
Eigen | 9.38% |
Cartesia ink-2 had the lowest unenhanced WER of the four ASRs on Real World Audio (11.05%). Krisp raised WER to 11.63% and ai-coustics to 16.72%, while Eigen reduced it to 9.38%. The smaller absolute Eigen gain here than for the other ASRs is an observation, not evidence that Cartesia has a fixed accuracy ceiling.
What this means for a Cartesia-based voice agent
If Cartesia is part of a latency-sensitive pipeline, compare the 1.67-point WER gain with separately measured processing time, time to first transcript, and end-to-end response time. The benchmark does not report those timing measures. Also test on your own calls: the Real World Audio set is an internally sourced three-hour English/Hindi subset, and its reference transcripts were manually audited from another ASR's output.
Related benchmark pages
Eigen, Krisp, and ai-coustics on the same Real World Audio set
How to evaluate real-time noise cancellation in a Python voice agent