// Case studies

What native speakers and trained reviewers found

Two programs running in production. Each one carries the real numbers, the defects found, and what the team changed because of them.

Voice AI · Text to speech Aug 2026 onwards

TTS evaluations via RLHF to improve precision and identify preferred voices

Improving model precision and selecting preferred voices, on simulated CX and BFSI calls in the Indian context. Pronunciation, tone, normalisation and voice preference, judged blind by native speakers.

10K+
human preference evaluations
10+
languages covered
300+
unknown defects found and treated
For a frontier Indic voice lab, 10+ languages
Read the case study →
Voice AI · Agent evaluation 2026

Making agent performance measurable

Over and above technical metrics, to evaluate whether an agent can hold a call. One subjective "was this a good call" broken into scores across the STT, LLM, TTS and extraction pipeline.

2.5×
improvement in failure identification
10%
of all call volume reviewed
15+
metrics across the pipeline
For a scaled voice AI orchestrator, 20+ agents across 8+ languages and 5+ industries
Read the case study →

Put human judgment on your voice AI

Send us the lines and calls your agent actually handles. We come back with the numbers, the failures that matter, and a retest date.

Contact us