Two programs running in production. Each one carries the real numbers, the defects found, and what the team changed because of them.
Improving model precision and selecting preferred voices, on simulated CX and BFSI calls in the Indian context. Pronunciation, tone, normalisation and voice preference, judged blind by native speakers.
Over and above technical metrics, to evaluate whether an agent can hold a call. One subjective "was this a good call" broken into scores across the STT, LLM, TTS and extraction pipeline.
Send us the lines and calls your agent actually handles. We come back with the numbers, the failures that matter, and a retest date.