// Case study · Voice AI · Text to speech

TTS evaluations via RLHF to improve precision and identify preferred voices

Frontier Indic Voice Lab 10+ languages Aug 2026 onwards
10K+
human preference evaluations
On simulated CX and BFSI calls in the Indian context
10+
languages covered
Indian English, Hindi, Marathi, Tamil, Telugu, Kannada, Punjabi, Bengali, Odia, Malayalam
300+
unknown defects found and treated
Every defect found feeds the next round, so the test set sharpens as fixes ship
80%
less time to identify issues at scale
88 voices screened in under 48 hours, 34 cloned voices the same day
01 · The problem

Two questions no automated score can answer

Model precision

Is the text spoken correctly and accurately for the use case?

  • Pronunciation. Names, places and code-switched words said the way a native speaker says them.
  • Tone. Firm on a collections line, warm on a refund, never the wrong emotion.
  • Suitability. The register the industry and the use case expect.
  • Normalisation. "Rs 1.5L", "12/03" and "S/o" have to become the right words. Word error rate cannot hear when they do not.

Needs: a machine checking every sentence, and native ears confirming what it flags.

Voice preference

Which voice is better, and for what use case?

No score tells you which voice a Tamil or Telugu speaker will trust. And a voice that wins a support call can lose a collections call.

Needs: native speakers judging blind, at scale, on the lines an agent actually says.

02 · Methodology

The loop

A dataset built from real world scenarios runs down two lanes at once. An LLM pipeline checks every sentence; a human pipeline puts native speakers on the high-risk cases. Both land in the same findings, which go back to the lab for the next round.

The RealLoop feedback loop for Indic voice models A dataset from real world scenarios splits into two parallel lanes, an LLM pipeline and a human pipeline for high-risk cases. Both converge on findings, which feed continuous iterations back into the dataset. 01 02 03 04 05 DATASET FROM REALWORLD SCENARIOS LLM PIPELINE HUMAN PIPELINE · HIGH-RISK CASES FINDINGS ITERATION
01Dataset
From real world scenarios
Support and collections lines, scrubbed, in ten languages
80% code-switched
02LLM pipeline
Every sentence checked
A machine reads the text against the audio and flags what went wrong
runs on the full set
03Human pipeline
Native speakers on high-risk cases
Blind pairs, a reason on every verdict, on the lines that carry money or intent
1,479 pairs · 7 judges
04Findings
Both lanes land here
Back to the lab as input: fine-tune, rules, and doubling down on the right voice
300+ findings · 70% P0
05Iteration
Continuous rounds
The same lines, the same panel, so every round makes the system better
a gain must clear zero
03 · Insights

Insights generated

Errors
Pronunciation and accent
500+errors across 10 languages in how words are said

Accent switching mid-sentence, from a regional language into Hindi, and Indian place names read wrong. Only humans hear these.

Andheri West "Andheri waist"60% of clips, 31 Indian English voices
PAN "pain" · KYC "kyz"nearly every voice, Telugu and English
Roman-script Hindi line on the production model17% acceptable, 81% of failures pronunciation
"week" inside a Tamil payment sentencewhole clip judged artificial
వేల, వందల (Telugu thousand, hundred)ending clipped on every amount
ನಮಸ್ಕಾರ (Kannada greeting)flagged on all 28 opening pairs
"two" "thoo", US accent mid-sentenceIndian English voices
Normalisation
40%of sentences carried the wrong normalisation for Indian languages

263 confirmed errors in 690 Indian English sentences. 70% were serious enough that a customer would act on wrong information.

Text that became the wrong words
Rs 1.5L "one point five rupees L"34 of 34 cases
12/03 "twelve over three"46 cases
2026-09-05 "two zero two six zero nine zero five"34 cases
S/o Mohan Lal "s o mohanlal"33 cases
00:00 and 12:00 both "twelve o'clock"16 cases
Lt. Col. R. S. Chauhan "lt nicole r s johan"12 cases
560034 "five lakh sixty thousand" · 24x7 a multiplicationpincode, shorthand
Benchmark

Why native speakers are the benchmark

A benchmark whose raters cannot hear the language measures the wrong thing.

The global English arena ranks the lab 44th of 90, judged by US and UK listeners on English prompts. On real Hindi lines, native speakers had the lab beating ElevenLabs v3 95 to 54 after a tie on demo text, and its voices top the Kannada and Telugu ladders. Native speakers are not a nicer benchmark. They are the only one that hears what the customer hears.

On the global English arena public

Rank strip of the global English text-to-speech arena #1 #45 #90 Cartesia Sonic 3.6 · #1 ElevenLabs v3 Conversational · #10 The lab · #44 The lab · #55
90 models, blind pairs, English only. Judged by people who cannot hear the lab's languages.

Judged by native speakers real

The lab
95
ElevenLabs v3
54
Where the lab's voices sit on native ladders
Hindi · wins head to headKannada · #1Telugu · #1
Decisive pairs won on sentences taken from real support calls. Same voices, same twelve native listeners, blind.

Every voice gets a verdict, not a score

Two of the sixteen Hindi voices kept after screening, shown the way they appear in the lab's dashboard, down to the sentence. Names removed. Reviewer notes are verbatim; "a" and "b" are the blind positions in the pair.

Voice A

female · lab model · hindi
1613fitted elo

The strongest female voice in the pool, firm and even from the first word to the last. She won 83% of her BFSI pairs, and reads slightly less warm on plain conversation.

where to use it

First pick for BFSI and any call where the agent has to sound in control.

Listen
BFSI collections, EMI reminder
0:00 / 0:00
देखिए, पच्चीस तारीख़ तक अगर EMI नहीं जाती तो पाँच सौ रुपये late fee लगेगी, और मैं नहीं चाहती कि आपको वो देना पड़े। क्या आज ही payment हो सकता है?
CX support, refund confirmation
0:00 / 0:00
असुविधा के लिए खेद है, मैं आपका refund अभी process करवाने में आपकी पूरी सहायता करूँगी, कृपया मेरे साथ बने रहें।
screening
45% great · 27% bad
pairs won by call family
BFSI 83% · CX 66% · Conversational 62%
defects, by cause
Normalisation 2

MONSOON2 in the Alphanumeric code line, scored bad, tagged pauses / rhythm.

24x7 in the Written shorthand (24x7) line, scored bad, tagged pronunciation, unnatural tone.

Pronunciation 3

Written shorthand (24x7) line scored bad at screening, tagged pronunciation.

Lost 1 blind pair on pronunciation.

Scripted call, turn 3: "thik hai" (pronunciation).

Tone and emotion 4

Empathy line: reviewers split 1 and 3 of 3, one tagged wrong emotion.

Lost 3 blind pairs on unnatural tone.

Scripted call, angry, frustrated caller: “agent doesn't seem genuinely sorry and gave weird response”

Scripted call, worried, anxious caller: “Agent is not empathetic at all”

Voice B

female · lab model · hindi
1586fitted elo

Warm, quick and natural, the most human-sounding voice when the script loosens up: she won 89% of her conversational pairs and reviewers called her natural even when she was fast. Weaker on firm BFSI asks, where the same warmth reads as soft.

where to use it

Conversational and CX. Not the voice for a collections call.

Listen
Conversational, open-ended turn
0:00 / 0:00
हम्म… सच कहूँ तो मुझे भी ठीक से नहीं पता। लेकिन मैं पता करके बताती हूँ।
CX support, refund confirmation
0:00 / 0:00
असुविधा के लिए खेद है, मैं आपका refund अभी process करवाने में आपकी पूरी सहायता करूँगी, कृपया मेरे साथ बने रहें।
screening
36% great · 18% bad
pairs won by call family
BFSI 54% · CX 58% · Conversational 89%
defects, by cause
Normalisation 1

24x7 in the Written shorthand (24x7) line, scored bad, tagged pronunciation.

Pronunciation 2

Written shorthand (24x7) line scored bad at screening, tagged pronunciation.

Lost 1 blind pair on pronunciation.

Tone and emotion 5

Inform / amount line: one reviewer scored it bad, tagged unnatural tone.

Lost 1 blind pair on unnatural tone.

Scripted call, turn 4: "wasn't apologetic enough" (wrong emotion).

Scripted call, worried, anxious caller: “some gramatic mistakes”

Scripted call, irritated, no intent to pay: “Agent is sighing too much”

Put native speakers on your voices

Send us the lines your agent actually says. We come back with a ladder, the errors that matter, and a retest date.

Contact us