Table of Contents
Every TTS vendor points to a benchmark score to prove its voices sound natural, but a leaderboard number rarely predicts how that voice performs on your own production traffic. In a 10-annotator listening study across 100 paralanguage queries, the MultiVox benchmark scored professional human recordings 4.5 out of 5 for naturalness against 2.1 for the open-source CosyVoice model.
Scores like these shift with the text, the listeners, and the comparison systems used, so a single figure predicts little about your deployment. A repeatable evaluation protocol starts with production metrics, a test-set recipe, and disclosure questions that separate evidence from advertising.
This article walks through that protocol: the metrics that hold up and the questions worth asking any vendor.
Key Takeaways
Use these findings to design your own benchmark.
- Most papers surveyed by Kirkland et al. didn't cite an ITU standard, though ITU-T P.800.2 rules out comparing MOS values produced by separate experiments.
- Performance varied sharply by content category in EmergentTTS-Eval, a NeurIPS 2025 benchmark.
- Tail latency under production load matters more than a vendor floor figure from a single warm request.
- Published numbers are only trustworthy when they come with versioned comparison targets, listener counts, and P95/P99 figures at your concurrency.
Provider Comparison at a Glance
This table compares streaming protocol, pricing model, HIPAA posture, and deployment options across Deepgram, ElevenLabs, Cartesia, Inworld, and OpenAI. Concurrency limits are the one criterion most vendors, including Deepgram, don't publish as a fixed number.
How to Read the Table
Where a cell says "not disclosed," treat that as a live gap in vendor documentation and a question to ask directly.
Comparison Data
| Criterion | Deepgram | ElevenLabs | Cartesia | Inworld | OpenAI |
|---|---|---|---|---|---|
| Flagship model | Flux TTS | Eleven v3 is cited | sonic-3.5 is cited | Realtime TTS-2 is documented | gpt-4o-mini-tts is documented |
| Streaming protocol | Speak v2 WebSocket | WebSocket and HTTP streaming | WebSocket streaming | WebSocket or WebRTC | HTTP chunked transfer |
| Concurrency limits | Not disclosed | Tier-dependent, not published as a fixed number | Tier-dependent, not published as a fixed number | Tier-dependent, not published as a fixed number | Not disclosed |
| Pricing model | Usage-based | Character-credit subscription tiers | Credit-based subscription tiers | Usage-based, tiered by plan | Token-based, pay-as-you-go |
| HIPAA | HIPAA support and BAA via sales and enterprise agreements | HIPAA-eligible via Enterprise tier and BAA | HIPAA compliant per Cartesia's own safety page | HIPAA and BAA available on select plans | HIPAA support and BAA available |
| Self-hosted deployment | Cloud, self-hosted, VPC | Cloud only | Cloud only | Cloud, plus on-prem via sales | Cloud only |
| Best fit | Real-time voice agents | Broad creative and content voice work | Low-latency conversational agents | Full-stack realtime voice with LLM routing | Teams already built on OpenAI's ecosystem |
- Deepgram
- Flux TTS
- ElevenLabs
- Eleven v3 is cited
- Cartesia
- sonic-3.5 is cited
- Inworld
- Realtime TTS-2 is documented
- OpenAI
- gpt-4o-mini-tts is documented
- Deepgram
- Speak v2 WebSocket
- ElevenLabs
- WebSocket and HTTP streaming
- Cartesia
- WebSocket streaming
- Inworld
- WebSocket or WebRTC
- OpenAI
- HTTP chunked transfer
- Deepgram
- Not disclosed
- ElevenLabs
- Tier-dependent, not published as a fixed number
- Cartesia
- Tier-dependent, not published as a fixed number
- Inworld
- Tier-dependent, not published as a fixed number
- OpenAI
- Not disclosed
- Deepgram
- Usage-based
- ElevenLabs
- Character-credit subscription tiers
- Cartesia
- Credit-based subscription tiers
- Inworld
- Usage-based, tiered by plan
- OpenAI
- Token-based, pay-as-you-go
- Deepgram
- HIPAA support and BAA via sales and enterprise agreements
- ElevenLabs
- HIPAA-eligible via Enterprise tier and BAA
- Cartesia
- HIPAA compliant per Cartesia's own safety page
- Inworld
- HIPAA and BAA available on select plans
- OpenAI
- HIPAA support and BAA available
- Deepgram
- Cloud, self-hosted, VPC
- ElevenLabs
- Cloud only
- Cartesia
- Cloud only
- Inworld
- Cloud, plus on-prem via sales
- OpenAI
- Cloud only
- Deepgram
- Real-time voice agents
- ElevenLabs
- Broad creative and content voice work
- Cartesia
- Low-latency conversational agents
- Inworld
- Full-stack realtime voice with LLM routing
- OpenAI
- Teams already built on OpenAI's ecosystem
Why Vendor Voice Quality Claims Don't Transfer to Production
Published comparisons often reflect sponsor-controlled choices because vendor-published tests usually let the vendor control or select the text, comparison systems, request conditions, and sometimes listener recruitment. Your deployment won't match those choices, so the score may not transfer.
Vendors Control the Demo Conditions
Floor figures are common across the industry. A published response-time spec is almost always a best case: the fastest model, ideal network conditions, and a single warm request, with no view into tail latency under concurrent production load.
Vendors rarely specify which model, tier, or voice produced that number, and switching any one of those variables moves the result. Demo pages rarely disclose which configuration generated the figure they're showing you.
Vendors Control the Text You Hear in a Demo
Vendors choose which content shows up in a demo, and that choice can flatter almost any system. In EmergentTTS-Eval, a 1,645-sample benchmark cited above, a 48.44-percentage-point gap separated the best and worst content categories for one system: gpt-4o-audio-preview won 88.84% of comparisons on emotional speech but only 40.40% on complex pronunciation: emails, phone numbers, URLs, street addresses, and tracking numbers.
A demo built on emotional dialogue shows the system winning 88.84% of comparisons. A demo built on order confirmations or account numbers would show it winning only 40.40% instead. Every open-source model the paper tested underperformed on the complex-pronunciation category, with errors including misread decimals and dropped digits.
Vendors Control the Comparison Methodology
If two labs test identical audio, the scores still diverge. ITU-T P.800.2 states plainly that MOS values from separate experiments aren't meaningfully comparable unless those experiments were specifically designed for that comparison.
Even wording moves the number. Kirkland et al., whose survey came up earlier, found that asking listeners about "quality" produced a significantly higher mean than asking about "naturalness," 4.46 versus 4.25.
The Metrics That Actually Predict Production TTS Performance
Same-experiment listening tests, round-trip intelligibility checks, and tail latency support defensible comparisons. Ask for every measure, then reproduce two.
Mean Opinion Score Under ITU-T Standardization
Only 1 of the 133 papers in Kirkland's survey cited any ITU standard, which tells you how rarely the rules get followed. The rules exist, though. ITU-T P.800 requires a 5-point absolute category rating scale and reference conditions in every experiment so results stay anchored. Its listener-selection criteria cover population, assessment work, and prior test participation.
It also specifies rooms with reverberation under 500 ms and background noise below 30 dBA. ITU-T P.808 extends the method to crowdsourcing: at least 8 raters per stimulus, validated headphone use, environment screening, and one gold-standard and one trapping question per 10 stimuli. When a vendor claims a MOS, ask which requirements the test met.
Word Error Rate as an Intelligibility Check
Score whether words survive synthesis without recruiting a listener. Synthesize your text, transcribe the audio with a fixed ASR model, and compute word error rate against the original.
One caveat has teeth: identical TTS outputs can rank in opposite directions under different ASR families, so run at least two ASR models with separate training lineages. WER catches garbled words. Prosody, including flat delivery, requires a separate measure.
Latency Percentiles Instead of Marketing Averages
When concurrency climbs, tails stretch. NVIDIA's TTS performance tables show P99 first-audio latency of 380.77 ms at 8 parallel streams against a 175.62 ms average, a 2.17× spread on dedicated A100 hardware.
Sherlock's TTFB study reports p95-p99 latency exceeding the median by 3-5× in production voice AI traffic, though these studies time different events and can't be lined up directly. Derive your latency budget from the tail at your peak concurrency, never from a vendor's floor figure or the median.
How to Test a Voice Quality Claim Against Your Own Content
Pull sentences your system will speak before auditioning a voice. Test with production text if you want a result that predicts production behavior.
Building a Test Set From Your Real Use Case
Sample agent turns and IVR menus from production logs. Include confirmation readbacks too. Deepgram's pronunciation-gap analysis is direct about sourcing: build the corpus from your own production data and exclude vendor demo scripts.
Then oversample the classes that fail most: currency, dates, addresses, phone numbers, homographs, and jargon such as drug names. If you need edge cases fast, TTSProof packages 817 curated ones across 39 categories. The set includes decimals, URLs, and medical vocabulary. Equivalence-aware scoring is built in.
Running Your Own Preference Tests
Once you have a shortlist, use blind pairwise comparisons. Camp et al. found that AB tests carry far less variance than separate MOS tests, especially between similar-quality systems, and should be your default for system comparison. Hide vendor identities and randomize playback order.
Wells et al. recommend repeating the same stimulus pair at the start and end as a consistency check. The open-source Microsoft P.808 Toolkit runs these designs on Mechanical Turk and includes reliability checks.
For pairwise comparisons specifically, Cooper's 2025 review recommends a z-test, t-test, or binomial test depending on sample size, reserving the Wilcoxon signed-rank test for MOS-style ordinal ratings.
Scoring Alphanumerics and Domain Terminology
Structured strings fail at rates prose never approaches. Deepgram's TTS evaluation guide reports 10-25% WER on alphanumeric identifiers against 2.8-5.7% for general text, and structured data shows 2-5× higher error rates overall. The same strings break recognition too.
Deepgram's vendor-published Five9 case study reports 2-4× higher alphanumeric accuracy than alternative STT options. Text normalization drives much of the gap. In EmergentTTS-Eval's ablation, an LLM-based normalizer raised the same model from a 51.69% to a 76.74% win rate with no change to the voice. For your own jargon, Deepgram's pronunciation-gap guidance recommends a nightly regression that adds phoneme error rate via the Montreal Forced Aligner.
Questions to Ask Before You Trust a Published Benchmark
Who listened, to what text, against which model versions, and on which date? If a vendor can't answer all four, treat the number as advertising.
- Who ran the test, and under what conditions. Ask for the evaluating organization, rater counts, ratings per condition, standard deviations, and the exact question posed to raters. Ask whether the vendor's number came from a single benchmark run or from repeated runs over time.
- Was it measured at production concurrency. One warm request tells you only the floor, so measure the tail separately. Pin down whether the timestamp means first byte, container header, or first playable audio, and record concurrency along with client and server regions and warm versus cold state.
- Does it name its comparison targets. Replication requires a named model on a dated API call. "A leading competitor" won't do. Deepgram's Flux TTS product page names its targets by model: Inworld 1.5 Max, ElevenLabs v3, and Cartesia sonic-3.5, and frames its results as self-reported testing. It doesn't name the ASR model behind its WER claim or state individual rater counts, so ask for those too.
Evaluating TTS Quality for Your Production Deployment
Run this protocol before signing. Trust a published number only when it came from your content, in a single blinded test, at your concurrency, with a method you could rerun.
Evaluation Checklist
Before you shortlist anyone, run down the following:
- Build a test set from production logs, and oversample structured strings and domain terms.
- Run blinded pairwise listening tests with randomized order and repeated consistency pairs.
- Compute round-trip WER under two ASR families, plus phoneme-level checks on your jargon.
- Load-test at peak concurrency and record P95 and P99, never the average alone.
- Put the disclosure questions to each vendor, and score unanswered ones against them.
Get Started With Deepgram
Named comparison targets and self-reported WER figures are exactly what a buyer should be able to challenge, and Deepgram publishes both. Hold them against the checklist above. The Flux TTS developer overview covers the endpoint if you want to run these tests yourself, and Deepgram's own speech-to-text model works as one of the two ASR families you'll want for round-trip WER testing.
Bring your hardest production text. Exclude demo scripts. Create a free account and put your $200 free credits toward the voice quality test set this article showed you how to build.
FAQ
What Is a Good MOS Score for a Production Text-to-Speech System?
There's no portable threshold. Edlund et al. found that framing the same evaluation task differently produced significantly different MOS responses across four TTS voices, so a raw score means little without knowing how the test was framed. The TTSDS2 study even found four systems scoring above real human recordings in the same test, so treat a standalone 4.x as uninformative without methodology.
Can You Compare MOS Scores Across Different TTS Vendors Directly?
Direct comparisons are invalid across separate experiments unless those experiments were explicitly designed for comparison. A lab score and a crowdsourced score may not be interchangeable even for one product.
How Do You Test Text-to-Speech Voice Quality on Your Own Content?
Synthesize your own production sentences, then score them blind against a rival's output on the same text. If your text streams from an LLM, test that path separately.
What's the Difference Between MOS and Word Error Rate for TTS Evaluation?
One measures listeners' impressions; the other counts surviving words. Synthetic audio also tends to be cleaner and more consistent than real-world speech, which can make it easier for an ASR system to transcribe and inflate intelligibility scores.
What Should You Ask a TTS Vendor Before Trusting Their Published Benchmark?
Get publication rights in writing before you invest in testing. AWS's service terms permit benchmark disclosure if you include replication information and grant reciprocal rights, while Deepgram's MSA restricts competitive use without written consent. Secure an addendum covering permission to run and publish benchmarks. It should also permit naming model versions.









