Voxvencer

Voice AI

What Makes a Text-to-Speech Voice Sound Natural on a Call

Why some synthetic voices sound human on the phone and others sound robotic: prosody, pacing, numbers, latency and how phone audio changes what callers hear.

Radio host speaking into a broadcast microphone
Photo: Tulane Public Relations, Wikimedia Commons (CC BY 2.0)

The short version

  • Naturalness comes mostly from prosody: rhythm, stress and intonation that match the meaning of the sentence.
  • Numbers, dates, addresses and names are where weaker voices give themselves away.
  • Judge voices through a phone speaker on real call scripts, not on demo sentences through headphones.

Modern text-to-speech can be hard to tell from a person in a polished demo. Put the same voice on a phone call reading appointment times and account numbers, and the differences between good and mediocre systems become obvious quickly.

Here's what separates a voice that callers trust from one that makes them reach for the zero key.

Prosody does most of the work

Prosody is the music of speech: rhythm, stress, pitch movement and pauses. It carries meaning. Compare:

  • "I can book you on Thursday." (not another day)
  • "I can book you on Thursday." (neutral)
  • "I can book you on Thursday." (as opposed to someone else)

A natural voice picks the right stress from context, raises pitch at the end of a genuine question, and pauses where a person would breathe or think. A weaker voice reads every sentence with the same shape, which sounds flat and makes long replies hard to follow.

For phone agents, the language model generating the text can help by writing the way people talk: short sentences, one idea at a time, punctuation that marks natural pauses. A sentence that reads well on a page can be a mouthful when spoken.

Pacing and pauses

On a call, the listener can't see your words. They have to hold them in memory. Good voices:

  • Slow down slightly for important details
  • Pause briefly before and after key information
  • Don't rush lists ("I have Tuesday at nine thirty, ... Wednesday at eleven, ... or Thursday at two fifteen.")

Too slow is also a problem. A voice that drags makes the caller impatient and makes the whole call feel long.

Numbers, dates and names

This is where weaker systems give themselves away. Text normalization (deciding how written text should be spoken) is hard:

  • "2:15" should be "two fifteen," not "two colon fifteen"
  • "Suite 700" is "suite seven hundred," but "Flight 700" might be "seven hundred" or "seven oh oh"
  • "1/2" could be a date, a fraction or a measurement
  • "$1,250.50" should be "one thousand two hundred fifty dollars and fifty cents" in some contexts and "twelve fifty, fifty" never
  • Phone numbers and account numbers should be read in groups with small pauses, not as one long number
  • "Dr." is "doctor" before a name and "drive" in an address

Names are their own challenge. A voice that mangles the caller's name or your own staff's names loses trust immediately. Look for systems that let you supply pronunciations for names and terms you use often.

When testing voices, write test scripts full of these cases. Demo sentences rarely include them.

What the phone line does to a voice

Phone audio is narrowband: it cuts off much of the high-frequency detail (we explain this in why phone audio is hard for speech recognition). That affects how synthetic voices sound too:

  • Breathy or very soft voices can turn muddy
  • Voices with a lot of high-frequency "air" lose their character
  • Sibilants ("s," "sh") can become harsh after compression
  • Small artifacts that are inaudible on headphones can become noticeable

A voice designed or tuned for telephone use often sounds slightly less impressive on studio speakers and noticeably better on an actual call. Always evaluate through a phone.

Latency is part of naturalness

A perfectly human-sounding voice that takes two seconds to start talking doesn't feel natural, because people don't leave two-second gaps in conversation. Streaming TTS, which starts playing audio as soon as the first words are ready rather than waiting for the whole reply, makes a large difference. We go deeper on this in voice agent latency.

Handling interruptions

People interrupt. They say "yes, yes" while you're still talking, or cut in with a correction. A natural-feeling agent stops speaking when the caller starts (called barge-in), listens, and responds to what was said. A voice that keeps talking over the caller feels robotic no matter how good it sounds.

Consistency matters more than variety

It's tempting to pick a voice with lots of expressiveness. For business calls, consistency is usually more valuable: the same calm, clear delivery on every call, whether the message is a booking confirmation or bad news about a delay. Overly emotive voices can sound insincere when the content is routine.

How to evaluate a TTS voice

  1. Write a realistic script. Include a greeting, a list of options, a date and time, a phone number, an address, a spelled-out name and an apology.
  2. Generate it with each voice you're considering.
  3. Play it through a phone. Call yourself and play the audio, or use the vendor's phone demo if they have one.
  4. Get several people to listen. Ask which voice they'd be happiest to hear from your business, and where each one stumbled.
  5. Check the details: were numbers read correctly, were pauses in the right places, was anything mispronounced?

You can run this test on our voice with the free text-to-speech tool: paste your script, generate the audio and play it through your phone.

Should the voice sound exactly human?

It should sound clear, warm and easy to listen to. Whether it's indistinguishable from a person matters less than you'd think, and the agent should say it's an automated assistant anyway. Callers judge the experience on whether they got what they needed without effort.

Writing text that sounds good spoken

The voice can only work with the words it's given. For phone agents, the text generation step should follow a few rules that make speech easier to follow:

Front-load the answer. "Your order shipped on Tuesday" before "it's with the carrier and should arrive Thursday."

One idea per sentence. Long sentences with multiple clauses are hard to follow by ear, because listeners can't glance back.

Avoid parentheses and abbreviations. "Approx." and "e.g." are read awkwardly or wrongly. Write out what you'd say.

Use contractions. "You're booked" sounds natural. "You are booked" sounds stiff.

Give numbers context. "Your balance is forty-two dollars and fifty cents" is clearer than "42.50."

Signal structure out loud. Instead of bullet points, say "There are two options. The first is... The second is..."

These rules belong in the agent's instructions, so every reply comes out ready to be spoken.

Woman with headphones speaking into a studio microphone
Photo: Kristin Hardwick, StockSnap (CC0)

Choosing a voice for your brand

A few practical considerations when picking a voice:

  • Clarity over character. A distinctive voice is memorable, but on a noisy phone line clarity wins.
  • Match your audience. A pediatric clinic and a debt collection line probably want different tones.
  • Consider speed. Some voices speak noticeably faster than others. Older callers and non-native speakers often prefer a slightly slower pace.
  • Keep it consistent. Use the same voice across your agents, greetings and on-hold messages, so callers recognize your business.
  • Test with real callers. Play a few options to customers or staff and ask which they'd trust.

Pronunciation lists

Every business has words a synthetic voice will get wrong: staff names, street names, product names, local place names. Keep a pronunciation list and test it whenever you add new entries. Typical entries might look like:

Written Should sound like
Dr. Nguyen "doctor win"
Siobhan "shiv-awn"
Kalispell "KAL-ih-spell"
SKU-4471 "S K U, four four seven one"
HVAC "H-vac" or "H V A C", whichever your team says

Depending on the TTS system, fixes might be phonetic spellings in the text, pronunciation dictionaries, or markup. The simplest fix, rewriting the text, often works fine.

Measuring how a voice performs

Beyond listening tests, a few signals from real calls tell you whether the voice is working:

  • Repeat requests: how often callers say "sorry?" or "what was that?" after the agent speaks
  • Interruptions: frequent barge-ins can mean replies are too long or too slow
  • Hang-ups after the greeting: a sharp drop-off right after the agent's first sentence can point to a voice that puts people off
  • Survey comments mentioning the voice, good or bad

Track these over time, especially after changing voices or speaking rates.

Frequently asked questions

Can I use a cloned version of my own voice?

Technically, voice cloning is possible with many systems. For outbound calls in the US, AI-generated and cloned voices are treated as artificial voices under the TCPA, so consent rules apply. Make sure you have rights to any voice you clone.

Male or female voice?

Pick the voice your callers respond to best and that fits your brand. Test both on real scripts if you're unsure.

How do I fix a word the voice mispronounces?

Many TTS systems support pronunciation hints or phonetic spellings. As a fallback, rewrite the text so the default pronunciation is correct, for example "Dee Ann" instead of "DeAnn."

Written by the Voxvencer editorial team. We build and run AI voice agents for call centers and small businesses, and we write about what we see on real phone lines. Questions or corrections: info@voxvencer.com.

Keep reading