
The short version
- Traditional phone audio carries roughly 300 to 3,400 Hz and is sampled at 8 kHz, which removes sounds that tell similar consonants apart.
- Compression, packet loss, speakerphones and background noise stack on top of that.
- Models trained on real call audio, plus confirmation of key details, close most of the gap.
A speech recognition model that transcribes a podcast almost perfectly can stumble on a phone call. That's not because phone callers speak differently. It's because the phone network throws away a lot of the sound before the model ever hears it.
If you're evaluating voice AI for a call center, understanding this explains a lot: why demos sound better than production, why some callers get misunderstood more than others, and why testing on your own recordings matters.
The phone line is a narrow pipe
Classic telephone audio was designed in an era when the goal was intelligible conversation over expensive copper lines. The standard "narrowband" voice channel carries frequencies from roughly 300 Hz to 3,400 Hz, and the audio is digitized at 8,000 samples per second (8 kHz). The common G.711 codec used across the public network works this way.
Human speech has energy well above 3,400 Hz. A lot of it lives in consonants: the hiss of "s," the "f" and "th" sounds, the burst at the start of "t" and "p." Those high-frequency cues are what help both people and machines tell "fifteen" from "fifty," "Smith" from "Swith," or "B" from "D" when someone spells their name.
On the phone, those cues are mostly gone. People cope because we use context and ask "sorry, was that B as in boy?" without thinking. Speech recognition has to do the same with much less to work on.
Newer "wideband" or HD voice calls carry more of the spectrum, but you can't count on it. A call from a mobile to a landline, or one routed through older carrier equipment, often falls back to narrowband.
Then the audio gets squeezed
Mobile networks and VoIP systems compress audio with codecs designed to save bandwidth. Each one discards information it judges less important to human listeners. When a call passes through several systems (mobile network, carrier interconnect, a cloud phone system, then into a recording), it may be decoded and re-encoded more than once, and each step adds artifacts.
On top of that, VoIP calls suffer from packet loss and jitter. A lost packet means a short gap in audio. Concealment algorithms fill the gap with a guess, which sounds fine to a human ear and can confuse a model in the middle of a word.
Real callers aren't in a studio
The environment adds its own problems:
- Speakerphones pick up room echo and make voices sound distant.
- Cars add engine and road noise, plus hands-free systems with aggressive noise suppression that can clip the start of words.
- Wind on a mobile outdoors can swamp the microphone completely.
- Background speech, like a TV or other people talking, is particularly hard, because it looks like speech to the model.
- Crosstalk, when the caller and the agent talk at the same time, mixes two voices on the line.
The words that matter are the hardest ones
The most important parts of a business call are usually the least predictable: names, account numbers, addresses, email addresses, dates and times. A language model can often guess a common word from context. It can't guess an order number.
That's why measuring transcription quality only with an overall word error rate can be misleading. A transcript can get 95% of the words right and still miss the one string of digits the call was about.
What actually helps
Models trained on phone audio
The single biggest factor is whether the speech model was trained on audio that sounds like your calls: narrowband, compressed, noisy, with a range of accents. A model trained mostly on clean wideband audio is being asked to work outside what it has seen.
Keeping the audio path clean
Where you control the setup, use the best codec your phone system supports, avoid unnecessary transcoding between systems, and keep jitter buffers and network quality in good shape. Feeding the model audio straight from the phone stream, rather than from a recording that's been compressed again, avoids an extra layer of damage.
Custom vocabulary
If callers regularly say product names, street names or industry terms the model rarely hears, supplying those terms can improve recognition. A clinic's agent should know the names of its doctors.
Designing the conversation for the medium
Good voice agents are built with the phone's limits in mind:
- Confirm critical details. "That's 4-4-7-1, 2-0-8. Is that right?" turns a misheard digit into a quick correction.
- Ask for spelling with examples. "Could you spell the last name? For example, B as in boy."
- Use channels with better fidelity for long strings. For email addresses, offering to send a text with a link is often more reliable than spelling it out.
- Ask one thing at a time. "What's your date of birth?" works better than "Can I get your name, date of birth and ZIP code?"
Handling uncertainty gracefully
When the model's confidence is low, the agent should ask again rather than guess. "Sorry, I didn't catch that. Could you say the order number once more?" is far better than looking up the wrong account.
How to test on your own calls
Take 30 to 50 recordings of real calls, including the difficult ones: speakerphone, cars, accents, noisy environments. Run them through the models you're considering and compare against a careful human transcript. Pay special attention to names and numbers. If you'd like a quick start, upload a few recordings to our free speech-to-text tool and compare the output with what you hear.
Diagnosing bad transcripts
When transcripts from real calls look worse than you expected, work through the likely causes in order:
1. Is the audio itself bad? Listen to the recording. If you struggle to understand it, the model will too. Check whether problems cluster on certain carriers, devices or locations.
2. Is the audio being degraded on the way in? Compare a recording captured directly from the call stream with one that's been exported, compressed and re-uploaded. Extra compression steps hurt.
3. Is the sample rate handled correctly? Feeding 8 kHz audio into a system expecting 16 kHz, or vice versa, without proper conversion causes problems that look like poor recognition.
4. Are both speakers mixed into one channel? Recording each party on a separate channel (stereo, or two streams) makes recognition and speaker attribution much easier than a single mixed track.
5. Are the errors concentrated in particular words? If most errors are names, product terms or numbers, vocabulary hints and confirmation steps will help more than switching models.
6. Is the model suited to phone audio? If audio is fine and errors are spread across ordinary words, the model may simply not be trained for this kind of input.

Mono, stereo and speaker separation
Many call recording systems save calls as a single mono file with both voices mixed together. That's fine for a person listening back, but it makes life harder for transcription. The model has to separate speakers itself (called diarization), and when people talk over each other, both voices end up tangled in one stream.
Where possible, capture each side of the call separately. For a live voice agent this happens naturally: the agent receives the caller's audio on its own stream. For call recording and QA, check whether your phone system offers dual-channel recording. It's one of the cheapest accuracy improvements available.
Accents and dialects
Speech models perform differently across accents, and the differences are larger on degraded phone audio. The honest approach is to test with recordings that reflect your real callers. If you serve a region with a strong local accent, or many callers who speak English as a second language, include plenty of those calls in your evaluation. If a model struggles with a group of your callers, that's a fairness problem as well as an accuracy problem, and it should weigh heavily in your choice.
What a voice agent should do when it can't hear
Even with the best model, some calls will be hard to understand. The agent's behavior in those moments matters more than its average accuracy:
- Ask once, naturally. "Sorry, I didn't catch that. Could you say it again?"
- Narrow the question. If an address keeps failing, ask for the ZIP code first, then the street.
- Offer an alternative. "I can text you a link where you can type it in, if that's easier."
- Hand off gracefully. After a couple of failed attempts, transfer to a person or take a callback number, with a note about what was hard to hear.
Callers forgive an agent that clearly tries and then gets them help. They don't forgive one that guesses wrong and books the wrong appointment.
Frequently asked questions
Why does my transcription software work on meetings but not on calls?
Meeting audio from a laptop or headset is usually wideband and much cleaner. Phone audio removes high frequencies and adds compression, which affects accuracy.
Does HD voice fix the problem?
It helps when both ends and every network in between support it. In practice many calls still end up narrowband, so models need to handle both.
Should I upsample 8 kHz audio before transcription?
Upsampling doesn't restore the missing frequencies. It's better to use a model that accepts 8 kHz audio natively, or that was trained on telephone audio.


