Voxvencer

Voice AI

Voice Agent Latency: Why One Second of Silence Feels Long

Where the delay in an AI phone call comes from, what targets make sense, and how streaming, turn detection and barge-in make voice agents feel responsive.

Data center with rows of illuminated server racks
Photo: rawpixel (CC0)

The short version

  • The pause before an agent replies is the sum of turn detection, speech recognition, the language model, speech synthesis and the network.
  • ITU-T G.114 treats under 150 ms of one-way network delay as acceptable for most uses, but the bigger delays come from the AI pipeline.
  • Streaming every stage and detecting end of speech quickly matter more than any single fast model.

In a normal conversation, people take turns quickly. Gaps between speakers are usually short, often well under a second, and long silences carry meaning: hesitation, confusion, a dropped line. Put an AI agent on the phone that takes two or three seconds to answer every question, and callers start saying "hello?" into the silence, talking over the agent, or hanging up.

Latency, the delay between the caller finishing and the agent starting to reply, is one of the biggest factors in whether a voice agent feels natural. Here's where it comes from and how to reduce it.

Where the delay comes from

Every turn passes through a chain of steps. The delays add up:

Step What happens
Network in Caller's audio travels from their phone to the agent
Turn detection Deciding the caller has finished speaking
Speech to text Final transcript of what they said
Language model Deciding what to say, and calling tools like a calendar lookup
Text to speech Generating the first audio of the reply
Network out Audio travels back to the caller

Teams often focus on the language model because it's the most visible. In practice, turn detection and the hand-offs between stages can contribute as much or more.

The network part

For telephone networks, the ITU-T recommendation G.114 says one-way transmission delay of up to 150 ms is acceptable for most user applications, and that delays above 400 ms are generally unacceptable for network planning. That's for the network alone, before any processing.

You can't control the caller's carrier, but you can control where your agent runs. Hosting the agent close to where calls enter your system, and avoiding unnecessary hops between cloud regions or vendors, keeps this part small.

Turn detection: the hidden delay

How does the agent know the caller has finished? The simplest method waits for a fixed period of silence, say 700 ms or a full second. Every one of those milliseconds is added to every reply.

Shorter silence thresholds respond faster but interrupt people who pause mid-sentence ("my account number is... let me find it... 4471"). Longer ones feel sluggish.

Better approaches combine silence with other signals: whether the sentence sounds complete, whether the caller's pitch dropped at the end, and context (if the agent just asked for a phone number, ten digits is probably the whole answer). This lets the agent reply quickly after complete thoughts and wait patiently during pauses.

Streaming everything

The biggest single improvement is streaming, so each stage starts work before the previous one finishes:

  • Streaming speech recognition produces a transcript while the caller is still talking, so the final text is ready almost immediately when they stop.
  • Streaming language model output produces the reply word by word, so the first sentence is ready long before the full reply.
  • Streaming text to speech starts playing audio as soon as the first phrase arrives.

Without streaming, delays are added end to end. With streaming, they overlap, and the caller hears the start of the reply much sooner.

Tool calls and slow systems

When the agent needs to look something up (an order status, calendar availability), it has to wait for that system. If your order API takes two seconds, the agent waits two seconds.

Options:

  • Fill the gap naturally. "Let me check that for you" covers a short lookup the way a person would.
  • Prefetch. If the caller's number identifies their account, load the account at the start of the call.
  • Speed up the backend. Sometimes the agent exposes slow internal systems that were already hurting human agents too.

Barge-in and interruptions

Latency isn't only about starting fast. It's also about stopping fast. When the caller starts talking while the agent is speaking, the agent should stop within a fraction of a second and listen. An agent that finishes its long sentence while the caller says "no, no, that's wrong" feels much slower than its reply time suggests.

Barge-in needs care: you don't want background noise or a cough to cut the agent off, so good systems look for actual speech, not just sound.

Shorter replies feel faster

A reply that takes 15 seconds to speak feels slower than one that takes 4, even if both started instantly. Keep agent turns short. One or two sentences, one question at a time. This also helps comprehension, as we discuss in what makes a text-to-speech voice sound natural.

Why owning the stack helps

When speech recognition, the language model and text to speech come from three different vendors, every turn crosses several network boundaries, and each provider buffers in its own way. Running all three together, close to each other, removes those hops and lets the stages stream into one another directly. That's one of the reasons we build our models in-house.

How to measure it

Don't trust demo videos. Measure on real phone calls:

  1. Call the agent from an ordinary mobile phone.
  2. Record the call.
  3. In an audio editor, measure the gap between the end of your speech and the start of the agent's reply, across 20 or more turns.
  4. Look at the median and the worst cases. Occasional long pauses are what callers remember.

Measure turns with tool calls separately from simple replies.

A latency budget

It helps to think of response time as a budget you allocate across the pipeline. The numbers below are illustrative, not measurements of any particular system, but they show how the parts add up:

Stage Example allowance
Detecting the caller has finished 200 to 400 ms
Final transcript after end of speech 50 to 150 ms
Language model to first words of reply 200 to 500 ms
Text to speech to first audio 100 to 200 ms
Network, both directions 100 to 200 ms

Add the low ends and you get a snappy reply. Add the high ends and you're well over a second before the caller hears anything. Every stage needs attention, because saving 300 ms in one place is easily lost in another.

Large clock hanging in a train station
Photo: Snufkin, StockSnap (CC0)

Perceived latency vs measured latency

What callers feel isn't exactly what the stopwatch says. A few techniques make an agent feel faster without changing raw processing time:

  • Short opening words. If the reply starts with "Sure" or "Got it," the caller hears a response immediately while the rest of the sentence follows.
  • Acknowledging before lookups. "Let me check that" turns a two-second database wait into a natural pause.
  • Shorter replies. Two short sentences beat one long one.
  • Consistent timing. Callers adapt to a steady rhythm. Erratic delays, sometimes instant and sometimes three seconds, feel worse than a consistent moderate delay.

These aren't tricks to hide a slow system. They're the same things people do in conversation.

Common causes of slow agents

When an agent feels slow in production, the cause is often one of these:

  • Fixed long silence timeouts for end-of-turn detection
  • Non-streaming components somewhere in the chain, often text to speech
  • Calls hopping between regions or vendors, adding network round trips at each stage
  • Large instructions or knowledge sent with every turn when only part is needed
  • Slow tool calls to internal systems, run one after another instead of in parallel
  • Cold starts where infrastructure spins up on the first call after a quiet period

Profile a real call stage by stage. The biggest delay is often not where you expect.

Testing latency before launch

Before putting an agent in front of customers:

  1. Run at least 50 test calls from ordinary mobile phones in different locations.
  2. Record them and measure the gap after each caller turn.
  3. Separately measure turns that include lookups.
  4. Repeat at different times of day, including your busiest hours.
  5. Set a target for the median and for the slowest 10% of turns, and don't launch until you meet both.

Then keep measuring after launch. Latency tends to creep up as instructions grow and integrations are added.

Frequently asked questions

What's a good response time for a voice agent?

As close to natural conversation as possible. Most people start to notice pauses somewhere around a second, and multi-second pauses feel broken. Measure your own and compare with how a person would respond.

Does a faster language model fix latency?

It helps, but often less than expected. Turn detection, streaming and network hops frequently matter just as much.

Should the agent use filler words to cover delays?

A short natural phrase before a lookup ("let me check") is fine. Adding filler to every reply sounds unnatural and wastes the caller's time.

Written by the Voxvencer editorial team. We build and run AI voice agents for call centers and small businesses, and we write about what we see on real phone lines. Questions or corrections: info@voxvencer.com.

Keep reading