Skip to content
Blog
· Glytos

How a voice AI agent works, and where it breaks on real calls

What happens inside a voice AI agent, the three ways to build one, and the failures that only show up on real phone calls.

A voice AI agent is software that holds a spoken conversation on its own. It listens to a caller, works out what they mean, decides what to do and answers out loud, on a phone line or in a browser. Within the same conversation it can look something up, book an appointment, transfer the call or end it.

That part is easy to describe. This post is about the rest: what sits between the caller's words and the agent's reply, the choices that decide how an agent behaves, and the failures we have run into on real calls while building Glytos. None of them is exotic, and anyone who has shipped voice will recognise most of them. They are worth writing down because almost none of them raise an error. The call connects, the agent talks, and something is quietly wrong.

What happens inside a voice AI agent

Whatever it is built from, every turn of a call goes through the same four jobs:

  1. Hearing. The caller's audio arrives as a stream of small frames.
  2. Deciding the caller has finished. This is turn detection. Too eager and the agent cuts people off mid-thought; too patient and the line goes quiet.
  3. Understanding and deciding. A language model reads the conversation so far, its instructions and anything retrieved for it, then either answers or calls a tool.
  4. Speaking. The reply becomes audio and plays back, ideally starting before the whole reply has been written.

Around those four sit the parts that make it an agent rather than a demo: tools, knowledge, call control (transfer, end, what to do about silence) and a record of what happened.

Three ways to build a voice AI agent

Cascaded: speech-to-text, a language model, text-to-speech

Three services in a chain. The strength is control. Every step produces text you can read, log, filter and test, you choose each component on its own, and the cost of each one is visible. The weakness is that it is a relay: every handoff takes time, and the caller's tone of voice is lost the moment speech becomes text.

Speech-to-speech

One model hears audio and answers in audio. It sounds more natural and keeps the tone of the conversation. What it gives up is control. There is no text step to inspect before the model speaks, and on some models you cannot make it say an exact sentence word for word: it is told what to say and says it in its own words. That matters more than it sounds when the line is a required disclosure, a closing sentence or a fixed fallback.

Full duplex

A speech model that keeps listening while it talks, so cutting in or adding a detail halfway through its answer becomes ordinary. It also breaks assumptions a cascaded stack relies on. Muting the caller while the agent speaks is a standard guard there, and on a full-duplex model the same mute simply deafens it.

Which one to choose

If the call follows steps that must happen in order, or needs exact wording, start cascaded. If the conversation is open-ended and naturalness matters most, try speech-to-speech. It does not have to be a permanent decision: on Glytos the speech engine is a setting on the agent, and its prompt, tools, knowledge and phone number stay as they are when you switch.

Turn-taking is where voice agents feel wrong

Most complaints about a voice agent that "feels robotic" are about turn-taking, not about the voice.

Knowing when the caller has stopped. A pause to think and the end of a sentence sound the same to a silence detector. Tuning this is a trade between interrupting people and leaving gaps, and the right point differs between a quick confirmation call and an open conversation.

Interruptions. Letting the caller cut the agent off sounds like an obvious yes. But on a phone on speaker, or a laptop without headphones, the agent's own voice comes back through the microphone. With interruptions on, the agent can interrupt itself, or have its own words transcribed as a new caller turn. On Glytos this is a per-agent choice, and agents start protected: the agent finishes its sentence unless you turn interruptions on.

Silence. Decide what the agent does when nobody speaks: prompt again in the agent's own language, and end the call after a limit you set rather than holding the line open indefinitely.

Where voice AI agents break on real calls

We have run into every one of these. For each one: what it looks like, why it happens and how to catch it.

The agent answers words nobody said

Symptom: fluent, confident replies to things the caller never said.

Cause, in our case: an audio format mismatch. A speech-to-speech model expected 24 kHz audio. It received 16 kHz audio passed through unchanged with the 24 kHz label on it, so it heard every caller 1.5 times too fast and about seven semitones high. It did not complain. It answered what it thought it heard.

Lesson: a model that receives bad audio does not fail, it guesses. Check sample rate and encoding at every boundary, and when replies stop making sense, listen to what the model received rather than what the caller said.

Silence turns into "thank you"

Symptom: the caller says nothing, and the agent replies to "thank you" or "thanks for watching".

Cause: Whisper-family transcribers are trained largely on subtitled video, and they are well known for producing stock subtitle phrases when handed silence or line noise. The agent then answers a sentence nobody spoke.

Lesson: filtering on the text alone does not work, because a caller can really say "thank you". What separates the two is how long the caller actually spoke: a phrase takes about a second to say, while a phantom comes out of a fraction of a second of noise. The filter also has to match the transcriber. A rule written for a transcriber that works in segments was wrong for a streaming one, and when we once applied it too broadly it threw away a real short answer.

The caller speaks and the line goes dead

Symptom: the caller says something, the agent says nothing, and it stays that way.

Cause: the transcriber could not commit a final transcript, most often because the caller spoke a language it was not set up for. The turn closed empty, nothing reached the model, and nothing anywhere reported an error.

Lesson: measure the audio itself, both its level and how many seconds of it were actually delivered. When there is proof the caller spoke but no transcript survived, the agent should ask them to repeat, a limited number of times, and then close the call politely instead of looping or sitting in silence.

The model understood, the transcript did not

Symptom: a Turkish caller asked "Bu arada sen kimsin?" ("By the way, who are you?"). The agent answered the question correctly. The stored transcript recorded the caller's words in Korean.

Cause: on speech-to-speech models, the conversation and the transcript are separate passes. The model hears the audio directly and takes its language from its instructions. The transcript comes from a separate recognition step that, left unconfigured, detects the language on its own, and short turns are easy to misdetect.

Lesson: a transcript is evidence of what the caller said, not proof. If your reports, analysis or extracted data are built on the transcript, give the transcription step the agent's language too, not just the model.

A placeholder that turns into nothing

Symptom: an outbound call opens with "Hello." where it should have said "Hello John."

Cause: the prompt said "Hello {{name}}" with the name taken from the contact list. A template that cannot find a value and silently replaces it with nothing produces "Hello." and raises no error. We had exactly this: the contact's data reached the call but not the prompt.

Lesson: test the rendered prompt with a real row from your list before a campaign starts. A template that fails quietly is worse than one that fails loudly.

Faster reasoning, worse flow

Symptom: the agent skipped a step its instructions required.

Cause: lowering a reasoning model's effort makes it answer sooner. On a live call, an agent with a multi-step script running at the lowest setting ended the call when the caller said they were not available, instead of offering a callback as its instructions told it to.

Lesson: speed was bought with flow-following. Give complex flows more reasoning, or move the flow out of the prompt and into a workflow, where each step is enforced by structure and the model only has a small job at each step.

How to test a voice AI agent before it takes real calls

  • Talk to it in the browser first. Same agent, same tools, no phone number needed.
  • Test the logic in text. Scripted conversations with checks: did it end on the right step, collect the right values, say a required phrase, and does a graded rubric such as "the agent offered a refund" hold? Text tests are fast and repeatable and catch regressions in the logic. They do not test the audio, so the failures above still need real calls.
  • Compare versions. Run two versions of the agent side by side on the same messages before you publish a change.
  • Read every real call. The transcript, the recording with caller and agent on separate lanes, the timeline of routing decisions and tool calls, and the cost. Watch calls live and step in when one goes wrong.

All of this is how Glytos works, and none of it is unique to voice: it is what any software going to production should have.

What a voice AI agent costs

There is no single number. The cost of a call is the sum of its components, each priced in its own unit: the language model by tokens, the voice by characters, the transcriber by minutes of audio, a speech-to-speech model by its own audio tokens or minutes, plus the platform.

That has practical consequences. A talkative agent costs more in voice. Long instructions are sent again on every turn and cost input tokens, which prompt caching softens. So a per-minute average hides what actually drives your bill. Glytos shows the cost of each call broken down by component, and with your own provider keys the provider bills you directly and only the platform fee is charged.

Questions to ask before choosing a voice AI agent platform

  • Can you see every turn: what the transcriber heard, what the model received, what it decided and what it cost?
  • Can you switch between cascaded and speech-to-speech without rebuilding the agent?
  • Can you control interruptions and silence per agent?
  • Can you test a change before it reaches callers, and compare versions?
  • Can you use your own phone numbers or SIP trunk, and your own provider keys?
  • Can you export your agents and take them with you?

Try it

You can build a voice AI agent on Glytos from a single prompt or a visual workflow, call it from your browser, and see every turn of every call. Start free, or read the quickstart first.

More from the blog