How to Evaluate an AI Voice Agent

Every AI voice agent sounds impressive in a demo the vendor controls, which is why demo-led buyers regret their choice. This guide covers the criteria worth scoring, five specific things to say on a demo call to see how an agent really behaves, why you should ask what the agent leaves behind (transcripts and analytics), and why a real proof on your own calls beats any demo.
Every AI voice agent sounds impressive in a demo. That is the problem. The vendor picks the scenario, writes the question, and controls the conditions, so what you are watching is not your business on the phone. It is a rehearsal. Then you sign, point it at real callers with real questions and real messes, and find the gap between the demo and the day-to-day. The businesses that choose well are not the ones who watched the best demo; they are the ones who knew how to test past it.
This guide is about doing exactly that: the criteria worth scoring, the specific things to say on a demo call to see how an agent behaves, and the questions that separate a system you can run your business on from one that only performs when the vendor is driving. If you are choosing an AI voice agent, this is how to judge one before it judges your callers.
The demo vs. the day-to-day
The vendor controls one of these columns. Nobody controls the other, which is exactly why it's the one worth testing.
Start with the criteria everyone agrees on, then move past them
There is a baseline every serious buyer scores, worth covering quickly so your evaluation is complete. These are table stakes, not differentiators.
- Response speed. A long pause after the caller stops talking breaks the conversation. Anything over about a second feels awkward and prompts the caller to talk over the agent.
- Voice quality and interruption handling. Natural speech, and the ability to handle a caller who cuts in mid-sentence rather than plowing ahead.
- Integrations. Whether it books into your actual calendar and writes to your actual CRM, or just claims to.
- Pricing you can predict. Flat and forecastable, versus per-minute billing that spikes in your busy months.
- Human handoff. Whether it knows when to route a call to a person instead of guessing.
Score these, but do not stop here. Every vendor's demo is built to pass exactly this list. The difference between agents shows up somewhere the checklist does not look: in what happens when the call goes off the script the vendor prepared.
Five things to say on the demo call that separate a real agent from a good demo
Here is the part no buyer's guide gives you: the specific, slightly adversarial things to actually say during a demo, and what each one reveals. Do not let the vendor drive. Ask to call the agent yourself, and try these.
Run these five on every vendor, and the field sorts itself out fast. The demo the vendor prepared looks identical across three tools. These five questions will not.
Ask to see what the agent produces, not just what it says
Most evaluations judge how the agent talks. Fewer ask what it leaves behind, and that is a mistake, because the output is where the long-term value is.
Ask every vendor two questions. First: does it give you a full transcript and analysis of every call, or just a count of how many came in? A system that only reports volume is a black box; a system that transcribes and analyzes every conversation turns your phone into a source of insight, showing you where callers drop off, which questions stump the agent, and what people actually ask for. Second: how does it improve over time? A static agent answers the same way forever. Ask whether it learns from the calls it handles, and how you would know it had.
These two questions separate a tool that just handles calls from one that helps you fix what is causing your calls to go wrong. The same record is what lets an agent qualify a caller and route the serious ones correctly instead of treating every call the same.
How this maps to what an AI voice agent should be
Put the tests together and a picture forms of what actually matters: an agent that answers the off-script question, adapts mid-call, knows its limits, and hands you back a record you can learn from. AI voice agents like Pesta are built around exactly these, which is why they hold up past the demo rather than only in it. The honest way to judge that is to put it through these tests on a real call yourself, not to watch a controlled walkthrough.
The off-script test is the one Pesta is built to pass. Most agents run on rigid scripts, so the unscripted question, the whole reason the test works, stalls them. Pesta answers from a knowledge base instead. Its BASA capability gives a real answer to a question nobody pre-wrote, the difference between an agent that resolves the off-script call and one that deflects it. Pesta is positioned as the first AI voice agent to work this way.
And on the "what does it leave behind" test, Pesta's call analysis reads every conversation and flags where calls break down, the questions that confuse callers, the point where they drop off, so the phone becomes something you can measure and improve rather than a black box. Powered by Deepdub, it also handles callers in a wide range of languages, which a mixed customer base needs more than most tools admit. None of this replaces your team; it carries the volume so your people handle the calls that genuinely need a person.
Run a real proof before you sign
The single most reliable rule here: never buy on the demo alone. Ask for a limited trial on your own calls, your real callers, your real busy hour, before you commit. A vendor confident in production will welcome it. One who only wants you to see the controlled demo is telling you something.
- Test on your own volume, not a script. The agent that shines on one clean call can fall apart when ten come in at once.
- Model the cost at twice your current volume. A per-minute plan that looks cheap today can invert the ranking as you grow.
- Read the trial's transcripts. Not the summary metrics, the actual conversations. That is where you find out what your callers really experienced.
The vendor that welcomes a real proof, hands you the transcripts, and improves over time is the one worth signing. The one that only wants you to admire the demo has told you how it will perform once the demo is over.
FAQ
[Q]What is the most important thing to test in an AI voice agent?[/Q]
[A]
The off-script call. Anyone can handle "what are your hours." The question that predicts real-world performance is a messy, unscripted one phrased the way a confused customer would actually say it. A scripted agent stalls or takes a message; a capable one answers. If you only run one test, run that one.
[/A]
[Q]Why shouldn't I just choose based on the demo?[/Q]
[A]
Because the vendor controls the demo: the scenario, the script, the conditions. It shows you the agent at its best under ideal conditions, not how it behaves with your callers on your busiest day. The demo is a rehearsal, and buyers who choose on the rehearsal tend to regret it within the first year.
[/A]
[Q]Should an AI voice agent give me call transcripts?[/Q]
[A]
Yes, and it is worth making non-negotiable. A system that only reports how many calls came in is a black box. One that transcribes and analyzes every conversation shows you where callers drop off and what they actually ask for, which is where the value beyond call-answering lives.
[/A]
[Q]How do I know if an AI voice agent will scale with my business?[/Q]
[A]
Test it at volume during the trial, not on a single clean call, and model the cost at roughly twice your current volume. An agent that handles one call beautifully can stumble when ten arrive at once, and a pricing model that looks cheap now can become the most expensive option as you grow.
[/A]