The market for AI phone agents got crowded fast, and the products vary far more than the marketing does. Every vendor site says roughly the same thing: natural conversation, books appointments, never misses a call.
The differences that matter are underneath that, and none of them come up unless you ask. This is the checklist we would use ourselves.
If a vendor is evasive on any of these, that is the answer.
Before you talk to anyone
Two things to establish first, because they determine which vendors are even relevant.
What your call mix actually is. Pull a month of logs and listen to twenty calls at random. You are looking for: what proportion are routine bookings, what proportion need judgement, when they arrive, and what information an engineer needs before attending. This takes an hour and changes the conversation entirely — you stop being sold to and start evaluating against a spec.
What you are actually trying to fix. “Missing calls” is not specific enough. Missing evening bookings, missing surge calls during storms, and missing calls while on the tools are three different problems with different solutions. If you cannot name the bucket, you cannot tell whether a demo addressed it.
The six questions vendors dislike
1. Who owns the call recordings and transcripts?
You should. Get it in writing. Recordings are your customer data, they are useful for training and disputes, and if you switch providers you want to take them with you. A vendor whose answer is vague here is telling you something about the exit.
Ask specifically: can you export them, in what format, and does that survive cancellation.
2. What happens on a failed handoff?
Every system fails on some calls. What you want to know is the fallback behaviour when a transfer does not connect — during business hours and at 2am, which should be different. “It transfers to a human” is not an answer if the human is asleep. The right answer involves capturing a callback number and setting an explicit expectation. We go through the mechanics in what happens when an AI receptionist doesn’t understand.
3. Is pricing per minute, per call, or per resolution?
These behave very differently. Per-minute punishes chatty callers and rewards a curt agent, which is not what you want. Per-call is more predictable and can encourage rushing. Per-resolution sounds fair until you read the definition of resolution.
Then ask the real question: what does a bad month look like? Model a storm week where volume triples. If the answer is uncapped, you need a cap or an alert.
4. Can we change the script ourselves?
Prices change. You add a service. A suburb comes into range. If every change is a support ticket with a five-day turnaround, the system will drift out of date within a quarter and you will stop trusting it.
Ask to see the editing interface, not a description of it. Ask how long a price change takes to go live.
5. What is the notice period, and what happens to the phone number?
Thirty days is normal. Twelve-month lock-ins on an unproven system are not, particularly for a category this young. And if they provisioned the number, establish now whether you can port it out. Losing a number that is printed on your vans is a genuinely expensive way to discover a contract term.
6. Where does the data live, and who can access it?
Relevant for compliance if you are anywhere with data residency rules, and relevant for everyone if your calls contain payment details or health information. Ask which sub-processors are involved — most vendors are built on someone else’s speech and language models, and that is fine, but you should know.
What to test in the demo
Vendors demo the happy path. Break it deliberately.
Interrupt it three times. Talk over it mid-sentence. Good systems yield and pick up the thread. Poor ones restart their sentence, which is unmistakable and irritating.
Give a partial address. Street, no number. It should notice and ask, not proceed.
Change your mind halfway. Book Tuesday, then say actually Thursday. Correction handling separates systems built by people who listened to real calls from ones assembled from a template.
Ask for something out of scope. A service they do not offer. The right answer is a clean no and a route to a human. An invented answer is disqualifying — that is a system that will confidently tell your customers something untrue.
Ring from a bad line. Outside, near traffic, on a mobile. This is where your real callers ring from and where transcription is genuinely tested. A demo on a headset in a quiet office proves nothing.
Ask what it does when the caller is upset. Say something distressed and see whether it continues cheerfully through the booking script. It should route to a person.
Get a recording of your own test call. Then listen to it a day later, when you are not in a demo mindset. Things you missed in the moment are obvious on playback.
The integration questions
This is where deployments quietly fail — not on the conversation, on what happens to the data afterwards.
Does it write into the calendar you actually use? Not “we integrate with major calendars”. Yours, specifically, with your engineers’ availability and your job durations. A booking that lands in a system nobody looks at is worse than no booking.
Does it respect real availability? Travel time between jobs, engineer skills, whether a job type needs two people. A system that books three jobs across town in one afternoon has created work, not saved it.
Where do the lead details go? Your CRM, an email, a spreadsheet. Establish the failure behaviour: if the CRM write fails, does the lead vanish or land somewhere recoverable? Persist-first is the only safe design — capture the lead, then attempt everything else.
What about existing customers? If someone rings whose details you have, does the agent know? For repeat-heavy trades like pest control or pool service, asking a regular customer for their address every time is a small insult that accumulates.
The questions about what is underneath
Almost every vendor in this category is assembling components rather than building from scratch, and that is fine. What matters is whether they will tell you which components.
Most systems combine a speech-to-text model, a language model, a text-to-speech voice, and a telephony layer. The telephony is usually Twilio or similar; the voices are frequently from one of a handful of specialist providers. A vendor who is straightforward about this stack is easier to trust than one implying it is all proprietary, because the second is either misleading you or has built something unusual that you should ask harder questions about.
Three things follow from knowing the stack:
You can estimate the floor on latency. More hops means more delay. A vendor running everything through several external APIs in sequence cannot be as fast as one that has collapsed some of those steps.
You know who else holds your call data. Each component in that chain is a sub-processor. If your calls involve health information, HIPAA guidance on business associates makes this a compliance question rather than a preference.
You can judge lock-in. A vendor built on standard components is easier to leave than one whose scripts exist only in a proprietary format. Ask whether you can export your call flow in any readable form.
None of this means avoid vendors who assemble. It means prefer the ones who will tell you.
Red flags
Guaranteed conversion numbers. Nobody can promise that; it depends on your close rate, your pricing, and your market.
“It learns automatically from every call.” In practice improvement comes from a human reviewing failures and changing the script. Vendors implying otherwise are describing a process that will not happen.
No mention of failure. If a vendor cannot describe how their system fails, they have not measured it, which means you will be the one who does.
Reluctance to let you test freely. A scripted demo you cannot deviate from is hiding something.
Setup that is entirely self-serve for a complex trade. Fine for a simple booking flow. Not fine for anything where a wrong answer dispatches a van or misjudges an emergency.
Pricing that only appears after a sales call. Sometimes legitimate for genuinely custom work. Often a filter for how much they think you will pay.
Scoring it
Rather than a gut call after three demos, score each vendor out of 5 on six axes and weight them for your situation:
- Conversation quality — interruption handling, correction, latency
- Failure behaviour — escalation, callback capture, honesty about limits
- Integration — calendar, CRM, availability logic
- Control — can you edit, how fast do changes go live
- Commercials — pricing model, cap, notice period, number portability
- Fit to your trade — do they understand what your calls involve
The weighting is yours. For an emergency-led trade, failure behaviour outweighs conversation quality. For a clinic, integration and data handling dominate. Writing the weights down before the demos stops you being talked into whichever axis the best salesperson emphasised.
What a sensible rollout looks like
Whoever you pick, do not go straight to full coverage.
Week one: overflow only. The agent picks up when your main line is engaged or unanswered. Your existing coverage is untouched, so the risk is bounded, and every call it takes is one you would otherwise have lost.
Week two: add out of hours. Evenings and weekends, with a clear callback expectation. This is where most of the value is.
Week three: review recordings. Actually listen to twenty calls, including the failures. This is the step people skip and it is the one that makes the difference. You will find at least one thing nobody predicted.
Week four: tune, then decide on scope. Whether to extend to daytime, whether to add call types, whether to keep going at all.
Set a decision point before you start — a date and a number. “If by the end of month one it has captured fewer than X leads, we stop.” Deployments without an exit criterion drift on for a year because nobody wants to admit it is not working.
Frequently asked
How long does setup take? For a straightforward single-trade booking flow, days. For anything involving emergency triage, multiple services, or real availability logic, a few weeks — most of which is you clarifying rules that were previously in someone’s head.
Should I get several quotes? Three is plenty, and make them run the same test calls so you are comparing like with like.
Can I use one agent across multiple locations? Usually, but check how it handles service areas and time zones. In border metros like Chattanooga or Toledo, getting this wrong means confirmed appointments an hour out or jobs quoted across a state line you are not licensed for.
What if we already have an answering service? Run them in parallel for a month before switching. Compare captured leads, not impressions. The comparison guide covers what to measure.
Is it worth it below a certain call volume? If your after-hours volume is genuinely small, probably not. Do the arithmetic first — the answer is sometimes no, and a vendor will not tell you that.
Want a second opinion on a vendor you are considering? Send us what they have quoted and we will tell you what we would ask them — including if the honest answer is that you do not need this.