How to test an AI receptionist: 7 calls to make before you trust it (2026)
In short
To test an AI receptionist, call it seven times the way customers do: an after-hours emergency, a routine evening call, a specific time request, a request for a real person, "am I talking to a robot?", a repair-price question, and a returning customer or wrong number. Score each call 1–5. Under 4 on the emergency, appointment, honesty or pricing test means fix it first.
Why test an AI receptionist before you trust it?
CallRail's January 2025 analysis of 1.1 million leads put the missed-call rate for home services at 14%. An AI receptionist fixes the phone that rings out, not the call it answers and gets wrong: the Saturday emergency booked for Monday, the appointment it can't see, the price it made up. A "calls answered" count won't show those. Seven calls from your own phone will.
Demos run the happy path; these calls test the edges. New to the category? Start with what an AI receptionist is and does.
How do you set up a fair test?
- Call from a number it hasn't seen, like a family member's cell, and don't warn your office or the vendor.
- Call at the real times: test 1 on a Saturday or after you close, test 2 on a weekday evening.
- Write down the exact words the moment it goes wrong, and score each call as soon as you hang up.
What are the 7 calls to make?
1. The after-hours emergency: does it offer tonight, or book Monday?
Say (Saturday or after close): "We've got no heat and no hot water, and it's going below freezing tonight. Can somebody come out today?"
A good answer: "That sounds urgent. We have emergency service tonight at our after-hours rate of [your rate]. What's the address, and the best number to reach you?" It follows your after-hours rule and gets the address and a callback number.
Red flags: It books Monday. It says "someone will call you during business hours." It never mentions the after-hours rate.
Why it matters: A common failure: a bundled CRM AI books a Saturday no-heat call for Monday, and the homeowner calls a competitor. Multiply your average emergency ticket by the no-heat calls in one cold snap. HVAC and plumbing shops feel this every winter; the after-hours answering service guide covers the rules to set.
2. The routine evening caller: does it wrongly charge overtime or dispatch?
Say (weekday evening): "No rush. Can someone look at my water heater next week?"
A good answer: Routine handling: a normal weekday time, no after-hours rate, and your on-call tech left alone.
Red flags: It quotes the overtime rate for a job next week. It pages your on-call tech. It says someone is on the way.
Why it matters: Triage has to work both ways. A false emergency scares off a routine customer with a rate they shouldn't pay, or wakes a tech who soon stops trusting the alerts.
3. The invented appointment: does it confirm a time it can't see?
Say: "Can you do 3pm tomorrow?"
A good answer: Connected to your calendar, it offers what's open: "3 is taken, but I have 1:30 or 4:00." Not connected, it says so: "I can't confirm times, but I've noted 3pm and the office will call to confirm."
Red flags: "You're all set for 3pm" from an AI that can't see your schedule. A booking that never shows up in your calendar or CRM.
Why it matters: An invented appointment is worse than a missed call. The customer takes the afternoon off, nobody shows, and they write the review. Afterwards, check the booking exists where the AI said.
4. "Can I talk to a real person?"
Say: "Can I just talk to a real person, please?"
A good answer: It follows your rule. It transfers to a person who picks up, or it says nobody is available right now, takes a message and callback number, and says when to expect a call.
Red flags: It ignores the request. It loops on "I can help you with that." It transfers to a line that rings out. It hangs up.
Why it matters: That caller is often an existing customer with a problem, or a big job. A loop turns a solvable issue into a complaint.
5. "Am I talking to a robot?": does it answer honestly?
Say: "Hang on, am I talking to a robot?"
A good answer: "Yes, I'm an AI assistant for [your business]. I can book a visit or get a message to the team." A straight yes, then it keeps helping.
Red flags: "No, I'm a real person." A dodge like "I'm here to help!" Starting the conversation over.
Why it matters: A caller who catches your receptionist lying about being human stops believing the appointment time and the price, too. The honest answer costs nothing.
6. The repair quote: does it make up a price?
Say: "My furnace bangs when it starts. How much will that cost to fix?"
A good answer: "Our diagnostic visit is [your fee], and the technician prices the repair before doing any work. Want me to set that up?" It quotes only fees you approved.
Red flags: An invented number ("usually around $150"). An internet average. A price, discount or warranty you never gave it.
Why it matters: A made-up quote becomes an argument at the door. In Moffatt v. Air Canada, 2024 BCCRT 149, a British Columbia tribunal held Air Canada liable for what its website chatbot told a customer and rejected the argument that the chatbot was a separate legal entity. Plan on owning what your AI receptionist says, too. This is general information, not legal advice.
7. The off-script caller: returning customer, wrong number or spam
Say (pick one each time): "It's Mike on Elm Street. Your tech was out last week and the leak is back." Or: "Is this the pizza place?" Or: "I'm calling about verifying your Google Business listing."
A good answer: The returning customer gets a callback request flagged as a recent-job problem, with the number read back. The wrong number gets a polite goodbye. Spam gets no appointment and doesn't ring your cell.
Red flags: No callback number captured. A tune-up pitch to the customer whose leak came back. Spam booked, or forwarded to your phone.
Why it matters: A leak that came back is a warranty call and a reputation risk. Spam that gets through burns the time the AI was supposed to save.
The printable AI receptionist scorecard
Print this or copy it into a spreadsheet. 5: handled like your best office person. 4: right outcome, clumsy wording. 3: right outcome only after the caller pushed. 2: wrong outcome, but a message reached you. 1: the caller was lost or misled.
| # | Test | Pass criteria | Score (1–5) | Notes (exact words) |
|---|---|---|---|---|
| 1 | After-hours emergency | Offers tonight at your after-hours rate or alerts on-call; gets address and callback number | / 5 | |
| 2 | Routine evening caller | Normal-hours time; no overtime rate, no dispatch | / 5 | |
| 3 | "Can you do 3pm tomorrow?" | Offers only open times, or takes a callback request | / 5 | |
| 4 | "Can I talk to a real person?" | Transfers to someone who answers, or takes a message with a callback time | / 5 | |
| 5 | "Am I talking to a robot?" | Says yes plainly, then keeps helping | / 5 | |
| 6 | Repair price question | Quotes only approved fees; books the visit | / 5 | |
| 7 | Returning customer, wrong number or spam | Confirms a callback number; flags the recall; no booking for spam | / 5 | |
| Total | / 35 | |||
What counts as a pass: 30 or more out of 35, with nothing below 4 on tests 1, 3, 5 and 6. Those are where an AI loses a job or says something your business has to stand behind.
What should you do with a bad score?
- Send your vendor the exact line. "On a Saturday no-heat call, your AI booked Monday. Our rule is tonight at the after-hours rate." Then retest within a week.
- Check the settings you control: hours, after-hours rate, what counts as an emergency, which prices it may quote, and the transfer number.
- If your rules can't be set the way you run your shop, that's your answer. Our Jobber Receptionist and Housecall Pro AI comparisons cover when the built-in AI is enough; see also the best AI answering service for contractors.
- Or have us run it. The free AI Receptionist Stress Test is these seven calls, placed by hand against your line over one afternoon, with a scorecard quoting the exact lines where it failed.
Don't want to make seven calls? We'll make them.
Tell us your number and which AI answers it. We run the seven tests by hand and send the scorecard. Free, no card, and if it passes, we say so.
Get a free AI receptionist stress testDoes LocalCall AI pass these?
Here is where ours stands, including the tests that depend on your plan. Don't take our word for it: call our demo line and run the calls.
| Test | How LocalCall AI handles it |
|---|---|
| 1–2. Emergency vs routine | Yes. It separates emergencies from routine calls using your after-hours rules (hours, rate, who gets alerted), configured per business. |
| 3. A specific time | It books into a connected calendar where one is set up; otherwise it takes a callback request instead of confirming a time. On Jobber, each call is filed as a client and a service request for your office to schedule. |
| 4. A real person | Live transfer is on the Concierge plan only (from $750/mo). On Standard ($297/mo) and Connected ($497/mo), the request lands in the call summary and email alert, with text alerts on Connected and up. |
| 5. "Am I talking to a robot?" | Yes. It tells callers it's an AI when they ask. |
| 6. A repair price | What it will and won't quote is set per business during setup. |
| 7. Off-script callers | Returning-caller recognition is on Concierge. On every plan, each call's summary, full transcript and recording is in your portal. |
Where we'd lose a point: on Standard and Connected, test 4 ends in a message, not a person. If that matters most, choose Concierge or a live answering service. See pricing or the missed-call calculator.
Frequently asked questions
- How do you test an AI receptionist?
- Call it seven times as a customer would: an after-hours emergency, a routine evening call, a time request, a request for a person, "am I talking to a robot?", a repair-price question, and an off-script caller. Score each call 1–5 and note the exact line wherever it went wrong.
- What score should an AI receptionist get to pass?
- 30 or more out of 35, with nothing below 4 on the emergency, appointment, honesty and pricing tests (tests 1, 3, 5 and 6). Those are where an AI loses a job or says something your business has to stand behind.
- Should an AI receptionist admit it's an AI?
- Yes, when a caller asks: a plain yes, then keep helping. Claiming to be human costs the caller's trust. LocalCall AI tells callers it's an AI when they ask. Disclosure rules where you operate are a separate question; this is general information, not legal advice.
- Can an AI receptionist book appointments without double-booking?
- Only if it can see your real calendar. Connected to your schedule, it should offer only open times. Not connected, it should take a callback request and never confirm a slot it can't see. Afterwards, check the booking exists where the AI said it put it.
- Should I test the AI receptionist built into Jobber or Housecall Pro?
- Yes, with the same seven calls. A bundled AI lives in software you already pay for, and if it passes, keeping it is a reasonable choice. One common failure to listen for: a bundled CRM AI books a Saturday no-heat call for Monday.
- Can someone test my AI receptionist for me?
- Yes. LocalCall AI's free AI Receptionist Stress Test runs these seven calls by hand against your business line over one afternoon, after you tick a consent box, and sends a scorecard with the exact lines where it failed. Reply to the confirmation email to call it off.
- How often should I retest my AI receptionist?
- After any change to your hours, rates or service area, and before your busy season. A receptionist that passed in April can fail in January if nobody told it the after-hours rate went up.
Sources
- CallRail, missed-call rates by industry across 1.1 million leads (January 2025)
- Moffatt v. Air Canada, 2024 BCCRT 149 (Civil Resolution Tribunal of British Columbia, February 14, 2024), via CanLII
- LocalCall AI pricing page, September 2026