Most businesses switch on an AI receptionist, call it once or twice to hear the greeting, and move on. The calls that matter, the confused caller, the after hours emergency, the question it was never given an answer to, rarely get tested until a real customer finds them.
This guide lays out the same method we use for client tests, so you can run a basic version yourself. It works with any AI receptionist, whatever product or platform it runs on.
Step 1: Decide what matters to your callers
Start with your real callers, not the receptionist's feature list. Write down the five or six reasons people call you most often, and the one or two calls that would hurt most if they went wrong.
- The most common call, such as booking or a price question
- The most valuable call, such as a new customer or a large job
- The most urgent call, such as an emergency or a same day problem
- The call you are least sure about, such as a question you never gave the AI an answer to
If you already have a concern, write it down. A test shaped around a specific worry finds more than a general one.
Step 2: Write call scenarios
A scenario is a short brief for the caller: who they are, why they are calling, and one or two things they will do that a script would not expect. Keep them realistic.
| Scenario | What the caller does | What it checks |
|---|---|---|
| New customer booking | Asks for the first opening, then changes to a different day | Understanding, booking accuracy |
| Price question | Asks for a price you have never published | Invented information |
| After hours problem | Calls late at night with an urgent issue | After hours behavior, urgency |
| Asks for a person | Politely insists on a human | Handoff to a human |
| Messy caller | Talks over the AI, pauses, spells a hard name | Latency, talk-over, message accuracy |
Step 3: Place the calls like a real caller
The point of a test call is to sound like a customer, not like a tester. Call from a phone the receptionist does not recognize. Talk at a normal pace, with the small hesitations people really have.
- Call during business hours and after hours
- Call as a first time caller and as a returning one
- Include at least one call with background noise, such as a car or a busy room
- Interrupt once, change your mind once, and ask one thing that is off script
Use fictional details
Never give the receptionist real customer or patient information during a test. Use a made up name and a number you control.
Step 4: Take notes on every call
Write notes during or right after the call, before you forget the details. Record the time, the scenario, what you asked, what the AI said, and anything that felt off.
- The exact words of any price, policy or promise the AI gave
- Any pause long enough that you wondered if the call dropped
- Whether you reached a person when you asked
- What the message or booking said when it reached the business
That last point is easy to skip and often the most revealing. Compare the call with what actually landed in your calendar, CRM or inbox.
Step 5: Score each call against a rubric
Scoring against fixed criteria keeps one memorable call from coloring everything. Our published rubric has nine categories, from the greeting to the message the business receives. Mark each category pass, partial or fail for each call, with a note explaining why.
Step 6: Write up what you heard
A good write up lists every call, what happened and the score, then a short set of findings in plain language: what the receptionist does well, where callers struggle, and which calls carry the most risk. Our sample report shows the format.
Take the findings to your vendor or whoever set up your receptionist. Specific examples, with the caller's words and the AI's reply, are far easier to act on than a general complaint.
Why outside callers find more
Testing your own receptionist has a limit: you know what it is supposed to say, so you tend to ask in ways it understands. Real people who do not know your setup call the way your customers do. That is why we built Real Human Feedback: the calls are placed by humans, the rubric is public, and we do not sell the fix.
Questions
Quick answers to what readers ask next.
How many test calls do I need?
Enough to cover each scenario that matters to you at least once, during and after business hours. A handful of well planned calls finds more than dozens of identical ones.
Can I use automated tests instead?
Automated tools are useful for teams building voice agents. They tend to miss the ways real people talk: interruptions, changed minds, background noise and questions nobody scripted.