The scoring rubric
What we test: the published rubric
R1
Greeting and identification
Greeting and identification of the business.
What the tester listens for
- The call is answered, and how long the caller waits before hearing a voice
- The business name is said clearly and correctly
- The caller can tell what kind of business they reached
- The opening invites the caller to speak rather than reading a long menu
Example test call
A first time caller who is not sure they dialed the right business asks, "Is this the place on Main Street?"
R2
Understanding intent
Understanding the caller's intent, including accents, interruptions and changed minds.
What the tester listens for
- The reason for the call is understood on the first or second try
- Accents, fast talkers and background noise are handled
- The caller can interrupt and be heard
- A caller who changes their mind midway is followed, not forced back to the start
Example test call
A caller starts asking about pricing, interrupts themself, and asks to book for next week instead.
R3
Booking accuracy
Booking and scheduling accuracy against the business's real availability.
What the tester listens for
- Offered times match the availability the business gave us
- Date, time, name and callback number are read back correctly
- The booking actually lands where the business expects it
- Reschedules and cancellations are handled without double booking
Example test call
A caller asks for "the first thing Tuesday morning, or else Thursday after three."
R4
Handoff to a human
Handoff to a human when a caller asks for one or when the situation requires it.
What the tester listens for
- A caller who asks for a person is offered a transfer or a clear alternative
- The transfer connects, or the caller is told honestly what happens next
- Situations the business marked as human only are routed that way
- The caller is not looped back into the AI after asking to leave it
Example test call
A caller says, "I would really rather talk to a person, is anyone there?"
R5
After hours and urgent calls
After hours behavior, emergencies and urgent requests.
What the tester listens for
- After hours callers hear accurate hours and next steps
- Urgent or emergency language is recognized
- The business's own emergency instructions are followed
- The caller knows when they will hear back
Example test call
A caller at 9:40 at night reports water coming through the ceiling.
R6
Edge cases
Edge cases: off script questions, wrong numbers, spam, angry callers.
What the tester listens for
- Off script questions get an honest answer or an honest "I do not know"
- Wrong numbers are handled politely and briefly
- Obvious spam is ended without wasting staff time
- An upset caller is acknowledged and not argued with
Example test call
A frustrated caller says it is the third time they have called about the same invoice.
R7
Invented information
Hallucinated or invented information: prices, policies, hours or services that are not real.
What the tester listens for
- Quoted prices match what the business actually charges, or none are quoted
- Policies, warranties and guarantees are not made up
- Hours and locations are correct
- Services the business does not offer are not promised
Example test call
A caller asks, "Do you guys offer a senior discount?" at a business that does not have one.
R8
Latency and talk-over
Latency, silences and talk-over.
What the tester listens for
- Pauses before each reply are short enough to feel natural
- Long silences that make a caller say "hello?" are noted
- The AI does not talk over the caller
- The call does not drop or restart
Example test call
A caller gives a long, winding answer with pauses in the middle of it.
R9
Message accuracy
Message taking and the accuracy of what reaches the business afterward.
What the tester listens for
- Names, numbers and addresses are captured correctly
- The summary the business receives matches what the caller said
- Urgency is carried through into the message
- Nothing the caller said is dropped or invented in the summary
Example test call
A caller spells an unusual last name and gives a callback number with an extension.
AI voice agent evaluation, from the caller's side
Most voice agent testing is done from inside the system, by the team that built the agent. This rubric looks at it from the outside, as a caller would, and adds the failure modes that matter most with AI: invented information, latency and talk-over, and handoffs that do not connect.
The same rubric is used for client tests and for the public benchmark methodology. Want to try it yourself first? Use the AI receptionist checklist.
Hear your AI receptionist the way your customers do
Tell us the number your AI receptionist answers and anything you want us to look at. Real people place the calls, and you receive a written report.