Skip to main content
Real Human Feedback

The scoring rubric

What we test: the published rubric

Every test call is scored against the same nine categories, and every score in your report traces back to one of them. The rubric is public so you can see exactly how your AI receptionist is evaluated, and so can your vendor.
Each category on each call is markedPassPartialFailwith a note explaining why.

R1

Greeting and identification

Greeting and identification of the business.

What the tester listens for

  • The call is answered, and how long the caller waits before hearing a voice
  • The business name is said clearly and correctly
  • The caller can tell what kind of business they reached
  • The opening invites the caller to speak rather than reading a long menu

Example test call

A first time caller who is not sure they dialed the right business asks, "Is this the place on Main Street?"

R2

Understanding intent

Understanding the caller's intent, including accents, interruptions and changed minds.

What the tester listens for

  • The reason for the call is understood on the first or second try
  • Accents, fast talkers and background noise are handled
  • The caller can interrupt and be heard
  • A caller who changes their mind midway is followed, not forced back to the start

Example test call

A caller starts asking about pricing, interrupts themself, and asks to book for next week instead.

R3

Booking accuracy

Booking and scheduling accuracy against the business's real availability.

What the tester listens for

  • Offered times match the availability the business gave us
  • Date, time, name and callback number are read back correctly
  • The booking actually lands where the business expects it
  • Reschedules and cancellations are handled without double booking

Example test call

A caller asks for "the first thing Tuesday morning, or else Thursday after three."

R4

Handoff to a human

Handoff to a human when a caller asks for one or when the situation requires it.

What the tester listens for

  • A caller who asks for a person is offered a transfer or a clear alternative
  • The transfer connects, or the caller is told honestly what happens next
  • Situations the business marked as human only are routed that way
  • The caller is not looped back into the AI after asking to leave it

Example test call

A caller says, "I would really rather talk to a person, is anyone there?"

R5

After hours and urgent calls

After hours behavior, emergencies and urgent requests.

What the tester listens for

  • After hours callers hear accurate hours and next steps
  • Urgent or emergency language is recognized
  • The business's own emergency instructions are followed
  • The caller knows when they will hear back

Example test call

A caller at 9:40 at night reports water coming through the ceiling.

R6

Edge cases

Edge cases: off script questions, wrong numbers, spam, angry callers.

What the tester listens for

  • Off script questions get an honest answer or an honest "I do not know"
  • Wrong numbers are handled politely and briefly
  • Obvious spam is ended without wasting staff time
  • An upset caller is acknowledged and not argued with

Example test call

A frustrated caller says it is the third time they have called about the same invoice.

R7

Invented information

Hallucinated or invented information: prices, policies, hours or services that are not real.

What the tester listens for

  • Quoted prices match what the business actually charges, or none are quoted
  • Policies, warranties and guarantees are not made up
  • Hours and locations are correct
  • Services the business does not offer are not promised

Example test call

A caller asks, "Do you guys offer a senior discount?" at a business that does not have one.

R8

Latency and talk-over

Latency, silences and talk-over.

What the tester listens for

  • Pauses before each reply are short enough to feel natural
  • Long silences that make a caller say "hello?" are noted
  • The AI does not talk over the caller
  • The call does not drop or restart

Example test call

A caller gives a long, winding answer with pauses in the middle of it.

R9

Message accuracy

Message taking and the accuracy of what reaches the business afterward.

What the tester listens for

  • Names, numbers and addresses are captured correctly
  • The summary the business receives matches what the caller said
  • Urgency is carried through into the message
  • Nothing the caller said is dropped or invented in the summary

Example test call

A caller spells an unusual last name and gives a callback number with an extension.

AI voice agent evaluation, from the caller's side

Most voice agent testing is done from inside the system, by the team that built the agent. This rubric looks at it from the outside, as a caller would, and adds the failure modes that matter most with AI: invented information, latency and talk-over, and handoffs that do not connect.

The same rubric is used for client tests and for the public benchmark methodology. Want to try it yourself first? Use the AI receptionist checklist.

Hear your AI receptionist the way your customers do

Tell us the number your AI receptionist answers and anything you want us to look at. Real people place the calls, and you receive a written report.