Skip to main content
Real Human Feedback

The AI Receptionist Benchmark

Benchmark methodology

The benchmark is only useful if anyone can see how it works. This page sets out what we profile, how products will be tested and scored, and the rules that keep the benchmark neutral.

1. Factual profiles

Every product in the benchmark has a profile built only from what its vendor publishes: the vendor, the category, published pricing, features, the industries it targets and its integrations. Where a vendor does not publish something, the profile says so, for example “Not published” for pricing, rather than estimating it.

Pricing is recorded as published by the vendor on the date checked, currently September 2026, with a link to the page it came from. Prices change, and every product page says so.

2. Reference setups

The benchmark tests products on its own reference setups, never on a client's installation. A reference setup is an account on the product, configured for a fictional business with published hours, services, prices and booking rules, so every answer the receptionist gives can be checked against a known truth.

Voice AI platforms such as Retell AI, Bland AI, Vapi and Synthflow are builder tools, not finished receptionists. For those, the benchmark scores a named reference receptionist built on the platform, and the platform's page discloses exactly which one.

3. Test calls

Every test call is placed by a real person. Callers work from written scenarios that cover business hours and after hours, first time and returning callers, and simple and messy requests: interruptions, changed minds, background noise, off script questions and requests for a human. Testers take notes on every call and record what the receptionist said in the moments that matter.

Test callers always use fictional identities. For medical and dental setups, no real patient information is ever used.

4. Scoring

Each call is scored in every rubric category it touches, with one of three results:

  • Pass The receptionist handled it the way a caller and the business would want.
  • Partial Part of it went well, part did not, and the note says which.
  • Fail The caller did not get what they needed, or was told something untrue.
  • Not yet scored No real test calls have been placed yet.

Every state carries a text label, never color alone. A product's published result lists the number of calls placed, the date of the last call and the reference setup used.

5. The rubric

The benchmark uses the same nine categories as our client tests, described in full on what we test.

The nine rubric categories
CodeCategoryWhat it covers
R1Greeting and identificationGreeting and identification of the business.
R2Understanding intentUnderstanding the caller's intent, including accents, interruptions and changed minds.
R3Booking accuracyBooking and scheduling accuracy against the business's real availability.
R4Handoff to a humanHandoff to a human when a caller asks for one or when the situation requires it.
R5After hours and urgent callsAfter hours behavior, emergencies and urgent requests.
R6Edge casesEdge cases: off script questions, wrong numbers, spam, angry callers.
R7Invented informationHallucinated or invented information: prices, policies, hours or services that are not real.
R8Latency and talk-overLatency, silences and talk-over.
R9Message accuracyMessage taking and the accuracy of what reaches the business afterward.

6. Neutrality rules

The rules that keep every profile factual and unranked.

  • The benchmark accepts no payment, affiliate commission or sponsorship from any vendor.
  • No product page carries an affiliate link. Vendor links go to the vendor's homepage and are marked nofollow.
  • Products are listed alphabetically or by category, never by quality, until real scores exist.
  • No star ratings, winners, “best overall” badges, or pros and cons that imply testing we have not done.
  • We do not sell AI receptionists, and we do not sell configuration or fixes, so there is no product for the benchmark to favor.
  • Client test results are never published. The benchmark uses its own reference setups only.

7. Current status

Testing is in progress. No test calls have been placed to any benchmark product yet, so every product shows as not yet scored. Scores will be added product by product as real calls are completed, each with its reference setup disclosed.

Back to the benchmark

Hear your AI receptionist the way your customers do

Tell us the number your AI receptionist answers and anything you want us to look at. Real people place the calls, and you receive a written report.