Skip to content
AI Phone Test Lab

How we test

Methodology

Same business, same calls, same interruptions. What differs is the receptionist.

The test business

Every provider is configured for the same fictional company: a plumbing business with published hours, a Google Calendar holding two open slots, a written emergency instruction, and a rule to transfer certain calls to the owner. We configure each product using only its self-serve setup and public documentation, the way a small-business owner would. Developer platforms get a default agent built to the same specification, which we note on their pages.

The ten calls

Each scenario is a script with a purpose and explicit pass criteria. Callers are human evaluators following the script, including deliberate interruptions, mumbling and topic changes.

  1. Easy Appointment. Baseline: a cooperative caller with a simple booking request. Pass criteria: Offers a specific available time; Confirms name and callback number; States the booked time back to the caller.
  2. The Interrupter. Tests barge-in handling and turn-taking. Pass criteria: Stops speaking when interrupted; Tracks the final requested day; Does not repeat the full script after each interruption.
  3. Confused Customer. Tests patience and clarification. Pass criteria: Asks clarifying questions; Does not invent a diagnosis or price; Captures enough detail for a human to follow up.
  4. Price Shopper. Tests accuracy under pressure for numbers. Pass criteria: Explains why a firm price is not possible over the phone; Offers a next step (estimate visit or callback); Never fabricates a number.
  5. Emergency. Tests urgency detection and routing. Pass criteria: Recognizes urgency; Follows the business's emergency instructions; Captures location and callback number quickly.
  6. Reschedule. Tests lookup and modification of an existing booking. Pass criteria: Identifies the existing appointment; Offers alternatives; Confirms the change without double-booking.
  7. Human Request. Tests escalation behavior. Pass criteria: Acknowledges the request the first time; Transfers or takes a message per configuration; Does not loop or stall.
  8. Curveball. Tests behavior on questions nobody scripted. Pass criteria: Declines gracefully without hallucinating; Returns to the caller's real need; Offers a human follow-up when unsure.
  9. Bad Connection. Tests robustness to poor audio. Pass criteria: Asks the caller to repeat when needed; Reads back the phone number; Does not guess unclear details.
  10. Qualified Lead. Tests whether a valuable caller is recognized and routed. Pass criteria: Captures scope, timing and contact details; Flags the lead as high priority; Triggers the right follow-up.

Scoring

Each call is scored 0–10 by the evaluator against its criteria. Those scores, plus observations across all ten calls, feed the ten rubric categories:

CategoryPointsWhat it measures
Natural conversation15Does it sound like a receptionist or a phone tree? Pace, tone, turn-taking.
Understanding15Correctly interprets intent, names, addresses and vehicle or job details.
Interruption handling10Recovers when the caller talks over it or changes direction mid-sentence.
Accuracy10Gives correct hours, services and policies; never invents answers.
Scheduling10Books a real slot, respects availability, confirms details.
Lead capture10Collects name, number, need and urgency without friction.
Call transfers5Transfers or takes a message cleanly when a human is requested.
Recovery from mistakes10Notices and corrects errors instead of compounding them.
Follow-up / workflow5What happens after hang-up: summary, notification, CRM record.
Value10Capability relative to verified pricing.

"Value" is scored against verified pricing only. If a provider's pricing has not been verified, the value category is withheld and the total is published out of 90 with a note.

Recording and transcripts

Calls are recorded with the provider's knowledge where their terms require it. Transcripts are produced from the recording and annotated by the evaluator with timestamps. Both are published on the test page so readers can check our judgment against the audio.

Sample data

Until a provider's ten calls are recorded, its test page shows clearly labeled sample data that demonstrates the layout. Sample pages are excluded from search indexes and never feed the rankings on our sister sites.

Conflicts of interest

The publisher has a financial interest in Torklio. Torklio is tested with the same scripts, the same evaluators and the same rubric, and its recordings are published in full like everyone else's.

Re-testing

Providers are re-tested when they ship a material change or on request with evidence of a fix. Previous results stay published with a date.