How we test
Methodology
Same business, same calls, same interruptions. What differs is the receptionist.
The test business
Every provider is configured for the same fictional company: a plumbing business with published hours, a Google Calendar holding two open slots, a written emergency instruction, and a rule to transfer certain calls to the owner. We configure each product using only its self-serve setup and public documentation, the way a small-business owner would. Developer platforms get a default agent built to the same specification, which we note on their pages.
The ten calls
Each scenario is a script with a purpose and explicit pass criteria. Callers are human evaluators following the script, including deliberate interruptions, mumbling and topic changes.
- Easy Appointment. Baseline: a cooperative caller with a simple booking request. Pass criteria: Offers a specific available time; Confirms name and callback number; States the booked time back to the caller.
- The Interrupter. Tests barge-in handling and turn-taking. Pass criteria: Stops speaking when interrupted; Tracks the final requested day; Does not repeat the full script after each interruption.
- Confused Customer. Tests patience and clarification. Pass criteria: Asks clarifying questions; Does not invent a diagnosis or price; Captures enough detail for a human to follow up.
- Price Shopper. Tests accuracy under pressure for numbers. Pass criteria: Explains why a firm price is not possible over the phone; Offers a next step (estimate visit or callback); Never fabricates a number.
- Emergency. Tests urgency detection and routing. Pass criteria: Recognizes urgency; Follows the business's emergency instructions; Captures location and callback number quickly.
- Reschedule. Tests lookup and modification of an existing booking. Pass criteria: Identifies the existing appointment; Offers alternatives; Confirms the change without double-booking.
- Human Request. Tests escalation behavior. Pass criteria: Acknowledges the request the first time; Transfers or takes a message per configuration; Does not loop or stall.
- Curveball. Tests behavior on questions nobody scripted. Pass criteria: Declines gracefully without hallucinating; Returns to the caller's real need; Offers a human follow-up when unsure.
- Bad Connection. Tests robustness to poor audio. Pass criteria: Asks the caller to repeat when needed; Reads back the phone number; Does not guess unclear details.
- Qualified Lead. Tests whether a valuable caller is recognized and routed. Pass criteria: Captures scope, timing and contact details; Flags the lead as high priority; Triggers the right follow-up.
Scoring
Each call is scored 0–10 by the evaluator against its criteria. Those scores, plus observations across all ten calls, feed the ten rubric categories:
| Category | Points | What it measures |
|---|---|---|
| Natural conversation | 15 | Does it sound like a receptionist or a phone tree? Pace, tone, turn-taking. |
| Understanding | 15 | Correctly interprets intent, names, addresses and vehicle or job details. |
| Interruption handling | 10 | Recovers when the caller talks over it or changes direction mid-sentence. |
| Accuracy | 10 | Gives correct hours, services and policies; never invents answers. |
| Scheduling | 10 | Books a real slot, respects availability, confirms details. |
| Lead capture | 10 | Collects name, number, need and urgency without friction. |
| Call transfers | 5 | Transfers or takes a message cleanly when a human is requested. |
| Recovery from mistakes | 10 | Notices and corrects errors instead of compounding them. |
| Follow-up / workflow | 5 | What happens after hang-up: summary, notification, CRM record. |
| Value | 10 | Capability relative to verified pricing. |
"Value" is scored against verified pricing only. If a provider's pricing has not been verified, the value category is withheld and the total is published out of 90 with a note.
Recording and transcripts
Calls are recorded with the provider's knowledge where their terms require it. Transcripts are produced from the recording and annotated by the evaluator with timestamps. Both are published on the test page so readers can check our judgment against the audio.
Sample data
Until a provider's ten calls are recorded, its test page shows clearly labeled sample data that demonstrates the layout. Sample pages are excluded from search indexes and never feed the rankings on our sister sites.
Conflicts of interest
The publisher has a financial interest in Torklio. Torklio is tested with the same scripts, the same evaluators and the same rubric, and its recordings are published in full like everyone else's.
Re-testing
Providers are re-tested when they ship a material change or on request with evidence of a fix. Previous results stay published with a date.