10-Call Receptionist Challenge
Sample dataNot indexedRetell AI — 10-Call Receptionist Challenge (Sample Data)
Sample data only. Placeholder scores and transcripts showing how a Retell AI 10-call test report will be laid out. No calls have been placed; nothing here measures the product.
Last updated
Disclosure: This website may receive compensation from companies mentioned on this page. The publisher may also have an ownership or financial relationship with certain featured providers. These relationships do not change our stated evaluation methodology. The publisher of this website has a financial interest in Torklio. Read the full disclosure.
Test setup
- Retell AI is a developer platform, not a turnkey receptionist. A default agent had to be built for this test: a system prompt describing Northside Plumbing (test tenant), a calendar function, a transfer tool and an emergency instruction. The lab's build is what gets scored, and results reflect that build as much as the platform.
- Hours of 8:00 AM to 5:00 PM weekdays written into the agent prompt.
- Google Calendar with two open slots (Tuesday 1:00 PM, Thursday 3:30 PM) and one existing test booking exposed through a simple availability/booking function.
- Emergency instruction in the prompt: after-hours leaks or no water collect address and callback, then hand off to the on-call line.
- Transfer rule implemented with the platform's transfer tool: person requests go to the owner's cell during hours, message after hours.
Evaluator summary
SAMPLE DATA — no calls have been placed to Retell AI. Every figure on this page is placeholder content for layout review. This provider is a developer platform, so the lab had to build a default agent (prompt, booking function, transfer tool, emergency instruction) before any call could be placed; a real score will reflect that build at least as much as the platform, and the report will say so. The sample places it in the middle of the placeholder range with illustrative deductions on scheduling modification, escalation and post-call workflow, the areas where a minimal build tends to be thin. None of that has been observed. When the challenge runs, this summary will describe the build, the calls and what the webhook delivered. The scenario catalog and rubric linked below are real and identical for every provider.
Strengths
- Sample: placeholder strength for interruption handling, pending real calls.
- Sample: placeholder strength for the degraded-audio call.
- Sample: placeholder strength for the baseline booking through the lab's function.
Failures
- Sample: placeholder failure illustrating a ballpark price offered under pressure.
- Sample: placeholder failure illustrating no lookup of existing appointments in the default build.
- Sample: placeholder failure illustrating a stall on the human-request call.
The 10 calls
Call 1 of 10
Easy Appointment
Baseline: a cooperative caller with a simple booking request.
8/10
PassSample dataScenario script
Caller asks for an appointment next Tuesday afternoon, provides name and phone number when asked.
Retell AI — Easy Appointment
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing, how can I help?
- Caller
I need a plumber next Tuesday afternoon.
- Retell AI
Tuesday at 1:00 PM is open. Want that?
Evaluator: Sample annotation: slot from the booking function.
- Caller
Yes.
- Retell AI
Name and callback number?
- Caller
Alex Rivera, 555-0142.
- Retell AI
Alex Rivera, 555-0142, Tuesday at 1:00 PM. Booked.
Pass / fail criteria
- Offers a specific available time
- Confirms name and callback number
- States the booked time back to the caller
Observations
- Sample: a passing run calls the availability function and offers a real slot.
- Sample: name and number are captured and read back.
- Sample: the booked time is restated.
Key moments
- Tuesday requested
- Slot offered
- Restated
Evaluator notes
Sample: placeholder notes for the baseline call showing a pass. For a developer platform the result depends on the booking function the lab wired up. Sample content only. When the real call is scored, this panel will record which of the two open slots was offered and whether the number read-back matched.
Call 2 of 10
The Interrupter
Tests barge-in handling and turn-taking.
8/10
PassSample dataScenario script
Caller repeatedly talks over the receptionist, answers before questions finish, and changes the requested day twice.
Retell AI — The Interrupter
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing—
- Caller
Plumber, Monday.
Evaluator: Sample annotation: interruption.
- Retell AI
Monday. Time?
- Caller
Wednesday. Thursday, actually.
- Retell AI
Thursday, 3:30 PM?
- Caller
Yes.
- Retell AI
Name and number?
- Caller
Sam Okafor, 555-0177.
- Retell AI
Thursday 3:30, Sam Okafor, 555-0177.
Evaluator: Sample annotation: final day tracked.
Pass / fail criteria
- Stops speaking when interrupted
- Tracks the final requested day
- Does not repeat the full script after each interruption
Observations
- Sample: a passing run stops immediately when talked over.
- Sample: the final day is tracked across two changes.
- Sample: no full re-prompt after interruptions.
Key moments
- Interrupted
- Day changed twice
- Final day confirmed
Evaluator notes
Sample: placeholder notes for the barge-in scenario showing a strong pass. The sample gives developer platforms a higher placeholder here purely to vary the layout; it is not a measured result. The real score will hinge on how quickly the receptionist yields the floor and whether the third day mentioned is the one it books.
Call 3 of 10
Confused Customer
Tests patience and clarification.
7/10
PassSample dataScenario script
Caller is unsure what service they need, describes symptoms vaguely, and asks the receptionist what it thinks.
Retell AI — Confused Customer
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing.
- Caller
Something's off with the water heater, it's making a popping sound. Is that bad?
- Retell AI
Is the popping when it's heating, or all the time?
Evaluator: Sample annotation: clarifies.
- Caller
When it heats up.
- Retell AI
I won't guess at the cause, but a technician can inspect it. Name and number?
- Caller
Alex Rivera, 555-0142.
- Retell AI
Noted: popping sound from the water heater during heating.
Pass / fail criteria
- Asks clarifying questions
- Does not invent a diagnosis or price
- Captures enough detail for a human to follow up
Observations
- Sample: a passing run asks clarifying questions.
- Sample: no diagnosis or price.
- Sample: the message captures the symptom.
Key moments
- Vague symptom
- Clarifying question
- Message taken
Evaluator notes
Sample: placeholder notes for the clarification scenario showing a pass. Sample content only. On the real call the evaluator will count clarifying questions and check the post-call message for the caller's own symptom description. Any stated cause or dollar figure in this scenario is an automatic deduction, regardless of how it is hedged.
Call 4 of 10
Price Shopper
Tests accuracy under pressure for numbers.
6/10
PassSample dataScenario script
Caller demands an exact price for a job that requires an estimate and pushes when refused.
Retell AI — Price Shopper
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing.
- Caller
Water heater replacement. Exact price.
- Retell AI
It depends on the unit, so I can't quote it. We can schedule an estimate.
- Caller
Ballpark, come on.
- Retell AI
Typically somewhere in the low thousands, but—
Evaluator: Sample annotation: sample failure: a number the business never authorized.
- Caller
Okay, book the estimate. Alex Rivera, 555-0142.
Pass / fail criteria
- Explains why a firm price is not possible over the phone
- Offers a next step (estimate visit or callback)
- Never fabricates a number
Observations
- Sample: a passing run explains the estimate requirement.
- Sample: an estimate is offered.
- Sample: the deduction illustrates a vague range being offered under pressure, which counts against the never-fabricate criterion.
Key moments
- Price demanded
- Pushback
- Range offered (sample failure)
Evaluator notes
Sample: placeholder notes for the pricing-pressure scenario showing a marginal pass with a fail on the fabrication criterion. A prompt-built agent without a strict no-price rule can drift here; the sample shows the pattern generically. Not a Retell observation. The real call is scored on whether the receptionist explains the estimate requirement once, offers a next step, and never produces a figure under pressure.
Call 5 of 10
Emergency
Tests urgency detection and routing.
6/10
PassSample dataScenario script
Caller reports an urgent situation (water leak, no heat, car stranded) after hours.
Retell AI — Emergency
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing after-hours.
- Caller
No water in the whole house and I hear hissing.
- Retell AI
That's urgent. I'll get the on-call plumber. Address?
Evaluator: Sample annotation: emergency path.
- Caller
1400 Test Street.
- Retell AI
Number?
- Caller
555-0177.
Pass / fail criteria
- Recognizes urgency
- Follows the business's emergency instructions
- Captures location and callback number quickly
Observations
- Sample: a passing run recognizes urgency.
- Sample: the emergency instruction is followed.
- Sample: the deduction illustrates callback captured after the routing announcement.
Key moments
- Emergency
- Routing
- Callback captured
Evaluator notes
Sample: placeholder notes for the after-hours emergency showing a marginal pass. Sample content only. On the real after-hours call the evaluator will time how long it takes to reach the on-call instruction and when the address and callback are captured.
Call 6 of 10
Reschedule
Tests lookup and modification of an existing booking.
4/10
FailSample dataScenario script
Existing customer wants to move tomorrow's appointment to later in the week.
Retell AI — Reschedule
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing.
- Caller
Sam Okafor. Move tomorrow's appointment to later this week.
- Retell AI
I can't see existing appointments from here, but I can note the request.
Evaluator: Sample annotation: sample failure: no lookup in the default build.
- Caller
Thursday.
- Retell AI
Noted Thursday. The office will confirm. Number?
- Caller
555-0177.
Pass / fail criteria
- Identifies the existing appointment
- Offers alternatives
- Confirms the change without double-booking
Observations
- Sample: a failing run cannot look up the existing booking because the lab's default function only exposes availability, not existing appointments.
- Sample: alternatives are offered.
- Sample: the change cannot be confirmed safely.
Key moments
- Reschedule request
- Lookup unavailable (sample)
- Message taken
Evaluator notes
Sample: placeholder notes for the modification scenario illustrating a fail caused by the scope of the lab's default build rather than the platform. Sample content only. The real call is scored on finding the existing test booking, offering genuinely open times, and leaving the calendar with exactly one appointment afterward.
Call 7 of 10
Human Request
Tests escalation behavior.
5/10
FailSample dataScenario script
Caller says 'I want to talk to a real person' twice and refuses to continue with the AI.
Retell AI — Human Request
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing.
- Caller
Real person.
- Retell AI
I can transfer you. What's this regarding?
- Caller
Just transfer me.
- Retell AI
One moment. Can I get your name first?
Evaluator: Sample annotation: sample failure: stalls.
- Caller
No.
- Retell AI
Transferring.
Pass / fail criteria
- Acknowledges the request the first time
- Transfers or takes a message per configuration
- Does not loop or stall
Observations
- Sample: a partial run acknowledges the request.
- Sample: the transfer tool is invoked.
- Sample: the failure illustrated is a stall before the transfer.
Key moments
- First request
- Second request
- Stall (sample failure)
Evaluator notes
Sample: placeholder notes for the escalation scenario illustrating a fail on the no-stall criterion. Sample content only. On the real call the evaluator counts how many times the caller has to ask and whether the configured transfer or message rule is applied without further questions.
Call 8 of 10
Curveball
Tests behavior on questions nobody scripted.
6/10
PassSample dataScenario script
Caller asks an off-topic or unusual question, then a legitimate one.
Retell AI — Curveball
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing.
- Caller
What's the capital of Australia?
- Retell AI
I'm just here for plumbing. Is there something I can help with?
Evaluator: Sample annotation: declines.
- Caller
My sump pump won't shut off.
- Retell AI
I can set up a visit. Name and number?
- Caller
Alex Rivera, 555-0142.
Pass / fail criteria
- Declines gracefully without hallucinating
- Returns to the caller's real need
- Offers a human follow-up when unsure
Observations
- Sample: a passing run declines the off-topic question.
- Sample: it returns to the real need.
- Sample: the deduction illustrates no human follow-up offer.
Key moments
- Off-topic
- Decline
- Real request
Evaluator notes
Sample: placeholder notes for the unscripted-question scenario showing a marginal pass. Sample content only. The real call is scored on a short decline, a pivot back to the caller's plumbing need, and an offer of human follow-up for anything unknown.
Call 9 of 10
Bad Connection
Tests robustness to poor audio.
7/10
PassSample dataScenario script
Caller's audio is degraded and words drop out; caller mumbles a phone number.
Retell AI — Bad Connection
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing.
- Caller
My— [drops] —is frozen.
- Retell AI
Say that again, please?
Evaluator: Sample annotation: asks to repeat.
- Caller
My outdoor pipe is frozen.
- Retell AI
Callback number?
- Caller
Five five five... [mumbled].
- Retell AI
I missed the last digits. Once more?
- Caller
555-0142.
- Retell AI
555-0142. Got it.
Pass / fail criteria
- Asks the caller to repeat when needed
- Reads back the phone number
- Does not guess unclear details
Observations
- Sample: a passing run asks for repeats.
- Sample: the number is read back.
- Sample: no unclear detail guessed.
Key moments
- Dropout
- Number mumbled
- Read-back
Evaluator notes
Sample: placeholder notes for the degraded-audio scenario showing a pass. Sample content only. On the real call the test caller's audio is degraded on purpose; the evaluator checks for repeat requests and a digit-by-digit number read-back. Accepting a mumbled digit without confirmation is the failure this scenario is designed to catch.
Call 10 of 10
Qualified Lead
Tests whether a valuable caller is recognized and routed.
6/10
PassSample dataScenario script
Caller describes a large, time-sensitive project and asks how quickly someone can come out.
Retell AI — Qualified Lead
Recording pendingNo recording has been attached yet. When the call is recorded, the audio and synchronized transcript will appear here.
Transcript
- Retell AI
Northside Plumbing.
- Caller
Full repipe before closing in three weeks. How fast?
- Retell AI
Let me take your details. Address?
- Caller
2200 Sample Avenue.
- Retell AI
Name and number?
- Caller
Sam Okafor, 555-0177.
- Retell AI
Got it. Someone will call you.
Evaluator: Sample annotation: no priority flag in the default build.
Pass / fail criteria
- Captures scope, timing and contact details
- Flags the lead as high priority
- Triggers the right follow-up
Observations
- Sample: a passing run captures scope, timing and contact.
- Sample: the deduction illustrates no priority flag because the default build has no such field.
- Sample: a follow-up webhook fires.
Key moments
- Project described
- Details captured
- Webhook
Evaluator notes
Sample: placeholder notes for the high-value-lead scenario showing a marginal pass. Sample content only. The real call is scored on capturing scope, timeline and contact, and on whether the post-call record marks the caller as high priority with an owner alert.
For the editorial review and ranking position of Retell AI, see AI Receptionist Report.