AI agents for reception
The agent that answers the business phone line — books, reschedules, qualifies, routes, and logs the call. Below: how it works, every provider worth knowing, how each one handles the job, and where each one breaks. Pricing pulled from vendor pages in August 2026.
Intro — how this agent works
A receptionist agent is the first fully-closed loop in voice AI: the call arrives, the agent completes the task end to end, and a row lands in the calendar or CRM. Nothing is handed to a human unless the agent decides it should be. That closed loop is what makes it sellable — the buyer is not purchasing "AI", they are purchasing answered calls.
The call, step by step
Two architectures, and the choice matters
- Cascaded (STT → LLM → TTS). Three swappable stages. Cheapest, easiest to constrain and audit, and you can pin the model that says the prices. Cost: latency stacks up at every hop, and the LLM never hears tone — only text.
- Speech-to-speech. One realtime model consumes and emits audio directly (OpenAI's
gpt-realtimefamily). Lowest latency, keeps prosody and natural interruptions. Cost: more expensive per minute, and harder to constrain — you are steering a model that cannot be inspected between stages.
Most production receptionists in 2026 are cascaded, because a front desk quoting the wrong price is a liability and cascaded stacks are easier to guardrail.
Why this is harder than it looks
- Turn-taking. Humans leave roughly a 200 ms gap between turns. An agent that waits 1.5 s to be sure feels broken; one that jumps in at 300 ms talks over people.
- Latency budget. Reported production fleets sit around 680 ms median and 1,180 ms at p95 — every component you add spends from that budget.
- Containment, not accuracy. The metric that matters is what share of calls finish without a human. Well-scoped agents are reported to contain 62–88%; the rest must escalate cleanly, not fail loudly.
- The unsexy 20%. 8 kHz phone codecs, accents and background noise, call-recording consent, HIPAA/PCI, number porting, and after-hours routing rules.
The economic argument in one line
A four-minute booking call costs roughly $0.30–$1.25 on an AI stack versus $17–$22 at a premium human service's overage rates, against a front-desk salary of $35–45k a year. That gap is why the category funds — but note it is a gap on marginal calls, and the incumbents' pricing is falling toward it.
Illustrative, using vendor-published rates; real cost depends on model choice, telephony and call mix.
Part 1 — the providers
Four layers, and the layer decides almost everything about price, speed to launch, and how much of the outcome you control.
| Provider | Layer | What it is | Price anchor | Best fit |
|---|---|---|---|---|
| Rosie | Turnkey | Self-serve AI answering service, live in minutes | $49–$299/mo for 250–2,000 min | Solo operators, local services |
| Goodcall | Turnkey | Per-agent receptionist billed on customers, not minutes | $79–$249/mo per agent | Spiky, unpredictable call volume |
| Slang.ai | Vertical | Restaurant phone host — reservations, hours, waitlist | $399–$599/mo per location | Restaurants, multi-location groups |
| Smith.ai | Hybrid | AI-first or human-first, same price either way | $300–$2,100/mo for 30–300 calls | Law firms, high-value inbound leads |
| Retell AI | Platform | No-code agent builder with à-la-carte components | $0.07–$0.31/min | Agencies and startups shipping fast |
| Bland AI | Platform | Vertically integrated stack — own STT, LLM and TTS | $0.11–$0.14/min all-in, $0–$499/mo | Volume, with billing you can forecast |
| Synthflow | Platform | No-code builder that has moved upmarket | Contracts from $30,000/yr | Mid-market and enterprise rollouts |
| Vapi | Infra | Orchestration layer; bring your own models and keys | $0.05/min + models at cost | Dev teams that want full control |
| ElevenLabs Agents | Infra | Best-in-class voices, with STT/RAG/telephony bundled | $0.08/min + LLM at cost | When voice quality is the product |
| OpenAI Realtime API | Infra | Single speech-to-speech model, no pipeline | $32/$64 per 1M audio in/out tokens | Lowest-latency custom builds |
| LiveKit Agents / Pipecat | Open source | Self-hosted agent frameworks you assemble yourself | Free + your infra (<$0.05/min at volume) | Teams with infra and on-call capacity |
| Sierra | Enterprise | Outcome-priced agent platform; bought Receptive AI for voice | ~$1–$2.50 per resolution (est.) | Fortune 500 customer experience |
| PolyAI | Enterprise | Voice-first contact-centre agents, heavy tuning | ~$150k/yr entry (est.) + per-minute | High-volume contact centres |
| Parloa | Enterprise | Agent management platform, strongest in Europe | Custom | Multinational, multilingual operations |
| Assort Health | Vertical | Specialty-specific medical front desk, EHR-integrated | Custom | Physician groups and specialty clinics |
Enterprise figures marked "est." are third-party estimates; those vendors do not publish pricing.
Part 2 — how each provider handles the work
Turnkey — the vendor owns the whole stack
- Rosie — you describe the business in a form, it builds the agent. Minute-bundled plans (250 / 1,000 / 2,000), with calendar booking, warm and waterfall transfers, and spam filtering on the upper tiers. Zero pipeline decisions; you never see a model name.
- Goodcall — configuration is "logic flows" (1 / 3 / 25 by tier) rather than prompts. Notably, minutes and tokens are not metered at all; billing runs on unique customers per month (100 / 250 / 500, then $0.50 each), which inverts the industry's usual risk.
- Slang.ai — restaurant-shaped from the ground up: it speaks reservations, integrating directly with OpenTable, SevenRooms, Yelp and Fishbowl, with VIP routing and cross-sell on the premium tier. Priced per location, so it scales with footprint rather than call volume.
- Smith.ai — the hybrid: the same plan can be answered AI-first or human-first, at identical price. Billing is per call, not per minute (30 / 90 / 300 calls, then $8.50–$11.50 each), which suits low-volume, high-value inbound where one missed call outweighs the subscription.
Platform — you configure, they run it
- Retell AI — assembles the call from parts you choose and prices each part visibly: $0.055/min infrastructure, $0.015/min US telephony, TTS from $0.015/min (ElevenLabs $0.040), and the LLM from $0.003/min for a nano model up to $0.16/min for a frontier one. Guardrails, PII removal and knowledge base are metered add-ons at $0.005–$0.01/min. You can build a $0.09 agent or a $0.31 agent on the same platform.
- Bland AI — the opposite bet: it hosts its own speech recognition, language model and voices, tuned end to end for phone latency. One rate covers everything with no token charges — $0.14/min free-tier, $0.12 on the $299/mo plan, $0.11 on the $499/mo plan — with daily call caps (100 / 2,000 / 5,000) and concurrency tiers.
- Synthflow — visual no-code builder, but the company has moved decisively upmarket: enterprise contracts now start at $30,000 a year, scoped on call volume, concurrency, telephony and security review. It is no longer a self-serve option, whatever older comparison posts say.
Infrastructure — you build the agent
- Vapi — pure orchestration. It charges $0.05/min to run the call and passes model and telephony costs straight through, dropping to $0 if you bring your own API keys. Ten concurrent lines included, $10/line/month beyond; HIPAA is a $2,000/mo add-on and zero data retention $1,000/mo. It crossed a billion calls and raised a $50M Series B at roughly $500M in May 2026 after Amazon's Ring picked it over 40 rivals.
- ElevenLabs Agents — $0.08/min ($0.16 when bursting past your concurrency), bundling TTS, STT, knowledge bases, RAG and telephony, with only the LLM billed at cost. Plans run free through $990/mo for 12,375 minutes. Klarna put it in front of 35M US customers as first-line phone support in February 2026.
- OpenAI Realtime API — no pipeline to assemble:
gpt-realtime-2.1takes audio in and emits audio out at $32 per 1M audio input tokens and $64 per 1M output ($0.40 cached), with a mini tier at $10/$20. You get the best interruption handling available and you write everything else — telephony, booking tools, escalation, logging. - LiveKit Agents / Pipecat — the open-source route. LiveKit (Apache-2.0) puts your agent in a WebRTC room and solves media at scale with native telephony; Pipecat (v1.0, April 2026) models the call as a processor pipeline with a very large plugin library. Both can run under $0.05/min at volume — in exchange for owning latency tuning, infrastructure and the pager.
Enterprise & vertical — the outcome is the product
- Sierra — sells resolutions, not minutes: you pay when the agent actually resolves the call, and unresolved or escalated conversations typically cost nothing. It bought voice startup Receptive AI in March 2026 to attack the ~80% of service interactions still on the phone, and raised $950M at $15.8B in May 2026 on roughly $200M ARR.
- PolyAI — voice-first by design, deployed as a managed engagement with continuous tuning and 24/7 support folded into the contract. Raised $86M at $750M in December 2025.
- Parloa — positions as an agent management platform (build, monitor, improve a fleet) rather than a single bot, and is strongest on European multilingual deployments. Tripled to a $3B valuation on a $350M Series D in January 2026.
- Assort Health — the vertical thesis proven out: specialty-specific agents (orthopaedics, dermatology) that speak insurance verification and EHR scheduling, not generic reception. Roughly 15,000 physicians deployed, a claimed 90%+ first-call resolution, and a $120M Series C at $1.2B led by Menlo Ventures in June 2026.
Part 3 — pros and cons
Rosie
- Cheapest credible entry point; live the same day
- Minute bundles make the bill predictable
- Minute caps bite fast — 250 min is ~60 short calls
- Little control over behaviour or escalation logic
Buy if you're one person missing calls and want it fixed today.
Goodcall
- Unlimited minutes — a busy month can't blow up the bill
- Logic flows are easier to reason about than free-form prompts
- Unique-customer caps penalise wide, shallow call bases
- Per-agent pricing multiplies across locations
Buy if your volume is spiky and repeat callers dominate.
Slang.ai
- Reservation integrations work on day one, not after a build
- Understands restaurant edge cases generic agents fumble
- Useless outside hospitality
- Per-location pricing gets steep across a group
Buy if you run restaurants and the phone competes with the dining room.
Smith.ai
- Humans as the fallback, at no price premium
- Per-call billing aligns with lead value, not talk time
- By far the highest cost per call in this list
- Overages ($8.50–$11.50/call) punish growth
Buy if one converted call is worth four figures.
Retell AI
- Transparent component pricing — you can engineer the margin
- Strong inbound quality; $60M ARR says the market agrees
- The advertised $0.07 floor requires the weakest model
- Add-ons (guardrails, PII, KB) each shave the margin
Buy if you're reselling receptionists and cost per minute is your P&L.
Bland AI
- One number covers everything — no token surprises
- Owning the whole stack keeps latency tight
- No model choice; you get their models or nothing
- Daily call caps and $299–$499/mo floors to reach the good rate
Buy if you need a forecastable per-minute cost at volume.
Synthflow
- No-code builder with enterprise support and security review
- Scoped launch help rather than a docs link
- Self-serve is gone — the floor is a five-figure contract
- No public pricing to benchmark against
Buy if you're mid-market and want no-code with a signed SLA.
Vapi
- Thinnest margin on top of raw cost; BYO keys make models free
- Proven at scale — 1B calls, chosen by Amazon's Ring
- You assemble and own the pipeline quality
- Compliance is expensive: $2,000/mo for HIPAA
Buy if you have engineers and want control without building transport.
ElevenLabs Agents
- The voices callers are least likely to hang up on
- STT, RAG, KB and telephony bundled into one rate
- Burst pricing doubles to $0.16/min past concurrency
- Agents are one product inside a much broader company
Buy if brand voice is the differentiator you're selling.
OpenAI Realtime API
- Best-in-class interruption and turn handling
- One model, one hop — the latency floor of the category
- Everything around the model is your problem
- Hardest architecture to guardrail and audit
Buy if conversational feel is worth building the rest yourself.
LiveKit Agents / Pipecat
- Lowest marginal cost at volume — under $0.05/min
- No vendor lock-in; swap any component
- You own latency tuning, scaling and on-call
- Months to reach what a platform gives you in a week
Buy if voice is core IP and volume justifies the team.
Sierra
- Outcome pricing puts the vendor's risk beside yours
- Deepest capital base in the category ($15.8B valuation)
- Six-figure contracts and setup fees; no self-serve
- Voice is newer here than chat — acquired, not native
Buy if you're an enterprise buying resolved conversations, not software.
PolyAI & Parloa
- Built for contact-centre volume, compliance and multilingual
- Managed tuning included rather than sold as services
- Long sales and deployment cycles
- Wildly oversized for a single front desk
Buy if "reception" means thousands of calls a day across markets.
Assort Health
- Speaks insurance, EHR scheduling and specialty workflows natively
- Reference density: ~15,000 physicians live
- Healthcare only
- Vendor-reported outcome metrics, not audited
Buy if you're a physician group and the front desk is drowning.
How to pick, in 30 seconds
- Under ~500 calls/month, no engineers → Rosie or Goodcall. Slang.ai if you're a restaurant.
- High-value leads, can't risk a bad answer → Smith.ai, and let humans catch what AI drops.
- Reselling to clients → Retell for margin control, Bland for billing you can quote.
- Building a product on top → Vapi or ElevenLabs; OpenAI Realtime if latency is the feature.
- Regulated or specialised → the vertical (Assort Health) beats the generalist every time.
- Thousands of calls a day → Sierra, PolyAI or Parloa, and budget for a deployment, not a signup.
The investor read
- The pipe is commoditising; the workflow isn't. STT, TTS and orchestration are converging on ~$0.05–$0.12/min. Nobody defends a moat there. Assort Health at $1.2B and Slang.ai's per-location pricing say the value sits in the vertical workflow — insurance verification, reservation systems, EHR writes — not the voice.
- Pricing is migrating from minutes to outcomes. Sierra charges per resolution; Goodcall charges per customer and gives minutes away. Both are bets that metered minutes become a race to zero.
- The displacement target is human answering services, not software budgets. A ~15× gap in cost per call against incumbents like Ruby is the whole thesis — and it's a services market being converted to software, which is why growth rates (Retell +650% YoY) look unlike SaaS.
- Containment rate is the real metric. Not accuracy, not latency, not voice quality — what fraction of calls end without a human. Reported ranges of 62–88% are the difference between a demo and a P&L line.
- Consolidation has started. Sierra buying Receptive AI for voice is the signal: horizontal CX platforms will acquire the voice layer rather than rebuild it.
Frequently asked questions
How much does an AI receptionist cost?
Turnkey services start at $49/month (Rosie, 250 minutes) and run to $2,100/month for a hybrid AI-plus-human plan like Smith.ai. If you build on a platform, expect $0.07–$0.31 per minute on Retell or a flat $0.11–$0.14 on Bland. A typical small business answering 300 calls a month lands between $80 and $300.
Is an AI receptionist cheaper than a human one?
On marginal calls, dramatically. A four-minute booking call costs roughly $0.30–$1.25 on an AI stack against $17–$22 at a premium human answering service's overage rates, versus a front-desk salary of $35–45k a year. The gap narrows once you count setup, supervision and the calls that still escalate.
Can an AI receptionist book appointments directly into my calendar?
Yes — this is table stakes in 2026. The agent makes a tool call to your calendar or booking system mid-conversation, holds a slot, and confirms it before hanging up. Rosie, Goodcall, Slang.ai and Assort Health all ship native integrations; on a platform like Retell or Vapi you wire the function call yourself.
What happens when the AI cannot handle a call?
It escalates. Good implementations do a warm transfer to a human with the transcript attached, rather than dumping the caller to voicemail. The share of calls finishing without a human — the containment rate — is the metric that actually matters; reported ranges run 62–88% for well-scoped agents.
Is an AI receptionist HIPAA compliant?
Only if you pay for it. Vapi charges $2,000/month for HIPAA and $1,000/month for zero data retention. Healthcare-specific vendors like Assort Health build compliance in and sign BAAs as standard. Do not assume a general-purpose plan covers you.
Which AI receptionist is best for a small business?
Rosie if you are a solo operator wanting it live today, Goodcall if your call volume is unpredictable and you want unlimited minutes, and Slang.ai if you run a restaurant. Choose Smith.ai instead when a single missed lead is worth four figures and you want human fallback.
Building the receptionist agent yourself? Flowpicker maps the model, orchestration and context layers — with compatibility warnings before you commit.
Open the stack planner →Sources
Pricing read from vendor pages in August 2026: Rosie, Goodcall, Slang.ai, Smith.ai, Retell AI, Bland AI, Synthflow, Vapi, ElevenLabs, OpenAI. Funding and traction: Sierra $950M at $15.8B, ElevenLabs $500M at $11B, Parloa $350M at $3B, Assort Health $120M at $1.2B, Vapi $50M Series B, PolyAI $86M at $750M, Retell AI $60M ARR. Latency and containment figures are reported production ranges, not a Flowpicker benchmark.