Custom AI Agents

Voice AI for Business: An Honest Platform Comparison

Rajat Gautam••12 min read•Updated
Share

Key Takeaways

  • →Advertised per-minute rates are a floor, not the bill: Retell's $0.055/min excludes the language model, telephony and text-to-speech.
  • →Vapi claims under 500ms latency, while an independent benchmark measured a 1,558 ms median over a real phone line.
  • →Synthflow moved from self-serve plans to enterprise-only contracts starting at $30,000 annually, so a pricing model can change entirely.
  • →Confirm the platform connects to your existing phone system, because a custom SIP trunk or a live contact center contract often decides the choice.
  • →No platform ships a working human handoff by default, so build and test the escalation path with a real off-script call before launch.
Voice AI for Business: An Honest Platform Comparison

Voice AI for business means software that answers or makes phone calls and talks like a person instead of a phone tree. It works by chaining four pieces together: speech-to-text, a language model, text-to-speech, and an orchestration layer that manages the back-and-forth. This guide compares the named platforms on pricing, latency, telephony, and what happens when a caller says something the script did not expect, using each vendor's own published pricing and one independent latency benchmark, not vendor claims taken at face value.

Every vendor in this category says the same three things: real-time, human-like, easy to set up. None of that tells you which one to buy for a call center, a booking line, or a support desk. What matters is underneath: what you pay per minute once the marketing tier runs out, how it connects to a real phone line, and what the caller hears after they say something the flow did not expect.

How voice AI works

Speech-to-text (STT) converts the caller's spoken words into text in real time. Engines used across these platforms include Deepgram, AssemblyAI, and Whisper.

A large language model (LLM) processes that text and generates a response: it follows the conversation, answers from a knowledge base, and decides what to do next, such as booking an appointment or transferring the call.

Text-to-speech (TTS) converts the model's response back into speech. Providers used across these platforms include ElevenLabs, PlayHT, and Cartesia.

An orchestration layer manages the live conversation: handling interruptions, deciding when to speak and when to listen, and connecting to outside systems like a CRM or a calendar.

How fast that whole loop runs, and how much slower it is in practice than the marketing page claims, is covered in the latency section below.

The platforms, in plain terms

  • Vapi and Retell AI are developer-first platforms. You assemble your own agent from a speech-to-text engine, a language model, and a text-to-speech voice, and pay for each piece separately.
  • Bland AI is a similar assemble-it-yourself platform with a more locked-down, all-in-one per-minute rate and its own telephony infrastructure.
  • Synthflow used to sell direct, self-serve plans and has moved to enterprise-only contracts.
  • ElevenLabs' Conversational AI (branded "ElevenAgents") is the consumer-to-enterprise route: a free tier, then paid tiers, built on ElevenLabs' own text-to-speech engine.
  • PolyAI is the enterprise-only option built for regulated industries handling complaint-heavy, compliance-heavy calls, and does not publish pricing at all.

Pricing, from each vendor's own pricing page

Every figure below was read directly off the named vendor's current pricing page. Where a vendor does not publish a number, this says so instead of guessing.

PlatformEntry costPer-minute rateWhat is bundledSource
VapiNo fixed platform fee on the Build plan, usage-based$0.05/min hosting, plus the actual cost of whichever speech and language models you pick (or free if you bring your own API keys)Call hosting, 10 concurrent calls included (extra lines $10/line/mo). No bundled minutes: minutes are usage-basedvapi.ai/pricing
Retell AIPay-as-you-go, no minimum and no platform fee$0.055/min infrastructure, plus text-to-speech ($0.015 to $0.040/min) and the language model you choose (as low as $0.003/min on a small model, up to $0.16/min on a top-tier one), plus $0.015/min US telephony20 free concurrent calls, custom SIP trunking at no extra chargeretellai.com/pricing
Bland AIFree tier, no card, 2 credits plus one inbound number$0.14/min (Start), $0.12/min (Build, $299/mo platform fee); Enterprise is customSpeech-to-text, language model, and voice bundled into the per-minute ratebland.ai/pricing
SynthflowNone published. "Enterprise contracts start at $30,000 annually"Not disclosed, scoped per contractCall volume, concurrency, telephony, integrations, security, and launch support, all scoped per dealsynthflow.ai/pricing
ElevenLabs Conversational AIFree tier: $0/month, 15 minutes$0.08/min standard overage, $0.16/min if you exceed your plan's concurrency ("burst"). Paid tiers run $6/month (75 min) up to $990/month (12,375 min) before EnterpriseText-to-speech, speech-to-text, knowledge base, and telephony connection. The language model and telephony carrier are billed separately at their own usage rateselevenlabs.io/pricing/agents
PolyAINot published. Demo request onlyNot publishedNot publishedpoly.ai

Read the Retell row carefully. Its headline infrastructure rate of $0.055/min is not what you pay. You pay that plus a language model fee that varies by roughly 50 times depending on which model you pick, plus telephony, plus text-to-speech. A platform's advertised per-minute number is the floor, not the bill.

Synthflow is the clearest lesson here: a pricing model can change entirely. A self-serve monthly plan today can become an enterprise-only annual contract by the time you are ready to buy. Check current pricing directly before budgeting against anything you read here.

Latency: what vendors claim versus what gets measured

Latency is the gap between the caller finishing a sentence and the agent starting to reply. Past roughly a second, callers start talking over the agent or hang up, because the pause reads as the call being broken, not as the system thinking.

Vendor marketing pages state very optimistic numbers. Vapi's own site claims "<500ms average latency" (vapi.ai). Synthflow's site claims "sub-100 ms latency" (synthflow.ai). Neither states the conditions, such as which model, which voice, or how many concurrent calls, that number was measured under.

An independent benchmark tells a different story. Openbenchmarks' voice agent latency benchmark (last call placed 2026-08-01) calls each platform's live agent over a real phone line and measures time to first audio byte, the silence between the caller stopping and the agent's audio starting, across several hundred turns per platform:

PlatformMedian (p50)95th percentile (p95)
Telnyx1,296 ms1,856 ms
ElevenLabs Conversational AI1,424 ms1,768 ms
Bland AI1,520 ms2,248 ms
Vapi1,558 ms2,008 ms
Retell AI1,740 ms2,259 ms

Synthflow and PolyAI were not in this benchmark run, so no independently measured figure exists for them here. Check current published benchmarks before trusting their marketing numbers under real call conditions.

Vapi is the one platform here with both a published claim and an independent figure, and the gap is large: a measured median of 1,558 ms against a marketing claim of under 500ms. A separate benchmark firm, DestiLabs, reviewing latency across a mix of platform deployments, put it plainly: "The platform is not your latency. As the results show, the same engine lands anywhere from ~560 ms to ~810 ms p50 depending on how the loop, tools, and streaming are built." The platform sets a floor. How it is configured, how many tools it calls mid-conversation, and how it streams audio decide where the number actually lands for a specific business.

Telephony integration

How a platform connects to an actual phone number matters more than most buyers realize before they try to switch providers later.

  • Vapi lists Twilio among its integrations for bringing in numbers, alongside its own hosted numbers.
  • Retell AI charges standard US telephony at $0.015/min and does not charge extra for a custom SIP trunk, which matters if a business already has a telephony contract it does not want to abandon.
  • Bland AI supports Twilio and SIP trunk connections, and separately lists direct integrations with Genesys, Five9, NICE CXone, Talkdesk, and Amazon Connect.
  • Synthflow runs its own carrier infrastructure and also connects to existing SIP trunks and PBX systems, naming Cisco, Avaya, Genesys, and RingCentral as supported.
  • PolyAI is built to sit inside large existing phone estates in regulated industries (published verticals: consumer services, financial services, healthcare, hotels, insurance, restaurants, retail, telecom, and utilities), but the integration detail is a sales conversation, not a published spec.

If a business already runs a contact center platform, that existing contract is often the deciding factor, not the vendor's feature list. A platform that cannot cleanly attach to the current phone system means replacing more than the AI.

Where voice AI actually earns its keep

These are the call patterns where a platform above is worth deploying, not a claim about what any specific deployment will produce.

  • Appointment booking. The agent checks live availability, books the slot, and sends a confirmation, without the caller waiting on hold. Works for dental offices, medical practices, salons, auto repair shops, HVAC companies, and legal consultations.
  • Support and FAQ handling. Business hours, order status, return policy, pricing questions: the predictable share of inbound calls that follow the same pattern every time. See deploying customer support agents for the broader pattern across channels, not phone alone.
  • Lead qualification. The agent asks the same screening questions every time, scores the response, and books a meeting only for prospects who meet the criteria, so a sales team is not spending time on unqualified calls.
  • Outbound appointment reminders and confirmations. The agent calls ahead of the appointment, confirms or reschedules based on the answer, and updates the scheduling system.
  • After-hours call handling. For plumbers, electricians, property managers, and veterinary clinics, the agent assesses urgency, dispatches for genuine emergencies, and books next-available slots for everything else.
  • Replacing a menu-based IVR. "Press 1 for sales, press 2 for support" becomes a conversation: the caller states what they need and the agent routes on the content of what was said, not a button press.

What happens when the caller says something unexpected

This is the question a vendor demo call is built to avoid, because a demo script never goes off script. A live caller does.

Bland AI's own product pages describe the agent as designed to "handle interruptions, accents, and topic changes" and to route difficult cases to a human "with full context" so the caller does not repeat themselves. Retell AI and Vapi both expose call transfer as a configurable action inside the agent builder, meaning the handoff to a human is something a business designs and tests itself, not something that happens correctly by default. None of the platforms above publish an independent, third-party figure for how often a live call actually gets correctly escalated versus mishandled. Treat any specific percentage quoted for that as unverified until the underlying test is visible.

The practical takeaway: assume every voice AI agent will eventually meet a caller who does not follow the expected path, and build the escalation path before launch. Test it by having a real person call in and deliberately go off script. If that test is not in a vendor's sales materials, ask for it directly.

Setting one up: the shape of the work

Regardless of platform, the deployment work follows the same shape:

  1. Pick one use case. Appointment booking, after-hours handling, or FAQ answering, not all three at once.
  2. Build the knowledge base. The frequently asked questions, hours, pricing, and policies the agent needs, written in the plain language it will speak aloud.
  3. Map the conversation. How the agent greets the caller, what it asks for each likely reason someone calls, and exactly when it hands off to a human.
  4. Configure and test. Pick the language model and voice, wire up the calendar or CRM integration, and call the agent dozens of times deliberately trying to break it: interruptions, background noise, edge cases.
  5. Deploy gradually. Start with after-hours calls only, listen to every transcript for the first stretch, and expand to peak hours once the agent is handling the routine cases correctly.

For the broader category this sits inside, see what AI agents are, and for cost comparisons across agent types beyond voice, see the AI agent cost and pricing guide.

Voice AI platforms serving healthcare, legal, and financial services offer HIPAA-compliant and SOC 2 configurations on request. What to ask a vendor for before any protected data touches the system: encrypted call recordings and transcripts, a signed Business Associate Agreement, access controls on call data, automatic redaction of personal information from logs, and a stated retention and deletion policy. Ask for the vendor's compliance documentation directly rather than taking a sales page's word for it.

When voice AI is the wrong choice

This is the section a vendor selling voice AI has no reason to write.

  • Call volume is too low to justify the setup. If a line only takes a handful of calls a week, the time spent designing prompts, testing escalation paths, and monitoring the agent may cost more than answering the calls directly.
  • The conversation depends on judgment calls with real consequences. Anything close to a medical, legal, or safety decision belongs with a person who can be held accountable for it, with AI assisting rather than replacing them.
  • The calls are outbound to people who did not ask to be called. Outbound calling in the US sits inside telemarketing law, the TCPA and related state rules, which governs consent and do-not-call records regardless of whether the caller is a person or a model. Get real legal guidance before an outbound voice AI campaign, not just a vendor's assurance that it is compliant.
  • The phone system cannot be touched. If IT will not authorize a new SIP trunk or number, or the contact center contract has years left on it, the integration cost can exceed the AI cost.
  • The real problem is a written policy nobody follows, not a missing phone agent. An agent that reads out a policy correctly does not fix a policy that is unclear or contradictory in the first place. Fix the source document first.

Where this fits if you sell or support over the phone

Choosing between these platforms is an architecture decision, not just a vendor decision. It touches the CRM, the existing telephony contract, the escalation paths, and who is accountable when the agent gets something wrong. That evaluation, done before signing with any single vendor, is the kind of vendor-neutral review our AI Sales and CX service is built around: pricing model, telephony fit, and escalation design checked against a business's actual setup, not a vendor's demo script.

The short version

Read the per-minute rate on the vendor's own pricing page, not a summary of it, and expect the language model and telephony line items to change the total by several times over. Do not trust a vendor's own latency claim; check an independent benchmark where one exists, and expect real-world latency to land higher than the marketing page says. Confirm the platform connects to the actual phone system before anything else. And build and test the human handoff directly, because no platform ships that correctly by default.

Keep Reading

Learn how to deploy customer support agents across every channel, not phone alone. See what AI agents are for how this category fits the wider picture, and the AI agent cost and pricing guide for cost comparisons beyond voice. Ready to evaluate this against your own phone setup? Book a strategy call or explore AI Sales and CX.

Sources

Frequently Asked Questions

How does voice AI work for business calls?+
It chains four pieces together. Speech-to-text turns the caller's words into text, a language model reads that text and decides the response, text-to-speech turns the response back into audio, and an orchestration layer manages the live back-and-forth, including interruptions and handoffs to outside systems like a CRM or calendar.
What is the best voice AI platform?+
There is no single best platform, because the choice is an architecture decision rather than a vendor decision. It depends on what you pay per minute once the marketing tier runs out, how the platform attaches to your real phone system, and how you design and test the escalation path. Compare the published pricing and an independent latency benchmark against your own setup before signing with anyone.
How much does a voice AI agent cost per minute?+
Published rates vary widely and rarely cover the full bill. Vapi lists $0.05/min hosting plus the cost of your chosen speech and language models, Retell AI lists $0.055/min infrastructure plus language model, text-to-speech and $0.015/min US telephony, Bland AI lists $0.14/min on Start and $0.12/min on Build with a $299/mo platform fee, and ElevenLabs charges $0.08/min standard overage. Expect the language model and telephony line items to change the total by several times over.
When should a business not use voice AI?+
When call volume is too low to justify the setup, when the conversation depends on judgment calls with real medical, legal or safety consequences, when outbound calls go to people who did not ask to be called, when the existing phone system cannot be touched, or when the real problem is a written policy nobody follows. An agent that reads out an unclear policy correctly does not fix the policy.
How do you test a voice AI agent before launch?+
Assume every agent will eventually meet a caller who does not follow the expected path, and build the escalation path before launch. Test it by having a real person call in and deliberately go off script, with interruptions, background noise and edge cases. If that test is not in a vendor's sales materials, ask for it directly.

Ready to deploy a voice AI agent for your business? Let's design your phone automation.

Book a Strategy Call

About the Author

Rajat Gautam

Rajat Gautam

AI Engineer and Consultant

My work goes far beyond recommending tools - I design AI systems that integrate directly into your workflows, eliminate inefficiencies, and deliver measurable business impact. Every solution I build is tailored, practical, and built with long-term scalability in mind.

Need help with this?

Related Topics

Voice AI
Call Center
Phone Agents
Customer Service
Automation

Related Articles

Ready to transform your business with AI? Let's talk strategy.

Book a Free Strategy Call