Vapi or Retell
Good first shortlist when you want phone-agent APIs, dashboards, and less media infrastructure to operate.
A practical comparison of hosted platforms and open frameworks, with public pricing normalized across the same phone-call workload.
There is no honest single winner. The useful split is how much of the stack you want the platform to own.
Good first shortlist when you want phone-agent APIs, dashboards, and less media infrastructure to operate.
Worth testing when one AI minute and predictable model pass-throughs matter more than mixing providers.
A natural shortlist when ElevenLabs voices are the product. Confirm the plan, LLM treatment, and telephony before budgeting.
Best fit for engineering teams that want agent code, transport, and model choice to remain portable.
Change the connected minutes below. Every row uses the same model and US carrier assumptions, so the comparison tests pricing shape instead of a provider's favorite demo.
Baseline: Nova-3, GPT-5 mini, a $0.015/min voice, US outbound Twilio, and no recording. Published rates checked before enterprise discounts.
Self-hosting is excluded from the ranking because it has no universal list price. Build a bill of materials for compute, networking, observability, capacity headroom, and on-call ownership. Start with the open-source framework directory.
A $0.05 platform fee is not a $0.05 phone call. The invoice is assembled across several meters.
The platform runs orchestration, agent sessions, transport, or call control. Vapi lists a $0.05/min hosting fee. Retell lists $0.055/min for voice infrastructure.
Cascaded agents pay for transcription, language-model work, and generated speech. The baseline uses Nova-3, GPT-5 mini, and a $0.015/min voice.
The US outbound baseline adds Twilio at $0.014/min. Inbound, toll-free, SIP, India, and other country routes need their own carrier rate.
Recording, transfers, phone numbers, denoising, knowledge bases, observability, and extra concurrency can move the total after the first pilot.
The cheapest row can still be the wrong system if your team cannot support its ownership model.
| If your constraint is | Start with | Pressure-test before buying |
|---|---|---|
| Fast managed launch | Vapi, Retell | Model freedom, concurrency, support, call logs |
| Predictable AI rate | Bland | Carrier pass-through, transfers, daily caps |
| Voice is the product | ElevenAgents | LLM billing, phone routes, plan commitment |
| Open agent code | Pipecat | Hosting profile, transport, on-call ownership |
| WebRTC and media rooms | LiveKit | Plan minimum, inference, observability |
| Infrastructure control | Self-hosted | Capacity, upgrades, monitoring, incident response |
A native speech-to-speech model does not have separate STT and TTS meters. Input audio and output audio are priced differently.
That difference matters because the caller and agent rarely speak for equal time. A sales qualification agent may listen more. A reminder agent may speak more.
Voice OSS models both sides instead of multiplying one advertised token rate by call length. Use the realtime voice cost calculator for OpenAI and Gemini native-audio scenarios.
Public pricing was checked against provider-owned pages on 2026-07-20. No affiliate ranking or paid placement changes the order.
We normalize connected minutes, model choices, carrier direction, and recording policy. The result is a list-price model, not a benchmark or negotiated quote.
We did not invent latency scores. Latency changes with the model, region, carrier path, endpointing, tools, and prompt. Run the same scripted calls from the countries you serve before committing.
Every platform name in the table links to its primary pricing page. The formulas and assumptions are documented in the methodology.
These pages unpack the pricing and ownership boundary between common shortlist pairs.
Compare platform fees, model flexibility, concurrency, and estimated call cost.
02Compare two open voice-agent foundations across cloud pricing and ownership.
03Compare a bundled minute rate with an unbundled, provider-flexible stack.
04Compare component pricing, included features, and predictable bundled costs.
05Compare voice-first bundled pricing with a configurable voice-agent platform.
The practical answers we use before a paid pilot.