Fully-loaded AI cost

What does a typical agent actually cost?

Not the price of a token. The fully-loaded cost of the whole thing: the prompt and tools billed on every turn, the clips you render and discard, the voice on top. Pick your use case, then any model on the stack for each step.

Models & assumptions
Channel
Fully loaded · per conversation
$0.048
At 25,000/mo
$1,202/mo
TTS · Flash v2.5 / Turbo v2.5 is 78% of it. Same workload across every LLM on the stack runs $0.039$0.981 (25× swing) — model choice is the biggest lever.
Pipeline breakdown · per conversation
TTS · Flash v2.5 / Turbo v2.5$0.038 · 78%
3× ~138 spoken chars/turn
LLM · input (prompt + tools)$0.0092 · 19%
3× turns · ~10,200 tok/turn · ~6,300 fixed overhead re-sent/turn
STT · Saarika$0.0009 · 1.9%
3× ~0.06 audio-min/turn
LLM · output$0.0003 · 0.7%
~45 tok/turn
typical consumptionestimate consumptionprices verified · your workload will differ
Get your workload benchmarked

The consumption above is typical. We measure yours on your models and hardware, then tell you the number.