What Does an Agent Actually Cost?

An interactive cost model for agentic LLM workloads · every input is yours to change

Illustrative rates

The workload

One run is a single agentic task: the model thinks, calls tools, reads the results, and thinks again. Each pass is a turn.

The lever almost nobody prices correctly. Cost does not scale linearly with this.
Resent on every single turn. Stable, so it caches well.
Files, specs, retrieved docs — whatever you hand it before turn one.
Tool results and the model's own replies, accumulating. This is what makes turns expensive.
Applies to the stable prefix only, and never to the first turn.

Rates

USD per million tokens

These defaults are illustrative round numbers, not any vendor's price sheet. Type your provider's current rates in and the whole page re-derives.

TierInputCachedOutput

Pick a tier to price the workload above. All three stay visible in the comparison below.

What is actually driving it

monthly, same tier

The same run under four different assumptions. Only the first one is real — the other three are the counterfactuals worth knowing.

Where the money goes, turn by turn

One run, broken out by turn. Watch the input stack climb while output stays flat.

cached input uncached input output

Which lever actually matters

±40% on one input at a time

Each bar is the change in monthly cost from moving that one slider 40% down or up, holding everything else still.

−40% on the input +40% on the input

Same workload, every tier

Identical run, priced three ways. The split between input and output shifts as rates change — which is why the cheapest tier is not always cheapest for a given shape of work.

input (incl. cached) output