AI
Finance
September 29, 2026

Cost Per Token Is a Vanity Metric: The Real AI Cost Optimization Metric You're Missing

Girish Bhat
SVP, Revefi

Every model release features the same headline highlighting another drop in price per million tokens. Vendors compete using this metric, procurement teams keep close track of it, and finance reports use it to demonstrate that running artificial intelligence is becoming less expensive. It stands as the simplest metric to report in the entire artificial intelligence stack, but for anyone working on actual cost optimization, it is the wrong metric to manage.

Cost per token measures the price of input and output data without indicating what that it achieved. A workflow that uses twice the tokens to reach an accurate answer on the first attempt costs less overall than a workflow that uses half the tokens per call but requires three retries to produce the same result. When you only monitor the price per token, both workflows appear equal, or even worse, the less efficient workflow looks like the better option.

Are Input Tokens and Output Tokens Priced the Same?

No, they are not, and the price gap is much larger than most budget discussions account for. Across every major provider, output tokens cost several times more than input tokens, often five to six times higher per million tokens. Models that appear cheap on a pricing page usually quote the lower input rate, and the headline figure used to compare vendors is almost always that lower input number rather than a blended rate.

This distinction matters because the structure of your workload determines actual expenses far more than the pricing page suggests. A workflow focused on short prompts and long outputs, such as reasoning steps, detailed explanations, tool call arguments, or draft-and-revise processes, pays higher rates on every single call even when input costs look low. Two teams comparing models based on headline token pricing are not comparing meaningful data unless they know the ratio of output tokens in each workload. A model that looks cheaper based on input rates can easily end up costing more in production if the workload generates a high volume of output tokens.

This dynamic explains why agentic workflows can become unexpectedly expensive compared to traditional chat applications. Intermediate planning steps, tool call parameters, and reasoning traces are all generated output tokens billed at higher rates, and those costs multiply with every step an agent completes before delivering a final result.

What Does Cost Per Resolved Task Actually Measure?

For AI cost management, the metric that truly counts is the cost per resolved task. This figure measures the complete expenditure across every token, tool call, retry, and escalation required to complete one unit of actual work accurately. It captures a reality that cost per token misses entirely, which is that a single user request almost never corresponds to a single model call. An everyday support inquiry might involve an information retrieval step, a reasoning pass, two tool calls to verify account status, a second attempt when the initial tool call yields malformed data, and a final response build. Every single action consumes tokens at its specific input or output rate. Standard model sticker prices fail to reveal any of this underlying spend.

This discrepancy grows more critical each quarter as token costs and overall AI expenditure head in opposite directions across many companies. While per-token rates continue to drop, total invoices keep climbing. This situation is neither a contradiction nor evidence of incorrect pricing reports. Instead, it demonstrates that usage volume, task complexity, and output-heavy operations are expanding faster than per-token discounts, a shift that goes unnoticed when organizations monitor the wrong metrics.

Why Does This Metric Gap Get Worse With Agentic Workloads?

Chat interfaces made this disconnect manageable because a user turn generally corresponded to a single model call. Agentic workflows completely disrupt that simple relationship. A single prompt can now split into planning phases, tool executions, sub-agent delegations, and retries before generating a final answer. Each intermediate step consumes tokens, particularly output tokens, that standard per-token pricing charts completely obscure. As agents gain greater autonomy and expanded tool access, the variance widens between the listed cost of a model and the real price of executing a task.

This concept serves as the core focus for the remaining articles in this series.

Token economics, AI financial operations, and agentic cost reduction all stem from a single underlying issue. The metrics that most teams rely on were created for an era when one user request equaled one model invocation, and that reality no longer exists for anyone deploying agents in production environments.

What Should You Measure Instead of Cost Per Token?

Three specific metrics are far more important than token rates when tracking actual AI expenditure. First, you must monitor cost per resolved task, an approach that demands tracing a request through every single operation rather than focusing solely on the final generation step. Second, you need to track tokens per task categorized by workflow type, separating input and output shares. This breakdown highlights processes that quietly drive up expenses through retries, context expansion, or lengthy responses instead of model selection. Third, you must plot total spending against unit cost as two distinct charts rather than combining them into one. A declining unit price alongside a growing total bill reflects operational reality rather than a reporting glitch.

These strategies do not require discarding per-token rates entirely. Token pricing remains a valid metric for evaluating baseline model costs, provided you compare equivalent input and output rates. The error lies in using a blended or input-only figure as a core management indicator to demonstrate efficiency to financial leaders. That figure does not reflect operational efficiency, it simply displays a standard price list.

The Takeaway

If the only AI cost optimization metric on your dashboard is price per token or per-million-token spend, you're measuring the vendor's pricing, not your own efficiency, and you're probably only seeing half the pricing table at that. Cost per resolved task is harder to instrument, because it requires tracing a request across every step it takes, input and output tokens alike, rather than reading a number off an invoice. It's also the only one of the two that tells you anything about whether your AI spend is actually working for you.

Girish Bhat
SVP, Revefi
Girish Bhat is a seasoned technology expert with Engineering, Product and B2B marketing, product marketing and go-to-market (GTM) experience building and scaling high-impact teams at pioneering AI, data, observability, security, and cloud companies.
Blog FAQs
What is the cost per token in AI, and why isn't it enough?
Cost per token is the price a model provider charges per unit of text processed or generated. It measures input pricing, not outcome efficiency, and often refers only to the input rate even when cited as a single headline number. A cheap-per-token workflow that needs multiple retries, or that generates long output, can end up costing more overall than a more expensive one that resolves correctly the first time.
Are output tokens more expensive than input tokens?
Yes. Across every major model provider, output tokens typically cost five to six times more per million tokens than input tokens. This means a workload's real cost depends heavily on its input-to-output token ratio, and two models that look similarly priced on a headline number can cost very differently in production depending on how much each workflow generates versus reads.
How do you calculate cost per resolved task?
Add up every cost driver involved in completing one unit of work correctly: model calls, tool calls, retries, escalations, and any human cleanup required, priced at each token's actual input or output rate. Divide that total by the number of tasks successfully completed. This traces spend across the full workflow rather than a single model call or a single blended rate.
Why is my AI bill going up even though token prices are falling?
Falling per-token prices lower unit cost, but total spend depends on usage volume and task complexity too. When usage grows faster than prices fall, and agentic workflows multiply the number of token-consuming steps per request, many of them output-heavy, total spend rises even as the per-token rate drops.
Does switching to a cheaper model always reduce AI costs?
Not necessarily. A cheaper model with a higher error rate can require more retries or human correction, raising the effective cost per resolved task even though the per-token price is lower. It's also worth checking whether “cheaper” refers to the input rate, the output rate, or both, since a model can be cheaper on one and more expensive on the other. Cost per resolved task is the metric that reveals whether a model switch actually saved money.