Note: This is based on a real customer case study.
The cost-to-autonomy curve is a framework. Frameworks are easy to nod along to and hard to feel. This is what it looks like in practice, walked through with a realistic scenario built on proportions common in production agentic systems at enterprise scale, not a real customer's numbers. Nothing below describes an actual incident. It's constructed to show how the math in the autonomy curve plays out when something goes wrong at full autonomy, and the shape of it holds up whether the specific numbers are exactly these or not.

What Does a Typical Automated Workflow Cost Before Anything Goes Wrong?
Picture a large enterprise retailer running an AI agent at full autonomy, L4 on the scale from the cost-to-autonomy framework, to handle order exceptions: address corrections, minor refunds, inventory substitutions. Out of 500,000 daily orders, about 10,000 trigger some kind of exception the agent handles automatically. Each one costs roughly 18 cents in compute, tokens, tool calls, a quick lookup or two, for a workflow running smoothly. That's around $1,800 a day. Nobody's worried about this line item. It looks like automation working exactly as intended.
What Triggered the Runaway Agent Incident?
A vendor-side change to an inventory API starts returning malformed responses for one product category. The agent has no circuit breaker for this failure mode, and because it's operating at L4, there's no human checkpoint to catch the pattern early. Each failed tool call triggers a retry. Each retry generates a longer reasoning trace as the agent tries to work around data it can't parse correctly. Over roughly fourteen hours overnight, the agent doesn't just fail gracefully, it misinterprets the malformed responses as a signal to reissue refunds on orders it believes weren't processed, when they already were.
By the time a finance team member notices something's off, not through the AI system flagging itself, but through an unrelated reconciliation check the next morning, about 1,200 orders have received duplicate refunds.
How Much Did the Incident Actually Cost?
Four cost categories, stacked on top of each other. Compute cost during the incident window balloons to roughly $18,500, as retry loops and lengthening reasoning traces drive each affected task from a normal 2,000 tokens to 40,000 or more, an 18x jump in compute alone. The erroneous refunds themselves total around $50,400, 1,200 orders at an average refund of $42. Human cleanup adds about $2,450, a finance team spending the equivalent of five people working seven hours each reconciling accounts and reversing incorrect refunds. Support follow-up on the resulting customer confusion adds another $2,400 or so.
Total cost of the incident: roughly $73,750. Expected cost for that same fourteen-hour window, had nothing gone wrong: about $1,050. The incident cost around 70 times what the workflow was supposed to cost for that stretch of time.

What Would This Have Cost at a Lower Autonomy Level?
Run the same scenario at L3, autonomous execution with after-the-fact spot-check review and a basic error-rate threshold in place. The malformed API responses still cause failed tool calls and retries, but an anomaly detector catches the elevated failure rate after ten or fifteen attempts, well before duplicate refunds go out, and flags the workflow for human review. Extra compute cost from the retries before the flag fires: maybe $400. No erroneous refunds, because a human catches the pattern before any go out. No hours-long cleanup. Total incident cost at L3: under $500.
This is the autonomy premium from the cost-to-autonomy framework made concrete. L4 looked cheaper on a per-task basis right up until the moment something went wrong, and then the error-correction cost that L3's checkpoint would have caught early turned into a $73,750 overrun instead of a rounding error, roughly 150 times more expensive than the same failure caught at L3.
What This Means for Your Own Agentic Workflows
The lesson isn't that L4 is always wrong. It's that L4 without a circuit breaker or anomaly detection is a bet that nothing unexpected happens upstream, ever, in a system that has no mechanism to notice when that bet fails. At enterprise scale, that bet gets more expensive to lose, not less, because volume multiplies the blast radius of a single undetected failure. The fix isn't necessarily reducing autonomy across the board. It's building the equivalent of L3's safety net into L4: automatic error-rate thresholds that pause a workflow and escalate before a failure pattern compounds, rather than relying on autonomy level alone to define how much oversight a workflow gets.
The Takeaway
A $73,750 overnight incident sounds extreme until you walk through how it happens: a single upstream failure, no circuit breaker, no human checkpoint, and fourteen hours before anyone notices, at a scale plenty of enterprises already operate at. This is a constructed scenario, not a documented incident, but the mechanism it illustrates, error-correction cost compounding invisibly at full autonomy, is exactly the failure mode the cost-to-autonomy curve predicts. The fix is cheap relative to the risk: error-rate thresholds and anomaly detection cost a few hundred dollars against incidents that can run into the tens of thousands.



