AI Agent Cost Tracking: Spend Cap vs Attribution

A cap stops a runaway bill and returns one number. Finance asks which agent, workflow or customer the money bought. Where caps, dashboards and tags fall short.

By Harinderpal Hanspal on July 2026. Updated October 2026

In a February 2026 survey for DoiT, 79% of 500 finance leaders reported AI cost overruns in the past year. A spend cap limits the damage but returns one number, and cannot say which agent, workflow or customer the money was spent on.

A spend cap returns one total, while finance asks which agent, workflow or customer it boughtSketch with two halves. On the left a single block of spend sits under a spending cap and is labeled one number. On the right three questions that finance asks of the same spend, which agent, which workflow and which customer, each marked as unanswered.a capceilingone numberstayed under. so what?what finance askswhich agent?which workflow?which customer?not in itnot in itnot in itthe total has no names in it
A spend cap returns one total, while finance asks which agent, workflow or customer it bought

A cap adds each call's cost to a running total and refuses the next call once the total reaches a ceiling. For a loop that will not stop, that is the right tool. It also produces exactly one number per account, and that number has no agent, campaign or customer in it.

How common the overrun is

DoiT commissioned Sapio Research to survey 500 finance leaders, manager level and above, at US and UK organizations with 1,000 or more employees that were spending on AI tools. Fieldwork ran in February 2026, with a margin of error of 4.4 points. In the past 12 months, 79% reported cost overruns (DoiT, AI spending survey, 2026). Two cautions apply. DoiT sells cost-management tooling, and the page does not define what counts as an overrun. Read it as 79% of those 500 leaders, not 79% of enterprises.

Attention is nearly universal. The FinOps Foundation's sixth annual survey, published in February 2026, covered 1,192 respondents representing more than $83 billion in annual cloud spend. It found that 98% now manage AI spend, up from 31% two years earlier (FinOps Foundation, State of FinOps 2026). That measures effort, not outcomes.

OWASP lists "unbounded consumption" as LLM10 in its 2025 list: an application that allows excessive, uncontrolled inference, with economic loss among the results, including attackers who run up the bill (OWASP GenAI Security Project, LLM10:2025).

Gartner goes further and predicts that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls. Those are alternative causes. The figure is a prediction built on a January 2025 poll of 3,412 webinar attendees, not a count of cancellations (Gartner, 25 June 2025).

What a cap does and does not tell you

Consider a fleet of agents that stays under its monthly ceiling. Nothing alarms. At month end finance asks which of them was worth what it cost, and the cap has no answer, because it never recorded who spent what.

Three weaknesses recur.

The estimate is a floor. A check before the call can price the prompt, since its length is known. The answer has not been written yet, so output tokens are guesses, and a cap built on the estimate undercounts exactly where reasoning-heavy agents spend the most.

Coverage follows the code path. A cap enforced in the place your own application calls the provider does nothing for a second path that calls a different provider, a retry layer or a tool that makes its own model calls. The unmetered path bills without limit.

Requests are not dollars. A rate limiter counts calls per minute and sees no cost at all. A flat limit lets an agent that sends large prompts spend many times what a chatty cheap one does.

Where dashboards and tags run out

Most teams fill the gap with provider dashboards, separate API keys per project, and billing exports sorted by tag. That is a fair start. It breaks when one key serves several agents, when some calls carry no tag, and when the bill covers only model tokens. Image and audio generation, paid external APIs and the hardware under a locally hosted model sit outside the token invoice. Some tools price a local model at zero, which is correct for vendor spend and wrong for the cost of serving it.

The test of a real cost record is that each event carries who asked and for what: workspace, workflow, agent run and the identity behind it. Without that, a report can say what was spent and never what it bought.

A month-end test for agent cost records

Ask for last month's spend, grouped by agent and by workflow, and time how long the answer takes. Then put these to the vendor:

Drawn from DoiT and Sapio Research (survey fielded February 2026), the FinOps Foundation's State of FinOps 2026, OWASP LLM10:2025 and Gartner's press release of 25 June 2025, all read on 6 October 2026. No figure here is a measurement of ours. The DoiT survey is vendor-commissioned.

Related notes

Related insights

Related paper: Governing agents in production: what to ask before an agent acts