Agent Loop Limits: Step Cap, Wall Clock, Cost Ceiling
A step cap is the default stop for a runaway agent loop. What it misses, and why a wall-clock limit, repeated-action detection and a cost ceiling each catch a different failure.
By Harinderpal Hanspal on October 2026
Mandiant reported an accounting agent that made over 15,000 API calls in under an hour and ran up about $50,000, with no attacker involved. A step cap counts steps; a wall-clock limit, repeat detection and a cost ceiling each stop a failure the cap lets through.
In September 2026, Help Net Security reported on a Mandiant study of an accounting agent that malfunctioned with no attacker involved. It made more than 15,000 high-cost API calls in less than an hour and ran up roughly $50,000 in cloud charges (Help Net Security, 16 September 2026). Mandiant sells AI security services, the article does not say which controls were missing, and this is one incident. It does show the speed: more than four calls a second, on average, for as long as nobody looked.
OWASP lists the general risk as LLM10, Unbounded Consumption. Its mitigations include rate limits and quotas, and timeouts and throttling for resource-intensive operations (OWASP LLM10:2025). The page is written for model-backed applications in general and never mentions agent loops or spending budgets. Agents make the exposure larger, because each step feeds the next. Anthropic reported that its multi-agent research system used about 15 times the tokens of a chat, and that early versions spawned 50 subagents for simple queries (Anthropic Engineering, 13 June 2025).
What the frameworks give you by default
Most teams stop a loop with a step cap, and the frameworks ship one. LangGraph raises GraphRecursionError past a recursion limit that has defaulted to 1000 steps since version 1.0.6 (LangChain docs). CrewAI's max_iter defaults to 20, while its max_execution_time and max_rpm default to no limit (CrewAI docs). AutoGen offers message, token and timeout conditions that combine with AND and OR (AutoGen docs).
So a time limit and a token limit exist. In at least one of these, the time limit is off until someone turns it on.
What a step cap misses
Picture an agent that reschedules a maintenance work order. The scheduling tool rejects the slot as a conflict. The agent tries another technician, gets another conflict, and asks the planning agent, which asks the scheduler again.
Three failures follow, and a step cap handles none of them well.
A slow loop. If each call hangs for a minute on a sluggish plant system, 1,000 steps take about 17 hours. The cap works, eventually, and the work-order queue is blocked the whole time. A wall-clock limit on the run ends it in minutes.
A fast loop. At four calls a second the clock limit is the wrong tool and the step cap fires too late. Suppose each step re-sends 100,000 tokens of context. At Sonnet 5.5's list price of $2 per million input tokens (Anthropic pricing), that is $0.20 a step and $200 in input alone across a 1,000-step cap, before a single output token. A cost ceiling that refuses the next call is the stop that matches the damage.
A loop that looks like progress. The agent varies one argument each time, so every step is different and the count keeps rising. Repeated-action detection catches it by comparing the tool, the arguments with timestamps and ids removed, and the result. Three identical outcomes in a row is a stuck agent.
The three limits fail in different places. The clock misses a burst, the ceiling fires only after money is spent, and the repeat check misses a loop that never repeats exactly. That is why they run together, with the cap as the fourth.
Questions to put to a vendor about loop limits
- Which limits apply to a run by default: steps, wall-clock time, tokens or dollars? Which are off until configured?
- Is the cost ceiling checked before the next model call, or reported afterward in a dashboard?
- How does the system detect an agent repeating itself with small changes to its arguments?
- When a limit fires, does the work order stay open with a clear stopped state, or does the run vanish?
- Do sub-agents count against the parent's limits, or does each get a fresh budget?
- Who is paged when a limit fires, and how fast?
A vendor who answers "we have a max iterations setting" has answered the first question and none of the others.
Drawn from Help Net Security (16 September 2026) reporting Mandiant's AI Risk and Resilience report, OWASP LLM10:2025, the LangGraph, CrewAI and AutoGen documentation, Anthropic Engineering (13 June 2025) and Anthropic's pricing page, all read on 6 October 2026. The sums on steps and dollars are our arithmetic on list prices.
Related notes
- AI agent cost tracking: a spend cap cannot say which agent earned its bill
- Test AI-written code by what a silent failure would cost
Related paper: Governing agents in production: what to ask before an agent acts