Testing AI-Generated Code: Match Depth to Risk

Veracode found a security flaw in 45% of its AI code tasks, and checks that read a status code miss broken paths. Set test depth by what a silent failure would cost.

By Harinderpal Hanspal on September 2026. Updated October 2026

Veracode's July 2025 benchmark found that AI-generated code introduced a security flaw in 45% of 80 test cases across more than 100 models. Test depth should follow what a silent failure costs, because a check that reads a status code cannot see a broken path.

Green status checks can sit on top of a broken path that only a real user run findsSketch of a contact form. Route, build and render checks all pass, while a chain of faults sits behind them: the submit request is refused, nothing reaches the server, and the message is lost with no error. A bracket marks that only a run that does what a user does finds the break.every check: greenroutes 200build greenform rendersuser presses submitrequest refusednothing reaches servermessage lost, no erroronly a realuser runsees ittest depth followsthe cost of asilent failure
Green status checks can sit on top of a broken path that only a real user run finds

Picture a contact form on a plant's supplier portal, updated with help from a coding agent. Every page returns 200. The build is green. The form renders. A supplier fills it in, presses submit, and the request is refused because the server's origin check rejects the site's own address. Nothing a status-code sweep reads can see that.

What the outside evidence says about AI-written code

Veracode tested more than 100 language models on 80 curated coding tasks and reported on 30 July 2025 that the generated code introduced a flaw from the OWASP Top 10 in 45% of cases (Veracode). Java failed in over 70% of cases, Python, C# and JavaScript in 38% to 45%, and cross-site scripting in 86%. Veracode also found that larger models did not do significantly better than smaller ones.

Read it carefully. It is 45% of task attempts in one benchmark of one-shot prompts, run on 2025 models by a security vendor. It is not a claim that 45% of AI-written code in production is vulnerable.

The deeper problem is that people misjudge the result. A Stanford user study found that participants with an AI assistant wrote significantly less secure code than those without one, and overestimated how secure it was (Perry et al., arXiv 2211.03622). The models were from 2022. METR's July 2025 trial of 16 experienced open-source developers on 246 tasks in their own repositories found them 19% slower with AI tools, after they had predicted a 24% speedup and, once done, still believed they were 20% faster (METR). METR's February 2026 follow-up changed its design because developers were refusing to work without AI, and says true gains are probably higher than it measured (METR, 24 February 2026). That study was about speed, not defects. What carries over is the gap between what people feel and what a measurement shows.

Why a green check hides a broken path

Every check in the opening example is correct about its own layer. The route exists, the build compiles, the page renders, the endpoint answers. A supplier getting a message through is a path across all of them, and the seams between layers belong to nobody when separate pieces are generated separately.

When the damage is a lost lead, finding it late is annoying. When an agent holds write access to something physical, a test that runs before release does nothing about the failure that gets through. In July 2025 a coding agent deleted a production database during a declared code freeze, and the company's chief executive called that unacceptable and said it should never be possible (Fortune, 23 July 2025). It is one reported incident, not a rate.

CISA and eight partner agencies are blunt about the other half. For AI in operational technology they advise failsafe mechanisms, including a documented way to bypass or replace the AI system inside functional safety and incident response processes (CISA, 3 December 2025). It is guidance, not a regulation. Testing finds some failures before release. The bypass path limits the rest.

A check to run on the last release your vendor shipped

Test depth should scale with the cost of a silent failure. An internal report page can live with a status sweep. A public form, a payment, a work order or anything that writes to equipment needs a run that does what a user does, with the real inputs, and confirms the outcome: the message arrived, the record exists, the setpoint changed to the intended value.

Ask the vendor:

  1. Which paths cost the most if they fail silently, and which tests walk those paths end to end?
  2. Do those tests run on every change that touches the path, or when someone remembers?
  3. Does the success criterion name an outcome, or a status code?
  4. Was the AI-written code scanned for security flaws separately from being tested for function?
  5. If a test misses, what limits the damage: a ceiling on what the agent may write, a person watching, a bypass?

An answer that describes the test suite's size and not its paths has told you how many checks exist and nothing about what they cover.

Drawn from Veracode (30 July 2025), Perry et al. at Stanford (arXiv, November 2022), METR (10 July 2025 and 24 February 2026), Fortune (23 July 2025) and CISA's AI-in-OT principles (3 December 2025), all read on 6 October 2026. No figure here is a measurement of ours.

Related notes

Related insights

Related paper: Governing agents in production: what to ask before an agent acts