Confidence in AI-generated code is rising in lockstep with its failure rate

Confidence in AI-generated code is rising in lockstep with its failure rate
Fitz Nowlan
  September 30, 2026

If you only read the headlines this year, you’d think AI makes shipping good software easier than ever. Yet, this was the year of very public AI-coding incidents. PocketOS’ production database deletion and a Vercel AI agent shipping unverified code are just two examples of AI-authored code shipped with confidence that turned out to be wrong. And, of course, these AI-coding incidents are distinct from Agentic AI orchestration incidents like the HuggingFace hack by OpenAI. 

SmartBear just finished research on this topic, the 2026 State of Software Quality and Testing Study, and two numbers stand out. 46% of surveyed teams have shipped AI-generated code that failed once it hit production. Yet 69% of those same teams say they’re highly confident their AI-generated code behaves as intended. You’d expect those numbers to move in opposite directions; they don’t. 

What’s actually missing isn’t more caution. It’s a holistic perspective that lives inside every experienced engineer’s head and doesn’t exist inside an AI system at all: the judgment to notice when something’s off before it ships. That’s the thread running through everything below: how that judgment gets lost, what it costs teams, and what it takes to build it back in as something explicit instead of assumed. 

There is no free lunch 

A year and a half ago, the instinct foisted upon most engineering orgs was simple: use AI everywhere and reap the gains; the proof will follow. That instinct is shifting as teams realize the full and ongoing responsibility of coding with AI. SmartBear’s report found that only 46% of teams check that a specification reflects actual intent before AI generates from it. Closing that gap between spec and intent is something humans did naturally, but the reliance on AI-coding exposes the need to be overly specific and concrete about what “correct” looks like. 

Furthermore, 55% of teams report application quality issues in the past year that they trace directly to development outpacing testing. Why do teams push ahead with AI-coded solutions when the quality can’t keep up? Most teams never made that call deliberately. AI is often the fastest, cheapest stand-in for something that would otherwise take years to build as durable, deterministic software. Making that trade makes sense when speed to market is the priority. But if the product gains traction, or in markets where timing isn’t the deciding factor in success, it’s usually still worth building the real, durable version instead of leaving AI standing in for it permanently. 

The known and unknown metrics of AI-driven development 

Integrating AI deeply into the software development lifecycle (SDLC) doesn’t actually change the criteria used for determining whether a product or business is successful. Revenue, usage, cost, and performance efficiency are still the right metrics to measure. 

What does change now with AI powering the SDLC is the human judgement lost between the metrics; a human could discern whether usage was legitimate and correlated to revenue. AI systems aren’t by default focused on the same connective tissue in the product broadly. Pass rates, coverage, and defect count still matter, but they aren’t enough if the human reviewer is not relating them to each other. AI’s programmatic intelligence is capable of fulfilling this role, but it must be directed to do so and provided with all of the correct information—data that a human would instinctively assemble on their own. 

We see this uncertainty materialize as 47% of teams have hit an AI-related incident they couldn’t fully explain. Only 50% can fully track AI output for audits, which is really a system-of-record problem. Without one, there’s no measurable assurance of what happened even after the fact, let alone a way to catch it while it’s happening. 

Underlying both of those stats is a related fundamental question: should AI author and test its own code? Only a quarter of respondents think it’s always appropriate for the AI that wrote the code to also test it, and many only think it’s okay for low-risk changes. Yet 92% accept AI as a primary tester in some form. That’s a real gap between the conditions teams say they want and the oversight actually implemented today. What’s missing is the tooling to close it. SmartBear’s BearQ™, an agentic QA platform for application integrity, is built to do just this. 

Disappearing guardrails in software development  

For decades, the guardrails on software lived in people, sitting in the heads of whoever had seen the failure before or remembered the customer feedback motivating a feature. That instinct is years of pattern-matching doing more governance work than any code review checklist ever did, and rarely was it ever written down. AI doesn’t have that. It has no customized and tailored sense of “this is off,” because it hasn’t spent the years acquiring the expertise in a domain that produces the instinct. 

Imagine sending a bowling ball down the lane with and without bumpers. Bumpers keep the ball in the region where a decent outcome is nearly guaranteed. That’s shaping the space it’s allowed to wander in, rather than trying to steer its momentum and tendencies head-on. The judgment an experienced engineer brought into their coding was those bumpers. AI runs off the lane without them.  

This relates to a disconnect between engineering practitioners and their own leadership over the use and ROI of AI. Most tension here comes down to the two sides not agreeing on the same reality: 73% of leaders report a lot of or complete confidence that AI-generated code behaves as intended, compared with 52% of practitioners saying the same. Leaders and teams often aren’t picturing the same target: is success tied to a percentage increase in revenue, a defined share of AI-reviewed code, or something else? Once leaders and practitioners agree on the metrics of success, disagreement stops being an ambiguous standoff and becomes a yes-or-no check against a shared number. 

Agents are moving faster than their guardrails  

All of the above indicates that the focus on governance trails the focus on programming in the AI-powered SDLC. As evidence, 96% of teams are piloting or already running AI agents, but only 34% have purpose-built oversight tooling to watch what those agents actually do. 

The specification constraints highlighted above and the governance tooling to enforce them are two sides of the same coin. Constraints define the rules; the harness is what actually enforces them. Without active and ongoing enforcement, the rules exist on paper and nowhere else. Development velocity went up 10x with AI; the harness that watches it hasn’t caught up to the same rate. But that is changing. Teams are already building agentic QA today to meet the gap and acquire operational maturity in matching the speed and scale of AI-powered development. 

Embrace AI coding with application integrity  

For a board member or a P&L owner, none of this is an abstract engineering debate. The result of only 46% of teams checking that a specification reflects actual intent before AI generates from it produces unpriced risk, the kind that only gets discussed in a board meeting after an incident.  

A gut-check on where your team actually stands with regards to shipping AI code with measurable assurance boils down to a few questions: 

  1. Do we check that the spec reflects intent before AI generates from it, or do we just assume it does? 
  2. If an agent did something we can’t explain, could we reconstruct what happened after the fact? 
  3. Is our oversight tooling scaling with how many agents we’re running, or falling behind? 
  4. Is the same AI that wrote the code also the one deciding whether it’s correct? 

The bugs are real, the production failures are real, and pulling back feels like the responsible move after a year like this one. But pulling back from AI in coding doesn’t offer a fix; AI will continue to be integral to how software gets written and it is valuable. There’s no way around it. The fix is enforcing integrity with AI at the same speed and the same breadth that AI is now writing the code. That is, validate intent before generation, track output well enough to actually audit it, and scale oversight with the agents. 

Delivering application integrity – continuous, measurable assurance that your software just works as intended, with the governance to operate at AI speed and scale – requires a disciplined approach to spec validation, constraint engineering and auditing.  

Nobody is going back to writing every line by hand. The most adaptive teams keep shipping fast, with the confidence numbers to back it up. The ones that don’t will only learn once an incident is finally expensive enough to notice, but by then it may be too late. 

Download the full 2026 State of Software Quality and Testing Study to see the complete data set behind the confidence paradox, or register for our BearQ event on October 22. See how BearQ extends beyond creating tests to coordinating quality work across Jira, GitHub, and AI-powered workflows – detecting issues, triggering the right next action, and carrying context from one agent to the next.

Frequently asked questions about the study 

What is the SmartBear 2026 State of Software Quality and Testing Study? 

The State of Software Quality and Testing research report examines how technology practitioners and senior leaders are navigating AI’s role in the software development lifecycle, specifically validation, governance, autonomous agent accountability, and scale. 

Who did we survey for this report? 

We surveyed almost 1,500 technology practitioners and senior leaders using AI in development, primarily based in the United States and United Kingdom, including 150 SmartBear customers. 

What is the report’s central finding? 

Leaders and practitioners remain confident in AI-generated code even after experiencing failures, and that same overconfidence extends to how they validate and govern it. 46% of teams have shipped AI-generated code that later failed in production, yet 69% of those same teams still report a lot of, or complete, confidence that AI-written code behaves as intended. 

Why did SmartBear conduct this research? 

We conducted this research to measure whether the tools and practices used to validate AI-generated code have kept pace with how much software development now relies on AI, and to give engineering and business leaders concrete data instead of anecdotes to make decisions from. 

You Might Also Like