46% Have Shipped Failed AI Code, Yet 69% Are Still Confident in It, New SmartBear Survey Finds

SmartBear's 2026 State of Software Quality and Testing report finds leaders' blind trust in AI code also shows up in how they validate and govern it

SOMERVILLE, Mass. - September 30, 2026 - SmartBear, helping teams build, test, and ship quality software at AI speed and scale, today announced survey findings that show tech teams maintain high confidence in AI-written code, despite frequent code failures, an inability to trace AI’s contribution to those failures, and mounting pressure on application quality.

SmartBear’s 2026 State of Software Quality and Testing surveyed 1,436 U.S. and U.K. leaders and practitioners who use AI in development. It reveals that 46% of teams have shipped AI code that later failed in production. Of those, 69% still have a lot or complete confidence that AI-written code behaves as intended. This blind trust in AI code increases the farther you get from the code. 73% of leaders have a lot or complete confidence that AI code works as intended compared to 52% of practitioners. 

“AI coding failures are already costing companies revenue, customers, and trust, yet leaders remain blindly confident that AI-generated code behaves as intended, even after watching it fail,” said Dan Faulkner, SmartBear CEO. “That same confidence-over-proof mentality runs through how organizations validate and govern their AI code. Closing the gap between perception and reality demands quality and testing built for AI’s speed, complexity, and scale.”

Other findings from SmartBear’s 2026 State of Software Quality and Testing research include:

  • Testing happens too late: Only 46% of teams validate the specification before AI generates code from it.
  • Limited visibility into AI coding issues: Almost half of organizations (47%) can’t explain how AI contributed to a bug or incident when one hits production. That number climbs to 73% once a team has already shipped a failure.
  • Confidence outpaces governance: 95% of leaders and 86% of practitioners say they’re confident their AI agents’ work is adequately reviewed. Yet only 25% of teams have humans reviewing more than 80% of what those agents actually produce, a gap that suggests oversight hasn’t kept pace with the confidence placed in it.
  • More oversight leads to better results: The report found teams that review more of their agents’ work ship fewer failures than teams that review less.

More AI, More Application Quality Issues

All of this is happening against a backdrop of increasing AI use to create code and continued concerns over software quality. 69% of U.S. software experts say AI writes or accelerates 41% or more of their code. That’s up from 43% of experts who said the same in SmartBear’s January survey.  

73% of experts are at least somewhat concerned their application quality is suffering, while 45% of U.S. respondents are very or extremely concerned, up from 36% in January. Meanwhile, more than half of all respondents, 55%, have experienced application quality issues in the past year because their testing can’t keep up with development, resulting in revenue loss, outages, and negative customer experiences.

Opportunity for Autonomous Testing

AI-powered testing and validation tools can help companies achieve application integrity, continuous and measurable assurance that software works as intended, and teams are already putting these tools to work. 65% of respondents say AI generates or maintains at least 41% of their test coverage. Also, 83% say autonomous testing, where AI agents independently generate, execute, adapt, and report on tests without manual scripting, would help them keep pace with AI code development. 

Yet the research finds that trust remains a barrier to adopting or scaling autonomous testing. About 1 in 4 respondents (23%) say trust is the top barrier, nearly double the 12% who name cost. Human review answers that directly. Agents flag what needs a closer look, and people provide the judgment agents can’t. Teams see value in a human-in-the-loop approach as just 3% rely on AI self-validation alone, while 84% use at least one form of human review to validate AI-generated tests.

To see the full survey data, visit: smartbear.com/state-of-software-quality-and-testing.

About SmartBear

SmartBear delivers application integrity for modern tech stacks, ensuring continuous, measurable assurance that software just works as intended – with governance to operate at AI speed and scale. SmartBear offers deep test automation, API lifecycle management, and observability capabilities. With integrations across the SDLC, it sets a new quality standard for application delivery teams.

SmartBear is trusted by developers, testers, and software engineers across 32,000 organizations, including 75% of the largest financial institutions and industry leaders such as Adobe, JetBlue, and Microsoft. SmartBear’s open source tools are downloaded more than 100 million times a month and have earned over 30,000 GitHub stars from the developer community. With its best-loved brands, including Swagger, TestComplete, Reflect, QMetry, Zephyr, and more, SmartBear meets customers where they are to make our technology-driven world a better place. Learn more at www.smartbear.com, or follow us on LinkedIn, X, and Reddit.

Share on