The State of Software Quality and Testing 2026
AI confidence is high, evidence lags behind
SmartBear's 2026 report offers a look at how software teams are navigating the AI-driven SDLC. Blind trust is the throughline, from the AI-generated code itself to the governance meant to keep it in check.
- 65%
- say AI writes or accelerates at least 41% of their code
- 73%
- are at least somewhat concerned application quality is suffering
- 46%
- have shipped AI-generated code that later failed in production
- 83%
- say autonomous testing would help them keep up with AI code volume
- 1,436
- respondents
- 49% / 51%
- practitioner / leader
- 5
- industries surveyed
01 — AI adoption
AI adoption skyrockets in just 7 months
AI took over the pipeline faster than traditional QA could keep up, and adoption shows no sign of slowing down.
About two-thirds of respondents (65%) say AI writes or accelerates at least 41% of their code. Nearly a third (31%) use AI for at least 61% of their code.
- AI writes or accelerates at least 41% of code 65%
- AI writes or accelerates at least 61% of code 31%
Focusing on the U.S., 69% of U.S. software experts say AI writes or accelerates at least 41% of code. AI usage has jumped dramatically over the past seven months. When we asked this same question in January 2026 for our Closing the AI Software Quality Gap report, only 43% said they used AI for that same amount of code.
U.S. respondents only. January 2026 figure from Closing the AI Software Quality Gap.
Testing and validation tools that are likewise powered by AI can counterbalance the velocity and abstraction of the AI-driven SDLC. Teams are capitalizing on the new technology, with 65% of respondents saying AI generates or maintains at least 41% of their test coverage.
- 0-20% 12%
- 21-40% 23%
- 41-60% 32%
- 61-80% 25%
- 81-100% 8%
02 — Blind trust
Blind trust in AI makes validation non-negotiable
As AI produces more code and applications faster, 73% of software experts are at least somewhat concerned that application quality is suffering. That includes 45% of U.S. respondents who are very or extremely concerned about application quality, which is up from 36% in January.
- Are at least somewhat concerned application quality is suffering 73%
- Saw quality issues in the past 12 months from development outrunning testing 55%
55% of professionals say their organization experienced application quality issues in the past 12 months that they attribute to development moving faster than testing could keep up.
The repercussions are both vast and costly for those experiencing application quality issues due to development outpacing testing. It’s becoming clear that AI can create big problems before teams even realize it, increasing downstream risk.
- Significant rework or remediation 54%
- Increased incidents or outages 46%
- Missed release deadlines 41%
- Negative customer experience 31%
- Breach of SLA or compliance 26%
- Reputational damage 22%
- Revenue loss 18%
- Job dissatisfaction 16%
Organizations are not letting concerns or negative outcomes get in the way of their embrace of AI, however. Perhaps no two findings illustrate blind trust in AI as much as these:
46% have shipped AI-generated code that later failed in production, yet of those teams, 69% still have a lot or complete confidence that AI-written code behaves as intended. Even among teams with high confidence in AI-generated code, half (50%) have shipped code that later failed.
This reflects a similar pattern seen in DORA’s research, which found higher AI adoption is associated with a nearly 10% increase in software delivery instability.
- Have shipped AI-generated code that later failed in production 46%
- Of those teams: still have a lot or complete confidence in AI code 69%
- Of high-confidence teams: have shipped code that later failed 50%
Slicing the data by role highlights disconnects between leadership and those closer to the code: leaders are blindly trusting AI, reporting more failures than practitioners and more confidence. Nearly three-fourths of leaders (73%) have a lot or complete confidence that AI-generated code and applications behave as intended compared with about half of practitioners (52%). While 50% of leaders report their organization has shipped AI-generated code that failed in production, only 41% of practitioners say so.
- Report their organization shipped AI code that failed in production 50% 41%
- Have a lot or complete confidence AI-generated code behaves as intended 73% 52%
The fact that people farthest from code trust it the most resurfaces throughout the dataset. For instance, more than half of leaders (57%) say they understand the day-to-day quality risks AI-generated code creates very well, but only 37% of practitioners say the same of leadership. What’s more, 16% of practitioners say leadership has a poor understanding of the risks.
This divide matters. When leadership believes AI works as intended, it falls to practitioners to make the results match expectations, whether or not they have the proper tools and processes to do so. Validation has never been more important.
While leaders are confidently scaling AI-coding investment, practitioners need more of that confidence. The data shows it’s a virtuous cycle: increased AI usage — seeing more of it at work, fails and all — is what actually builds confidence. But what about all those failures? Teams don’t interpret them as indicators that they shouldn’t be using AI, but rather as signs that they need more capacity to check it.
03 — Independent checks
Both humans and AI have roles to play in today’s testing
A frog in a hot pot doesn’t know it’s in trouble until it’s too late. Organizations with too much trust in AI and too little independent validation can find themselves in hot water of their own.
While AI can help test software, organizations need the right checks in place to independently verify its work at every stage of the development lifecycle. The closer to code, the more skeptical: 81% of leaders and 64% of practitioners believe AI can reliably catch its own errors.
That 17-point gap on self-checking captures the core tension running through the SDLC. Closing the distance between leaders’ perceptions and practitioners’ reality starts with organizations investing in quality and testing built to deliver application integrity.
- Believe AI can reliably catch its own errors 81% 64%
- Can identify all or most AI-written code in the codebase 87% 73%
Confidence keeps outpacing proof. On every measure in this figure, leaders report more confidence than the practitioners closest to the code.
Practitioners are more cautious about AI-generated tests than leaders. Only 56% of practitioners trust AI-generated tests more than human-written tests compared with 72% of leaders. About 20% of practitioners trust them less than human-written tests, which is twice the leader rate of 10%.
- Trust AI-generated tests more than human-written tests 72% 56%
- Trust AI-generated tests less than human-written tests 10% 20%
The majority of respondents aren’t yet ready to keep humans out of the loop. Only 3% rely on AI self-validation alone, and 84% use at least one form of human review to validate AI-generated tests. In addition, while 92% accept AI as the primary tester of its own code in some situations, about half (48%) think it’s appropriate only with human oversight.
- Use at least one form of human review to validate AI-generated tests 84%
- Rely on AI self-validation alone 3%
- Always appropriate
- 25%
- Only with human oversight
- 48%
- Other conditions, or never
- 27%
Shares of all respondents. Other conditions, or never includes those that say it's only appropriate for low-risk code, never appropriate, or not sure.
This skepticism about AI testing its own code is warranted. How can teams validate the code if they don’t even know what specification the AI is writing against?
Having the same system responsible for generating the code and validating it results in a testing black box — neither the logic behind the tests nor the assumptions baked into them are independently legible to the humans nominally overseeing the process. Teams have no separate signal to confirm that what is being tested reflects what actually matters, and no clear line of sight into whether a passing test suite represents genuine coverage or simply AI confirming its own work.
When teams independently review AI code — either with a human or a different AI than the one that wrote it — they employ distinct strategies at three key checkpoints.
- 46% Before generation
- 49% Before commit
- 60% Before release
- Before generation
- check that the specification reflects intent before AI generates code from it
- Before commit
- independently review more than 60% of AI-written code before committing it
- Before release
- independently test more than 60% of AI-written code before releasing it
Columns share a 0–100% scale. Checking increases the closer code gets to shipping — the opposite of shift-left.
Teams are more likely to check the code the closer it gets to shipping, which runs counter to the shift-left approach teams have pursued for years. Validating the specification before AI ever starts coding from it is the earliest and cheapest point to catch a problem. Bugs found later in the SDLC are the most costly and time-consuming to fix. So while AI is helping some testing teams achieve more coverage, it’s not necessarily solving the problem of cost. Thorough testing is more than just catching the issues; it’s doing so before they become a drain on the work.
04 — Governance
AI’s black box undermines governance
Most organizations can point to what AI wrote. Far fewer can prove what parts of their code were AI-generated, and even fewer can explain it when something breaks.
Unlike traditional scripted automation, where a failing test points straight at the offending code, an AI-generated failure leaves teams investigating blind, which is exactly why the cost of finding it climbs. Ultimately, vibe coding and AI-generated code operate in black boxes, and no system of record means no governance.
80% of software experts say they can identify all or most of the AI-written code in their codebase.
- Can identify all or most AI-written code in their codebase 80%
- Have been unable to explain how AI contributed to a bug or incident 47%
Here again, leaders and practitioners have strikingly divergent perspectives: 87% of leaders say they can identify most or all AI-written code, versus 73% of practitioners. In addition, 43% of leaders say AI outputs are not fully tracked, but 57% of practitioners say the same.
Meanwhile, nearly half of teams (47%) have been unable to explain how AI contributed to a bug or production incident. That number climbs to 73% among teams that have already shipped a failure. Full audit tracking doesn’t bridge the divide either. Teams with complete tracking (72%) are about as likely to struggle to explain AI’s role in an incident as teams without it (75%).
- Teams with complete audit tracking 72%
- Teams without complete audit tracking 75%
Tracking alone is not explanation. A record of what AI did is not the same as an account of why it broke.
But when something related to AI-generated code or decisions goes wrong, respondents say a person or team still has to answer for it. Having a clear trail to point to is so crucial for this reason. Blind spots are particularly precarious for organizations in regulated industries, where compliance requirements around AI-generated outputs are tightening.
- The team or team lead 42%
- The developer who wrote or accepted it 29%
- The organization broadly 19%
- The vendor of the AI tool 6%
- Not sure 2%
- No clear owner 2%
Accountability does not shift when the incident cannot be explained: among those who have been unable to explain AI’s role, 42% still name the team or team lead and 29% the developer — the same as everyone else.
05 — AI agents
Adoption of AI agents has outrun oversight
At the most advanced stage of the AI-disrupted SDLC, AI systems can work autonomously with limited-to-no human intervention. AI agents that can write code, run and generate tests, and modify production configurations with limited human review are already operating in a meaningful number of engineering environments.
are using or evaluating AI agents in their development and testing environments
- In dev, testing and production 43%
- In dev and testing, not production 26%
- Piloting or evaluating 22%
- Only production 5%
- Not using AI agents 4%
48% are running agents in production.
Agents have moved from pilot to production faster than oversight has kept pace with them. A fundamental shift happens when agents operate autonomously across development and production pipelines: teams cede upfront control of requirements and specifications and only subsequently learn what the agent’s code has actually delivered.
- A human reviews 60% or less of what their agents produce 44%
- A human reviews more than 80% of agent work before use or merge 25%
- Oversee agents with a purpose-built commercial tool 34%
are confident their agents are adequately reviewed and overseen.
That confidence doesn’t line up with what review actually looks like. Nearly half of respondents (44%) say a human reviews 60% or less of what their agents produce. That unreviewed work is exactly the kind of blind spot that could be driving the failures this report has uncovered. Code nobody independently checks is also code nobody can later explain when it breaks.
Only 25% of teams have a human review more than 80% of their agents’ work before it is used or merged.
Only a quarter of teams are checking most of what their agents produce, and leaders are still nearly unanimous that their AI agents are adequately overseen — a case of blind trust in governance that isn’t actually happening. It begs the question of whether organizations know what real governance actually looks like.
A company’s governance approach has a measurable effect on what ships. Organizations with more extensive oversight of agent output are shipping fewer failures: 49% of teams that review more than 60% of their agents’ work have shipped AI code that later failed, against 58% of teams that review 60% or less.
- Teams reviewing more than 60% of agent work 49%
- Teams reviewing 60% or less of agent work 58%
Review narrows the odds, but it doesn’t close them — highlighting the need for testing and validation.
Even among teams reviewing more, roughly half still shipped a failure. Review narrows the odds, but it doesn’t close them, highlighting the need for testing and validation.
When asked more broadly how their organizations oversee and govern AI agents in their development and testing environment, only 34% say they use a purpose-built commercial tool. The rest use internal guardrails they built themselves, working on a case-by-case basis, or operating with no formal oversight. Given this, it’s no surprise AI agent governance is one of the biggest areas for improvement. Closing that gap will take more than internal experimentation. It calls for the guardrails an experienced software quality partner can help put in place, so AI agents act safely and as intended.
06 — Autonomous testing
Autonomous testing could alleviate the AI coding bottleneck
Despite the intensity of today’s AI-driven output, 63% of teams say their testing and verification capacity is keeping pace with AI’s code volume. Less surprising is that 36% are starting to fall behind or already have.
- Testing and verification capacity is keeping pace with AI code volume 63%
- Starting to fall behind, or already behind 36%
The data shows that being able to match the speed of development enables organizations to be better positioned at the deployment stage.
- Scaling AI-generated testing 57% 46%
- Leaning more on AI to check its own output 45% 37%
- Narrowing test coverage 23% 29%
- Shipping with reduced review 23% 28%
Most respondents see autonomous testing, a step beyond AI-assisted testing, as a promising remedy for the chaos of today’s AI-accelerated SDLC. With this software quality approach, AI agents independently generate, execute, adapt, and report on tests — without requiring manual scripting. Humans stay in the loop, reviewing what the agent flags to provide the judgment agents can’t. It operates as a continuous, living function of the software delivery process, not a black box that grades its own work.
of software experts say autonomous testing would improve their ability to keep up with the volume of AI-generated code.
Leaders are more bullish on autonomous testing than practitioners, representing one more area where the confidence gap shows up. Nearly nine in ten (89%) leaders say autonomous testing would help them keep up with AI code volume compared with 77% of practitioners.
- Would help us keep up with AI-generated code volume 89% 77%
Yet while autonomous testing offers so much promise for organizations, the teams thinking about adoption face different impediments. Topping the list is trust in the results. About a quarter of respondents (23%) name it as the top barrier to adopting or scaling autonomous testing, nearly double the 12% who name cost.
- Trust in the results 23%
- Integration with existing tools 18%
- Governance or compliance concerns 16%
- Internal resistance to autonomous agents 12%
- Lack of skills or expertise 12%
- Cost 12%
- No major barriers 6%
Human oversight can help organizations overcome trust concerns. QA engineers and leads can interpret results, connect agent findings to business context, and make the quality decisions that agents can’t make.
07 — Conclusion
Turning confidence into something you can stand behind
Gaps in the areas of trust and validation, governance and accountability, AI agent use, and scale come down to the same root cause: an operating model built for a slower kind of software delivery that hasn’t caught up to what AI-generated development actually demands.
But what’s unexpected is that many organizations, and leaders especially, don’t seem aware of their gaps, pushing forward with little proof to back up their confidence. The combination of increasing trust and AI failures can even have personal impacts on livelihoods: it’s humans who bear the brunt of AI errors.
As for the confidence gap between practitioners and leaders, what will close it is autonomy, backed by human oversight.
Fortunately, many organizations are already handling this transition well. Rather than eliminating risk, which is an unattainable aspiration, many have built quality and verification into how AI development works. It’s not a gate at the end of the process but an ongoing capability running alongside it.
Continuous, measurable assurance that software works as intended, application integrity, is what turns confidence into something those closest to the code can stand behind — not just something they and leadership hope holds.
08 — Sidebar
In their own words
Leaders and practitioners share their biggest concerns about AI-generated code and applications.
AI-generated code increases development speed faster than we can independently test, verify, and govern it.
My biggest worry is that AI code is being churned out much faster than our testing team can check it, leaving us vulnerable to hidden bugs slipping into live systems.
AI-generated code may contain hidden defects that go undetected, leading to production failures and damaging user trust.
Honestly, my biggest concern is trust. AI can write clean looking code, but I don't always know what's happening under the hood.
Hidden technical debt multiplies as AI refactors without architectural governance.
Limited human oversight over AI code increases the chance of non-compliant logic entering production.
Maintaining accountability when automated decisions introduce unexpected operational risks concerns me most.
Limited transparency of AI-written code raises the chance that edge-condition bugs reach production.
It's hard to fully trust without human review.
09 — Regional sidebar
How the U.K. is managing the AI-driven SDLC
The U.K. numbers tell the same story as the global report: trust in AI-generated code and testing is running ahead of what teams can prove.
Confidence doesn’t track experience, and most of the teams keeping pace are trading review, coverage, or release timing to get there. That’s not a case for slowing AI adoption. It’s a case for building continuous, independent verification into the pipeline, so confidence has something real to stand on.
Blind trust in AI holds in the U.K., too
AI hasn’t just added more code to the pipeline; it has become a standard part of how that code gets checked as well. Just over half of U.K. organizations (54%) say AI writes or accelerates at least 41% of their code, and a nearly identical share (55%) say AI generates or maintains at least 41% of their test coverage.
- AI writes or accelerates at least 41% of their code 54%
- AI generates or maintains at least 41% of their test coverage 55%
With code and applications now shipping at AI-driven speed, 68% of U.K. respondents are at least somewhat concerned that application quality is suffering, including 37% who are very or extremely concerned. Yet 60% still have a lot or complete confidence that AI-generated code and applications behave as intended.
That confidence hasn’t been earned by results. Among U.K. respondents with a lot or complete confidence in AI-generated code, 36% say their organization has shipped AI code that later failed in production.
- At least somewhat concerned application quality is suffering 68%
- Very or extremely concerned application quality is suffering 37%
- A lot or complete confidence AI-generated code behaves as intended 60%
- Of those confident respondents, shipped AI code that later failed 36%
Confidence hasn’t been earned by results: more than a third of the most confident U.K. respondents have shipped AI code that later failed.
Verification is real, but it clusters late and isn’t always audit-ready
Most U.K. teams check AI’s work, but the checks cluster near the end of the pipeline rather than the start: 51% review more than 60% of AI-produced code before it’s committed, compared with 64% who test more than 60% of it before release.
Belief in AI’s ability to check itself hasn’t displaced human review as the default validation method. Slightly more respondents rely on an independent reviewer (56%) or the original prompter (54%) than on AI validating its own output, which only 12% cite as one of their methods and just 1% rely on exclusively.
- Review more than 60% of AI code before commit 51%
- Test more than 60% of AI code before release 64%
- An independent reviewer 56%
- The original prompter 54%
- AI validating its own output (as one of several methods) 12%
- AI validating its own output, exclusively 1%
Respondents could select more than one method, so shares do not sum to 100%.
AI agents have reached production faster than oversight has caught up
Documentation is the weak point. Only 58% of U.K. organizations say AI outputs are tracked and documented to a standard that would satisfy compliance or audit requirements. That lack of governance can lead to major software or compliance issues that affect customers and the business — and 40% already report an incident where they couldn’t fully explain how AI contributed to a bug or production issue.
AI agents are already mainstream in the U.K.: 49% of organizations run them in production, and 97% are using or evaluating them. Oversight hasn’t scaled at the same rate — only 23% review more than 80% of what agents produce before it reaches production.
Most of that oversight is homegrown. Internal guardrails that teams built themselves are the most common model (46%), ahead of purpose-built commercial tools (39%). Confidence in that oversight differs slightly among leaders and practitioners: 96% of leaders are very or somewhat confident their agents are adequately reviewed, compared with 91% of practitioners.
- Run AI agents in production 49%
- Using or evaluating AI agents 97%
- AI outputs tracked to a compliance or audit standard 58%
- Had an incident they could not fully explain 40%
- Review more than 80% of agent output before production 23%
- Internal guardrails teams built themselves 46%
- Purpose-built commercial tools 39%
- Very or somewhat confident agents are adequately reviewed 96% 91%
Keeping pace still means trading something
Nearly three-quarters of U.K. respondents (73%) say their testing and verification capacity is keeping pace with AI’s code volume. But for most, something had to give to keep pace: 63% of teams keeping pace still report at least one deployment trade-off, whether that’s reduced review, narrower test coverage, delayed releases, or knowingly accepting more risk. Teams increasingly face this choice — give up quality or give up speed. In reality, speed and quality are both attainable.
Most see autonomous testing as the way to make that possible: 90% agree it would help their organization keep up with AI’s code volume, and only 3% disagree. What’s holding it back isn’t a single obstacle but a cluster of related ones. Integration with existing tools (20%), trust in the results (20%), and governance or compliance concerns (18%) are roughly tied for the top spot.
say autonomous testing would help their organization keep up with AI’s code volume.
- Integration with existing tools 20%
- Trust in the results 20%
- Governance or compliance concerns 18%
Roughly tied for the top spot — what holds autonomous testing back in the U.K. isn’t a single obstacle but a cluster of related ones.
10 — Industry comparisons
Not every industry is running the same race
Telecommunications and media lead on AI adoption and absorb the most failures because of it, while healthcare and banking move slower and carry more of the compliance weight. The following findings paint the picture.
- Telco & media 80%
- Retail & e‑commerce 67%
- Banking & financial 62%
- Software & SaaS 62%
- Healthcare 58%
- Healthcare 69%
- Software & SaaS 68%
- Retail & e‑commerce 66%
- Banking & financial 59%
- Telco & media 50%
- Telco & media 59%
- Software & SaaS 40%
- Banking & financial 38%
- Retail & e‑commerce 38%
- Healthcare 36%
- Telco & media 73% 68%
- Retail & e‑commerce 65% 47%
- Banking & financial 61% 46%
- Software & SaaS 59% 39%
- Healthcare 59% 32%
Telco also runs the most AI-written code: 80% at 41% or more, against 58% to 67% in every other industry.
- Software & SaaS 59%
- Retail & e‑commerce 59%
- Banking & financial 55%
- Telco & media 55%
- Healthcare 42%
- Telco & media 88%
- Retail & e‑commerce 82%
- Banking & financial 79%
- Software & SaaS 78%
- Healthcare 74%
Combines “yes, for all of it” and “for most of it”.
| Option | Software & SaaS | Banking & financial | Telco & media | Healthcare | Retail & e‑commerce |
|---|---|---|---|---|---|
| Scaling AI-generated testing | 57% | 53% | 49% | 54% | 45% |
| Leaning more on AI to validate its own output | 44% | 42% | 37% | 39% | 41% |
| Knowingly accepting more risk | 36% | 33% | 32% | 30% | 35% |
| Narrowing test coverage | 26% | 23% | 30% | 20% | 25% |
| Shipping with reduced review | 24% | 23% | 30% | 19% | 26% |
| Delaying releases | 18% | 21% | 26% | 22% | 24% |
| None of the above | 7% | 4% | 1% | 6% | 5% |
Multi-select, so columns sum past 100%. Bars are the share of each industry selecting that option. Telco is highest on narrowing coverage and reduced review, at 30% each against 20% and 19% in healthcare. Differences of fewer than about 8 points between industries sit inside the 95% intervals.
- Healthcare 57%
- Banking & financial 55%
- Software & SaaS 47%
- Retail & e‑commerce 46%
- Telco & media 44%
Share saying NOT fully.
- Telco & media 68%
- Retail & e‑commerce 47%
- Banking & financial 45%
- Software & SaaS 41%
- Healthcare 40%
Telco is also the industry most sure it can identify AI-written code (88%).
- Software & SaaS 52%
- Retail & e‑commerce 51%
- Healthcare 50%
- Telco & media 45%
- Banking & financial 42%
Share running agents in production. Banking also doubts AI most (19%).
11 — About the report
Who we surveyed
SmartBear surveyed 1,436 technology professionals whose organizations use AI in development. Each owns testing and quality, either as a practitioner or a senior leader at organizations with more than 500 employees and more than $50 million in annual revenue. The survey was conducted in Q3 2026.
- Senior leaders
- 51%
- Practitioners
- 49%
- Software & SaaS 31%
- Banking & financial services 22%
- Telecommunications & media 18%
- Healthcare 15%
- Retail & e‑commerce 14%
Industry mix of the 1,436 respondents.
- United States
- 72%
- United Kingdom
- 19%
- Rest of world
- 9%
Application Integrity
Give your confidence something to stand on
Continuous, independent verification built into how AI development actually works — not a gate at the end of the process.