New research from Pentest-Tools.com suggests that AI-assisted penetration testing tools are generating vulnerability findings faster than most security teams can verify them, creating a validation backlog that is offsetting the time AI was meant to save.
The company surveyed 158 security practitioners in June 2026, including penetration testers, security engineers, AppSec and DevSecOps professionals, consultants, and MSSP practitioners, all of whom use AI-assisted tools in their vulnerability assessment and validation work.
Among the 147 respondents who have used AI to generate findings, 87.8% said the output required significant manual validation. 61.2% said this happened with between 5% and 25% of findings, while 26.5% said more than a quarter of AI-generated findings needed rework before they could be trusted.
The capacity problem becomes more acute at scale. Asked whether their team could triage and validate more than 500 AI-generated vulnerability candidates from a single engagement, only 20.3% of respondents said they already had a workflow in place to do so. 38.6% said the volume would strain their team, and 29.7% said it would be unmanageable.
Respondents described tools returning hundreds of findings that turned out to be duplicates, false positives, unexploitable issues, or fabricated CVEs that do not exist. One practitioner summarised the pattern: an AI tool produced 300 findings, of which 250 were later found to be junk, including “potential SQLi that is not exploitable” and “AI-made-up CVEs that do not exist.” The respondent added: “I bought the tool to save time, but I did more manual work than before.”
Fabricated and hallucinated findings emerged as the leading frustration in the survey overall. Roughly 30% of free-text responses on the biggest frustration with AI pentesting tools cited false positives, hallucinated exploits, or fabricated findings, ahead of cost, integration issues, or any other complaint.
Practitioners also described a trust effect: once a hallucinated finding was caught, teams became more cautious about the rest of a tool’s output, increasing the verification workload even for genuine findings. One security manager at a mid-market company described the effect of encountering fabricated results as “confidence that turns out to be just a big lie.”
Where AI is deployed within the pentest lifecycle reflects this caution. Usage is highest in vulnerability scanning and discovery (74.1%), report writing (69%), and documentation and findings tracking (66.5%). It drops in phases requiring live judgement: exploitation and attack path chaining (36.7%), remediation validation and retesting (34.8%), and post-exploitation and lateral movement (25.3%).
Business logic testing emerged as the area practitioners consistently said AI struggles with most, ahead of exploit chaining and creativity. Respondents gave concrete examples: AI tools can identify SQL injections, but do not reliably catch that a discount coupon should only work once per customer, that adding a negative quantity to a shopping cart can produce a free purchase, or that changing a user ID in a URL can expose another customer’s data without triggering any error or anomaly.
The survey also points to AI-related testing scope expanding on a separate front. 75.3% of practitioners said they already test AI-powered systems or LLM-integrated applications as part of current engagements, with a further 17.1% expecting to do so within 12 months, bringing total adoption or planned adoption to 92.4%. Around one in three (33.5%) already include risks from unauthorised employee use of AI (“shadow AI”) in their assessment scope, while 53.2% have discussed doing so but have not yet formalised the process.
Stakeholder pressure is rising in parallel. 37.3% of respondents said internal stakeholders now expect more frequent testing than they did 12 months ago, driven by awareness of AI-assisted attacks, while a further 31.6% said stakeholders were aware of the risk shift but had not yet changed buying behaviour.
When evaluating AI-driven pentesting platforms, accuracy-related criteria outranked cost. False positive rate and signal quality were the top consideration, cited by 63% of respondents, followed by proof of exploit and verified attack paths (53%). Cost and licensing came third, at 47%.
A spokesperson for Pentest-Tools.com said: “AI is speeding up what practitioners can find. The issue is in what follows after that. When you have 300 findings from a tool and 250 of them are invalid, the time savings in discovery are spent on triage. The value of AI in pentesting lies in actionable findings.”
The survey data also suggests testing cadence, rather than organisation size, is the strongest predictor of how well a team copes with AI-generated volume. Teams testing more frequently were more likely to already have workflows for handling large numbers of findings, while teams testing less than five times a month were the most likely to describe the volume as unmanageable. Pentest-Tools.com notes that the sample size for the highest-frequency testing groups is small, and treats this finding as directional rather than conclusive.
The full survey report, “AI pentesting in 2026: why testing cadence decides who copes,” is available here.

