Only 9% Trust AI Pentesting Alone. Here's What the 91% Want
A security leader who signed an autonomous penetration testing trial in mid-2025 was doing the obvious thing. The category was new, the demos were extraordinary, and roughly three in ten of their peers said they would happily let automation carry the whole testing load. The same person signing the same contract today is going against the room, because that number is now nine percent.
Twelve months is not long enough for a technology to get worse. The models did not regress, the agent frameworks got better at chaining steps, and adoption of AI inside testing workflows went up, not down. A twenty point collapse in stated trust cannot be a verdict on capability. Something else moved.
Here is the thesis, stated plainly: the market did not lose faith in AI doing the testing. It lost faith in nobody being accountable for the result. Those are different failures with different fixes, and conflating them leads buyers to reject tooling they should be adopting. This post separates them, then hands you what the headline number does not: a four level autonomy specification, drawn from a standard published this year, that you can paste into your next penetration testing RFP.
What the 9% Number Actually Measured (and What It Did Not)
The figure comes from Cobalt's AI and Pentesting Pulse Report 2026, published 25 June 2026. The sample was 455 security and engineering leaders at organisations with more than 500 employees, spanning software, healthcare, financial services and information services. It is a real sample with a disclosed sector mix, which already puts it ahead of most vendor surveys.
The question it asked was narrow and specific: would your organisation rely entirely on automation for its penetration testing needs? In 2025, 29% said yes. In 2026, 9% said yes. Meanwhile 47% said they now prefer a hybrid model, humans and automation together, over either extreme.
That is the whole finding. It measures willingness to remove humans from the loop. It is not a benchmark of how many vulnerabilities an AI agent can find, not a comparison of detection rates, and not a claim that AI testing tools got worse.
Set that as a guardrail before you read another paragraph, because the same report shows AI usage climbing over exactly the period when stated trust in full autonomy fell. Adoption of AI somewhere in the testing workflow rose among both professional penetration testers and independent researchers. The people closest to the work are using more AI and trusting unsupervised AI less, simultaneously, and that combination is the actual story.
- It is not a measure of AI capability. Detection rates, coverage and exploit-chaining ability are not what was asked about. Nothing in this survey benchmarks what an agent can find.
- It is not about AI-assisted testing. The question was about relying entirely on automation. AI assistance sits everywhere else on the spectrum and is now near-universal among practitioners.
- It is a preference survey, not an outcome study. It records what 455 people say their organisation would be willing to do, not what those organisations purchased or what happened when they did.
Why Confidence Collapsed: Three Mechanisms
A twenty point swing needs a mechanism, not a mood. Three things happened between the 2025 survey and the 2026 one, and they compound.
False Negatives, Not False Positives
The finding that should have led every write-up of this report got buried: 78% of organisations said fully automated scanning tools had missed critical vulnerabilities in their environment.
Most buyers arrive at automated testing braced for the opposite problem. False positives are the well known tax on scanning: the tool cries wolf, an engineer spends two hours proving there was no wolf, and everyone grumbles about signal to noise. That failure is annoying, expensive, and completely visible. You can measure it and hold a vendor to a number.
False negatives behave nothing like that. A tool that reports nothing looks exactly like a clean environment. There is no alert to triage, no ticket to close, no metric that moves. The dashboard is green because the scanner had nothing to say, and "nothing to say" and "nothing to find" render identically. You learn the difference from an attacker, from a customer's security questionnaire, or from an auditor asking why your test scope excluded the one application that got breached.
That asymmetry is why 78% is the number that moved the market. It is not a satisfaction score. It is a count of organisations that discovered, after the fact, that a green result had been meaningless.
Coverage Is Narrower Than the Demo
The second mechanism is scope. Autonomous tools perform best where a known, working, public exploit exists, because that is the terrain their reasoning and their payload libraries were built for. They stall where it does not: novel business logic, chained conditions, authorisation flaws that depend on understanding what a role is supposed to be able to do.
How much of a real estate that leaves covered is a different conversation with its own mechanics, and it depends far more on what your applications do than on any headline coverage percentage. We have written about where automated testing is genuinely sufficient and where it is not; the point here is narrower. The demo runs against the terrain the tool is good at. Your estate is not that terrain.
The Expectation Reset
The third mechanism is the simplest and probably the largest. In 2025, buyers answering this survey were describing an expectation. Autonomous testing was a category description, a roadmap slide, a pilot someone's team was about to start. Optimism about a thing you have not run yet is cheap.
In 2026, the same buyers were describing trial results. They ran the tools, compared the output against a human test on the same target, and watched what happened when the agent hit an authenticated workflow it did not understand. The gap between a category description and a completed pilot is, near enough, the entire twenty points.
None of these three mechanisms is an argument that AI cannot test. Held honestly, fully autonomous testing has a genuinely strong column, and pretending otherwise is how buyers talk themselves into paying human rates for work a machine does better.
- Runs in hours, not the two to six weeks a scheduled human engagement takes
- Cadence can match your deploy frequency rather than your budget cycle
- Cost per cycle is low enough that testing frequency stops being a procurement decision
- Excellent breadth against known patterns, public CVEs and misconfiguration classes
- No scheduling lead time, no waiting for a tester to come free
- Silent false negatives: a clean report and an untested target look identical
- Stalls where no public exploit exists, which is where business logic and authorisation flaws live
- No accountable signer, so nobody's name and credential stand behind the result
- Scope drift risk is unbounded unless boundaries are enforced technically, not just agreed
- Produces no attestation artefact an auditor or enterprise customer will accept on its own
Read that ledger again and notice what none of the cons says. Not one of them says the engine reasons badly. They say you cannot see where its coverage ended, you cannot bound where it goes, and you cannot produce anyone who will stand behind what it found. That is a different complaint from "the technology does not work", and it takes a different fix.
The Distinction the Headline Destroys: Assisted vs Autonomous
Read the 9% as "AI is bad at penetration testing" and you will draw precisely the wrong operational conclusion, which is to slow down AI adoption in your security programme at the moment your peers are accelerating it.
The survey asked about relying entirely on automation. That is one specific configuration at the far end of a spectrum. AI assistance sits everywhere else on that spectrum, and it is now close to universal among practitioners. Two thirds of professional testers use it. Four fifths of independent hackers use it. Those are not people who think the technology is unreliable. They are people who have found the tasks it is excellent at: enumerating an attack surface faster than a human can type, correlating findings across hosts, drafting reproduction steps, triaging thousands of low signal results down to a reviewable set, and translating a raw finding into language a developer can act on.
Think of it as a dial rather than a switch, with three positions that behave very differently in production.
Our guide to agentic penetration testing walks through how these systems actually plan and execute. The relevant point for this post is that the spectrum is not a marketing framing. As of 2026 it is a specified one, with numbered levels, and that changes what you can ask for.
See what your external surface exposes, mapped to the controls it touches.
Run a free External Security Check →What the 91% Actually Chose (and What "Hybrid" Has to Mean to Be Worth Anything)
Forty-seven percent said they prefer hybrid. Which would be useful, except that "hybrid" is a word any vendor can put on a page. On its own it tells you nothing about how much of the judgment a person actually did, or at which step they did it.
A preference is only actionable once you convert it into properties you can test at procurement. Read across what buyers rejected and what they kept, and the hybrid model resolves into four:
One: machines cover breadth, humans own judgment. Automation is better than people at scale work. It enumerates faster, runs more checks, and never gets bored on host 400 of 500. Humans are better at the four things automation reliably fumbles: business logic flaws that require knowing what the application is for, chaining several medium findings into one critical path, validating that a reported finding is real, and framing impact in terms a board understands. A hybrid model that does not explicitly assign those four to a person is not hybrid.
Two: there is a named, independent person accountable for the output. Not a company, not a platform, a person, with credentials, who signed, and who does not work for the organisation being tested. Independence is not a nicety bolted onto accountability, it is the reason the signature clears anything. PCI DSS 11.4 names an independent qualified tester in the standard's own words, and the security questionnaire your sales team keeps forwarding asks for a third-party pentest for the same reason. This is the property the 9% number is really about, and it fails in two directions rather than one. Autonomous testing removes the signature. A tool your own team points at your own application removes the independence, however good the tool is, because that is self-assessment with better tooling. Both failures are load bearing for reasons that have nothing to do with detection quality.
Three: the evidence is reproducible by a third party who was not present. Somebody who did not run the test, and who has no reason to take your word for it, must be able to follow the report and reproduce the finding: captured requests and responses, timestamps, and steps precise enough to re-execute. That is the property auditors care most about, and we cover it in depth in our piece on what auditors accept from AI-generated pentest reports.
Four: the scope was agreed before the test, not inferred during it. An agent that discovers an adjacent subnet and decides to explore it has exceeded its authorisation, regardless of how useful the finding is. Scope is a legal boundary before it is a technical one.
Those four properties reduce to four questions you can ask in a single scoping call.
Notice what all four have in common. Not one of them is about the engine. Every one is about the artefact the engine produces and the controls around its production. The market did not re-rate the model. It re-rated the report.
That is also why the shift shows up hardest in regulated buyers. If your test exists to satisfy SOC 2 evidence requirements or PCI DSS 4.0 requirement 11.4, the deliverable is the product. An unsigned, unreproducible finding is not cheaper testing. It is a document your auditor will hand back.
Which raises the obvious question: if these four properties are what buyers now want, has anyone written them down in a form you can cite in a contract? As of this year, yes.
The Standard Nobody Is Connecting to the Stat: OWASP APTS
Here is the part that has gone almost entirely unremarked in the coverage of this survey. In the same year that buyer sentiment converged on "yes to AI, no to unaccountable AI," OWASP published a standard that specifies exactly that position, requirement by requirement.
The Autonomous Penetration Testing Standard (APTS) is a governance framework. OWASP is explicit that it is not a testing methodology and does not replace PTES, the Web Security Testing Guide, or OSSTMM. Those tell you how to test. APTS tells you how to govern a system that tests on its own: what it must be prevented from doing, who must be able to stop it, and what it must be able to prove afterward. The structure is worth knowing in detail, because the detail is what makes it usable in procurement.
APTS defines 173 tier-required requirements across eight domains: Scope Enforcement (26), Safety Controls and Impact Management (20), Human Oversight and Intervention (19), Graduated Autonomy Levels (28), Auditability and Reproducibility (20), Manipulation Resistance (23), Third-Party and Supply Chain Trust (22), and Reporting (15). A further 19 advisory practices are documented in an appendix but do not count toward conformance at any tier.
Conformance comes in three tiers. Tier 1 Foundation covers 72 requirements. Tier 2 Verified covers 157 cumulatively. Tier 3 Comprehensive covers all 173. A tier claim is all-or-nothing: OWASP requires every MUST at the claimed tier and all lower tiers, with no deviation. The standard deliberately does not tell you which tier to demand, leaving that to each buying organisation, though it describes Tier 3 as meeting the highest assurance bar for critical infrastructure. Our own recommendation, not OWASP's, is that regulated buyers treat Tier 2 as the floor and reserve Tier 3 for critical infrastructure and for any deployment running at the highest autonomy level.
Those levels are the second axis. APTS specifies four autonomy levels: L1 Assisted, L2 Supervised, L3 Semi-Autonomous and L4 Autonomous. At L1 the operator commands every action, one technique per command, with no chaining and no inference. L2 chains techniques within a phase but stops for operator approval at every phase boundary. At L3 the operator sets boundaries and intervenes on exceptions while the platform executes complete attack chains inside them. L4 runs multi-target campaigns with dynamic scope, and this is the level buyers most often misread: it does not mean nobody is watching. APTS still requires periodic operator review, dedicated monitoring staff covering the platform's operating hours, kill-switch authority delegated across all of those hours, and a tested incident response procedure before you may claim it. Each level carries its own containment requirements and oversight obligations rather than being a label on a slider.
The two-axis model exists specifically so a vendor that is highly capable and poorly governed cannot hide the second fact behind the first.
The Human Oversight and Intervention domain is where the standard maps most directly onto why buyer confidence fell. Its 19 requirements cover approval gates before exploitation and lateral movement, real-time monitoring of agent activity, defined approver-timeout behaviour chosen in advance rather than improvised (exploitation and lateral movement default to deny, unexpected findings default to pause and isolate), kill-switch authority and who holds it across a named chain from operator to CISO, dual control requiring a second independent approver for the most severe exploitation, escalation paths for findings the system did not expect, operator qualification standards tied to the autonomy level assigned, and shift-handoff procedures for teams running 24/7 coverage.
Read that list against the failure modes people actually hit with autonomous agents and it reads like a post-incident action list. It is not hypothetical: we covered what happens when an AI agent's containment boundary is the only thing standing between a test and production, and every control in that domain exists because some version of that scenario has already played out.
The point worth carrying away is this. In June 2026, the 91% looked like a sentiment. It is now a specification with requirement IDs, domain counts, and conformance tiers. Sentiment you can argue with. A requirement ID goes in a contract.
Write It Into the RFP: An Autonomy Level Specification
None of the above helps if it stays a reading exercise. Here is how to turn it into a procurement artefact this week.
Here is a paragraph you can lift directly into an RFP:
The vendor shall specify the maximum autonomy level (OWASP APTS L1 through L4) applied to each asset class in scope, and shall state its conformance tier under OWASP APTS. Testing against assets classified Critical shall not exceed the autonomy level agreed in writing prior to engagement. All findings shall be accompanied by reproduction steps, captured request and response evidence, and a named, credentialed reviewer accountable for the report, together with a coverage statement identifying assets and functionality not tested.
Three sentences. They do more procurement work than any demo.
If you are weighing a self-hosted open-source agent against a managed service, the same specification applies to both, and it is usually where the real cost difference shows up. We compared the two models directly in our analysis of open-source versus managed autonomous testing: the tooling is often comparable, and the governance burden is not.
For teams running this as a continuous programme rather than a point-in-time engagement, the autonomy specification slots directly into a CTEM validation stage, where it governs what the continuous layer is permitted to do between formal assessments.
Before you build a procurement decision on one survey, though, it is worth knowing what that survey can and cannot support.
Read the Source Critically (Including Who Published It)
A post that leans this hard on one survey owes you the caveats.
It measures stated preference, not behaviour. The sample is 455 respondents at organisations with more than 500 employees, with a disclosed sector mix, self-reporting what their organisation would be willing to do. That tells you where the market's stated position sits. It does not tell you what those organisations actually purchased, and stated preference and signed contracts diverge routinely, usually in the direction of whatever is cheaper.
The publisher has a stake in the answer. Cobalt is a penetration testing as a service company whose model puts human testers in the loop, so a finding that buyers want humans in the loop is commercially favourable to it. That is worth stating plainly, and it is not an accusation: vendor-published research is most of the research this industry has, and the methodology here is more transparent than most. Note also that the two headline numbers differ in kind. The 9% is a preference. The 78% false-negative figure is experiential, a report of something that happened to the respondent's organisation, and experiential data is harder to steer with question wording. If you discount one, discount the preference number first.
The publisher then built to its own finding. On 23 July 2026, roughly a month after the report, Cobalt launched an autonomous penetration testing product of its own, announced with its own pentesters directing every engagement. That is arguably worth more to you than the survey is. A company publishes evidence that buyers want a human answerable for the result, then puts its engineering budget behind autonomous capability with a human directing it, and the two signals agree. The market conclusion was never "no autonomy." It was "no unaccountable autonomy." Stated preference and capital allocation landing in the same place is a stronger read than either on its own.
Which leaves you with the only test that matters. Your decision about autonomy level should turn on your asset classes, your regulator and your audit calendar. It should not turn on the aggregated preference of 455 people who do not work at your company and have never seen your production environment. Use the survey to understand the direction of travel. Use the specification to make the decision.
Where CyberOrbit Sits
Our model is straightforward to state, and the fair way to read it is against the four questions above rather than against our description of it.
You configure which systems you want tested. CyberOrbit's platform scopes and schedules the assessment, runs it, and captures evidence as it goes: real requests, real responses, timestamps, reproduction steps. An experienced security professional, independent of your organisation, then reviews the findings and signs the report you select. Breadth from the machine, judgment and accountability from a named human who does not work for you.
Where that model stops is worth saying as plainly as where it works. A reviewer who validates findings and signs the report is not the same thing as a specialist spending a fortnight modelling one application's business logic, and novel logic flaws in a complex workflow still need that. We handle the systematic 80% and put an independent signature on the end of it. For the rest, we would tell you so rather than sell around it.
We did not adopt this model in response to the survey; it is the model the survey describes buyers converging toward. We are not claiming conformance to any OWASP APTS tier, because that is a formal assessment and we will say so when one exists.
Step four of that specification applies to us as much as to anyone else, and it is worth being direct about it. "A security professional" is a role, not a name, and this paragraph is marketing copy rather than evidence. The name, the credential and the scope of what that person personally reviewed belong on the report and in the scoping call, which is where you should insist on having them from any vendor including this one. Put the four questions to us and we will answer them in the same order you would ask a competitor. The pricing page sets out the scope and sign-off model in full.
Frequently Asked Questions
Why do only 9% of security leaders trust AI pentesting alone?
What did Cobalt's AI and Pentesting Pulse Report 2026 find?
Is AI penetration testing reliable in 2026?
What is the difference between autonomous and AI-assisted penetration testing?
What is OWASP APTS and what does it require?
What are the four autonomy levels in autonomous penetration testing?
How do I specify an autonomy level in a pentest RFP?
Sources
- 78% of Security Teams Experience Critical False Negatives From Automated Scanning Tools, Business Wire, 25 June 2026, the launch release for Cobalt's AI and Pentesting Pulse Report 2026 (n=455)
- Decline in Confidence in Autonomous Penetration Testing, Dark Reading
- Trust in AI Vulnerability Scanning, Infosecurity Magazine
- OWASP Autonomous Penetration Testing Standard (APTS), OWASP, the source for the 173 requirements, eight domains, four autonomy levels and three conformance tiers
- OWASP Web Security Testing Guide, OWASP, one of the testing methodologies APTS complements rather than replaces
- Cobalt adds autonomous pentest to scale application security testing, Help Net Security, 23 July 2026