Skip to main content

Agentic penetration testing: what a regulated EU buyer should actually check

Yanis Grigy, CEO6 min read

Agentic penetration testing is now a procurement category, not a demo.

From April to June 2025, XBOW ranked first on the US HackerOne leaderboard, the first time an autonomous system topped every human researcher in that program. Over that window it submitted nearly 1,060 vulnerabilities: 54 critical, 242 high, 524 medium, and 65 low. Independent reporting confirmed the ranking and noted the nuance, first in the US, sixth globally.

If you buy security testing for a regulated business, the headline is not "AI beat humans." The headline is that two things changed in 2026 that your RFP probably has not caught up with: analysts renamed the category, and OWASP started treating the testing agent itself as an attack surface.

Gartner renamed the category in March 2026

On 24 March 2026, Dhivya Poole, Mitchell Schneider, and Eric Ahlm published the Gartner Market Guide for Adversarial Exposure Validation, which defines AEV as "technologies that deliver consistent, continuous and automated evidence of the feasibility of an attack."

The important part for a buyer is what it replaces. AEV formally absorbs two earlier Gartner categories: breach and attack simulation, and automated penetration testing and red teaming technology. Where BAS proved your detection controls fire and automated pentesting proved a vulnerability was exploitable, AEV covers both and asks for proven exploitability against your actual environment rather than a theoretical risk score. Gartner projects that by 2029, 60% of organisations will run a structured exposure validation practice as part of CTEM.

Practical consequence: if you write "automated penetration testing tool" in a requirements document, you are shopping in a category the analyst who defines it has retired. Ask vendors where they sit against AEV, and specifically whether they produce evidence of a successful attack path or a list of findings.

The lanes still decide what gets covered

The market is not one thing. A 2026 landscape review by Astra sorts the tools by where they actually work: web applications and APIs (XBOW, Terra, Aikido, Astra), internal networks (NodeZero from Horizon3.ai), and external attack surface (Hadrian). What makes them agentic rather than a faster scanner is that they reason about results, adapt, and chain techniques instead of running a fixed rulebook. For the mechanics, we wrote the agentic pentest explainer separately.

A tool that is excellent on web apps can be blind on your internal network. "Autonomous" is a capability claim scoped to a surface, not a guarantee across your whole estate.

The tool you buy is itself an agentic application

This is the due-diligence axis most 2025-era shortlists missed. On 9 December 2025, OWASP published its top 10 for agentic applications, built by a global community of security experts to cover autonomous systems moving from pilot into production. The ten ranked risks include ASI02 tool misuse and exploitation, ASI03 identity and privilege abuse, ASI06 memory and context poisoning, and ASI10 rogue agents.

Now read that list as a buyer. An agentic pentest engagement means granting a vendor's autonomous agent credentialed, adversarial access to your estate, with a mandate to chain techniques and escalate. That is an agentic application operating on your production systems, and OWASP's top three risks describe precisely what goes wrong when one exceeds its brief.

An autonomous agent that finds everything and attests nothing has moved your problem, not solved it.

So ask the containment questions. How is scope enforced, by configuration or by a hard technical boundary. Does the agent hold its own identity with its own least-privilege credentials, or does it borrow a human's. Is there an immutable log of every action the agent took, sufficient to reconstruct the engagement for an auditor or after an incident. What is the documented behaviour when the agent finds a path out of the agreed scope.

Breadth is still not attestation

Speed and coverage are solved. Audit-grade assurance is not.

The Astra review is blunt that human pentesters remain essential for complex gaps and novel techniques that have not been automated, and it flags that an autonomous agent can miss context-specific rules that cause a compliance breach. A supervisor under DORA, NIS2, or an ISO 27001 audit does not accept "the AI found 524 medium issues." They accept a scoped, validated report tied to your critical functions, with a human accountable for the method. If DORA is your driver, the testing requirements are specific about independence and scope; under NIS2 the question is cadence rather than a single annual exercise.

The buyer's checklist

Before you sign, run six questions at every vendor, agentic or not:

  1. Where does my data go. For an EU regulated entity, a tool that processes your traffic on US infrastructure re-imports the Cloud Act exposure you were trying to remove. Sovereignty is a procurement line item, not a slogan.
  2. What does the report attest. Can it produce output your auditor accepts against your framework, or only a raw vulnerability dump you still have to translate.
  3. Who validates. Is a qualified human in the loop before findings reach a report, or is false-positive triage now your team's unpaid job.
  4. Does it map to my obligations. A finding that is not tied to a control or a critical function is noise to a compliance team.
  5. How is the agent contained. Enforced scope, dedicated least-privilege identity, immutable action log, documented behaviour on scope escape.
  6. Where does it sit against AEV. Evidence of a proven attack path against your real environment, or a probability score.

Agentic tools are a real gain on the first half of the problem. The second half, the part your auditor signs, is where sovereignty, containment, and human validation still decide the winner.

Fleuret builds agentic pentest that is EU-sovereign by design and validated before it reaches your report. If that is the lane you are buying in, See Fleuret in action.

Sources


Share this postShare on LinkedIn

Privacy Settings

This site uses third-party website tracking technologies to provide and continually improve our services, and to display information according to users' interests. I agree and may revoke or change my consent at any time with effect for the future.