Skip to main content

XBOW alternative: the EU shortlist under DORA

Yanis Grigy, CEO55 min read

TL;DR

XBOW is the reference architecture for autonomous offensive security in 2026, and as of August it is no longer the best-funded name in the category. It raised a $120 million Series C in March 2026 at a valuation above $1 billion, added $35 million from strategic investors in May 2026, and runs a multi-agent architecture from Seattle. For EU regulated buyers it is still hard to onboard: DORA third-party reporting puts the vendor on a register that feeds supervisory designation, NIS2 supply-chain duties apply, and the CLOUD Act gives US authorities reach over data held by US providers regardless of physical location.

Two corrections to earlier versions of this article, because both change the decision. The EU AI Act high-risk deadline that was set for 2 August 2026 has been deferred. And the funding line above no longer makes XBOW the leader: Horizon3 raised a $250 million Series E on 3 August 2026 at a valuation above $2 billion. See the two September updates below.

A second thing changed since April, and it sits on the buying side rather than the product side. XBOW is now sold through AWS Marketplace, the Microsoft Security Store and Accenture's own delivery, which means it can enter a regulated entity without ever passing a security vendor review. The September update below covers what that does to your third-party register.

One more clock starts this month, and it applies for a different reason than the others: the Cyber Resilience Act begins asking manufacturers of products with digital elements for reports on 11 September 2026, within 24 hours of becoming aware of an actively exploited vulnerability. If you sell software or connected products into the EU, that duty is yours rather than your customer's, and the September update below covers what a 24-hour window does to the way you read an agent's findings.

The European agentic-pentest set is therefore the relevant shortlist. The five tools EU mid-market CISOs actually consider in 2026, with the trade-offs that matter:

  • Fleuret AI (FR), sovereign full-stack agentic pentest, €4k per test and continuous on quote, NIS2 / DORA-ready PDF.
  • Escape (FR), DAST and API-focused agentic engine, $18M Series A, BOLA / IDOR / access control specialty.
  • Aikido Security (BE), Series B, compliance plus AI pentest in a unified suite, broad mid-market reach.
  • Patrowl (FR), exposure management and continuous testing, named in Gartner Market Guide for Preemptive Exposure Management 2026.
  • Pentera (IL), automated security validation, mature large-enterprise comp, ~€46k/yr too steep for most mid-market.

Each section below details architecture, scope, ICP fit, and the reason a particular buyer picks one over the others.

Updates since publication

This comparison is refreshed as the rules move, so the entries below are dated. The evergreen shortlist starts further down at Why CISOs search "XBOW alternative europe".

August 2026: what changed since this comparison was published

Three things moved since April, and two of them cut against the original framing.

XBOW got a lot bigger, and it now arrives through consultancies. It raised $120 million in March 2026, led by DFJ Growth and Northzone, at a valuation above $1 billion, then took a $35 million strategic extension in May 2026 from Accenture Ventures, NVIDIA's NVentures, Samsung Ventures and SentinelOne Ventures. Accenture is folding XBOW into its own cyber delivery. Neither announcement changes the EU data-path question, but it does change how you will meet the product: through a systems integrator on an existing framework contract, not through a direct sales call. That route bypasses the vendor review the tool would otherwise trigger, which is worth flagging to whoever owns your third-party register.

The EU AI Act clock moved, and the April version of this article said otherwise. It cited high-risk obligations landing on 2 August 2026. Under the Digital Omnibus on AI, Annex III high-risk obligations now apply from 2 December 2027 and Annex I from 2 August 2028, with Article 50 transparency duties still starting on 2 August 2026. If the AI Act was the argument you were taking to procurement this year, drop it and use DORA and NIS2 instead. Those clocks did not move.

DORA third-party reporting became routine, and stricter. Financial entities filed their register of information for the 2026 cycle between 11 February and 31 March 2026, covering contracts in place on 31 December 2025, against validation rules tighter than the first year. Those registers feed the ESAs' annual designation of critical ICT third-party service providers. A pentest platform is an ICT third-party service provider. The constraint is administrative and recurring, not theoretical.

September 2026: NIS2 enforcement reached the Court of Justice

The regulatory argument you take to procurement changed again, and this time it moved in the buyer's favour.

On 8 July 2026 the European Commission referred Ireland, Spain, France and the Netherlands to the Court of Justice of the European Union for failing to notify full transposition of NIS2, with a request that the Court impose financial sanctions in the form of a lump sum and daily penalties. The transposition deadline was 17 October 2024, letters of formal notice went out on 28 November 2024, and reasoned opinions followed on 7 May 2025. The directive covers 18 sectors of high criticality.

Read what that does to your timeline rather than the headline. A member state facing daily penalties transposes fast, and national supervisors that were quiet during the drafting period start asking essential and important entities for evidence. If you are in France, Spain, Ireland or the Netherlands, the practical effect is that your national obligations firm up on a schedule you no longer control. Buying a pentest platform whose reports are not mapped to NIS2 Annex I duties is a bet that the deadline slips again.

ENISA published the third edition of its sector assessment in May 2026. NIS360 2026 reports improving maturity across critical sectors while criticality stays broadly flat, with banking, electricity, aviation, space and digital-by-default services such as telecommunications, cloud and data centres the most critical, and trust services, financial market infrastructures and aviation joining the high-maturity high-criticality group. If you sit in one of those sectors, your supervisor has a published view of what good looks like around you. That is the bar an autonomous pentest report has to clear.

So the ordering of arguments for 2026 is now: DORA third-party reporting first, NIS2 second and rising, EU AI Act last and deferred.

September 2026: XBOW now arrives through your cloud bill, not a sales call

The August update noted that XBOW reaches European buyers through consultancies. That was half the picture. Across 2026 the company built several more routes into an enterprise, and most of them never pass a security vendor review.

It listed on AWS Marketplace in February 2026, where customers can procure or renew through public and private offers and burn down AWS committed spend, including private pricing agreements. It joined the AWS ISV Accelerate co-sell programme on 13 May 2026, then reached AWS Security Competency status on 7 July 2026 and signed a Samsung SDS partnership in June 2026. On the Microsoft side it shipped a Pentest Manager Agent for Security Copilot, a Sentinel data lake connector and a Pentest Analysis Agent, announced at RSAC 2026 as a public preview and distributed through the Microsoft Security Store, Microsoft Marketplace and the Security Copilot agent gallery.

Read the committed-spend line twice, because it is the one that changes behaviour inside your own company. When a tool is bought against an existing AWS or Microsoft commitment, the purchase reads as cloud consumption on its way through finance. It rarely triggers the new-vendor questionnaire, the transfer assessment or the legal review that a direct contract would. The security team often meets the tool after it is already running.

None of that changes the obligation. DORA attaches the register entry to the contractual arrangement for ICT services and the provider behind it, not to the payment rail it travelled on. A marketplace purchase is still a contractual arrangement with a US-headquartered ICT third-party provider, and it still has to be described, with entity, country and the sub-processor chain beneath it, in the filing your entity makes each February. The only thing that changed is that nobody told your third-party risk owner it happened.

If you own third-party risk at an entity in DORA or NIS2 scope, the useful move here is dull and quick: ask your cloud administrator for the list of marketplace subscriptions running against your committed spend, and reconcile that list against your register. Autonomous pentest platforms are not the only category selling this way. They are just the category that will hold credentials to your production applications.

September 2026: your delivery route is now a supervised provider, and that cuts both ways

The marketplace section above has a second half that only became legible this year. On 18 November 2025 the European Supervisory Authorities designated the first set of critical ICT third-party service providers under DORA. The published list, issued under Article 31(9), names 19 firms, and three of them are exactly how an EU buyer meets XBOW: Accenture plc, Amazon Web Services EMEA Sarl and Microsoft Ireland Operations Limited. Google Cloud EMEA Limited, IBM, Oracle Nederland, Kyndryl, Equinix and Deutsche Telekom are on it as well, and the ESAs update the list every year, so it gets longer rather than shorter.

Designation moves those firms into direct ESA oversight with a Lead Overseer assigned per provider. In the ESAs' own words, they "will assess whether CTPPs have appropriate risk management and governance frameworks in place to ensure the resilience of the services they deliver to financial entities". That sounds like good news for a buyer whose pentest tool arrives on an AWS bill. It is not, for three reasons worth taking to the vendor call.

Designation does not move the obligation off your entity. You keep the register entry, the contractual terms, the audit rights, the incident notification path and the exit support clause, because designation does not shift responsibility away from the financial entity. "We bought it against committed cloud spend and that provider is supervised" is not an answer to a question about your own arrangement.

Routing an offensive tool through a designated provider makes your concentration answer worse. Concentration risk is measured on how much you depend on one provider for critical or important functions. If your hosting, part of your IT delivery and now the platform holding credentials to your production applications all sit behind the same designated firm, that entry reads differently than it did last year, and it is one of the few things on a register a supervisor can compare across entities.

A designated provider can be pulled into your own testing. Financial entities have to support threat-led penetration testing that involves designated providers. Buying offensive tooling through a firm that is itself in scope of that exercise is not disqualifying, but it is a coordination cost nobody prices at signature.

The useful question at the vendor call did not change, it just got harder to dodge: which legal entity signs, where does the traffic and the findings get processed, and does the answer sit inside or outside a relationship your supervisor is already reading.

September 2026: the best-funded name in the category is now Horizon3, and it has an Amsterdam address

On 3 August 2026 Horizon3 closed a $250 million Series E at a valuation above $2 billion, co-led by NightDragon and NEA, with EDBI, SAIC and Qualcomm among the strategic investors. The same report puts the company at roughly 7,200 customers, approaching $100 million in annual recurring revenue with 120% year-on-year growth, and 310,000 production security tests run without disruption. That is more capital than XBOW's $120 million Series C and $35 million extension combined, and a much longer production record.

It matters to this article for one reason. Horizon3 opened an EMEA headquarters in Amsterdam on 24 June 2026, which means the best-funded vendor in the category now has a European address to put in a pitch deck. Read the announcement before you put it in a vendor file. It describes a regional hub, hiring and customer proximity. It does not name a European legal entity as the contracting party, it does not state an EU hosting region, and it says nothing about where customer data is processed. The company is San Francisco-headquartered.

That gap is the single most common way an EU buyer gets a shortlist wrong in 2026. "European headquarters" in a press release and "EU ICT third-party provider" on a register of information are different claims, and only the second one survives a supervisory question. The test is dull: which legal entity, incorporated where, signs the contract, and which regions process the traffic and the findings. A regional office answers neither.

Horizon3 is also in a different category from XBOW for most of this article's readers. NodeZero is built for internal network and Active Directory validation, a surface XBOW's own product page does not cover. If you shortlisted it as a web-application alternative because it appeared next to XBOW on a comparison page, you have bought the wrong tool with the right logo.

September 2026: XBOW's own benchmark stopped separating XBOW alternatives, and XBOW says so

Most pages that rank XBOW alternatives still quote a score on XBEN. XBOW published that benchmark itself: 104 Jeopardy-style capture-the-flag challenges under the Apache 2.0 licence, built to mirror the vulnerability classes its own security team met on real engagements. It was a good yardstick for two years.

Read the repository today and the first thing it tells you is to stop using it. The notice says the benchmarks are outdated as of mid-2026, that industry-standard performance on the set is now approximately 100%, and that they are no longer useful for discriminating between models or attack frameworks, partly because the vulnerabilities have since been absorbed into model training. The vendor who owns the scoreboard has retired it.

The announcements have not caught up. Penetrify reports 104 out of 104 on a run dated 2 September 2026, black-box, running application only, no source code and nobody at the keyboard. Pentestkit, an MIT-licensed multi-agent framework on GitHub, reports the same 104 out of 104, every role driven by one open-weight model, with a verifier agent that refutes any candidate finding it cannot reproduce. Two more results at a ceiling its author says means nothing.

A perfect score on a saturated benchmark tells you a vendor can run the suite. It does not tell you it can test your application.

What survives saturation is everything around the score, and that is where those two runs are genuinely useful. Penetrify publishes the run harness, a results file and all 104 per-challenge transcripts, so anyone can rebuild the targets and check the claim. It also publishes the figures a buyer can hold a vendor to: 9.7 minutes average solve time per challenge, roughly $29 per pentest at its fast-tier list price, 16.8 hours of wall-clock compute for the whole suite. Pentestkit stores a full transcript and replayable evidence per challenge for the same reason.

So the benchmark question changed shape. Do not ask for the score, because everyone has it. Ask for three things: the transcripts behind the runs you are being shown, the harness so your own team can re-run the suite unattended, and the time and cost per target. A vendor that validates its findings can hand over all three inside a day. A vendor quoting a leaderboard position on a set XBOW retired is quoting the only number it has.

September 2026: two more web-application entrants, and a coverage claim with no baseline

Two announcements since July move the web-application row of the table, which is the row where XBOW actually competes.

Reflectiz launched a multi-agent web pentest product on 8 September 2026. The release describes specialised agents that discover, attack and validate across the web layer, and states the output as findings with reproduction steps and evidence, plus a coverage map of what was tested and cleared. Those two artefacts are checkable and worth asking for. The headline claim, up to ten times more coverage than conventional pentesting tools, names no baseline and no measurement method, so read it as positioning rather than as a result. The dateline is Boston, which puts it on the American side of your first procurement question.

Pentera entered the same row on 29 July 2026. Its AI-native web application testing discovers application behaviour, adapts payloads in real time and validates exploitable risk, including authenticated testing behind SSO, MFA and OAuth, running in production inside customer-defined scope and guardrails with a complete audit trail. Two caveats before it goes on a shortlist. It is in beta with selected customers, with general availability rolling out in the fourth quarter of 2026, so it is not something you can buy against an audit date this quarter. And it does not move the eligibility line: this page still files Pentera as a non-EU provider on a register of information.

The practical effect is that "shortlist by the surface you are replacing" got one line shorter. Pentera used to be the internal-network and Active Directory name in this comparison, a surface XBOW does not cover at all. Once its web application testing reaches general availability it competes with XBOW head on, which helps a buyer who wanted one vendor for both surfaces and changes nothing for a buyer whose first filter is where the contracting entity is incorporated.

The 2026 alternatives list got longer, and most of it is American

Search "XBOW alternatives" today and the names have changed since April. RunSybil, MindFort and Terra Security now appear on most comparison pages as the like-for-like autonomous application pentest set, alongside Horizon3 and Ethiack. That is a useful signal about the category, and a trap for an EU buyer, because RunSybil, MindFort and Terra Security are US-headquartered. Swapping XBOW for one of them solves nothing on the DORA register: you file a US ICT third-party provider either way.

Two of the newer entrants are worth knowing about for reasons other than eligibility.

MindFort publishes its prices. Its plans start at $199 per month for up to two pentests, $999 per month for up to four, with enterprise on quote. Whatever you think of the depth at that price, it sets a reference point, and it is a fair thing to put in front of any vendor who will not quote before a discovery call. XBOW's own product page still describes the input as a URL and does not print a figure.

Ethiack is the EU name most of these lists miss. Coimbra-based, it raised a EUR 4 million seed round in December 2024 led by Explorer Investments to build its Hackbot continuous pentest agent, combining autonomous testing with a curated ethical-hacker network. For a buyer who wants an EU legal entity and a human in the loop on validation, it belongs on the shortlist next to the French set.

The general rule this produces: when a comparison page calls something an XBOW alternative, check the headquarters line before you read the feature table. Roughly half the 2026 list fails your first procurement question.

September 2026: the alternatives you run yourself, and what they cannot hand an auditor

Every list above assumes you are buying a vendor. The other half of the "XBOW alternatives" search is people who would rather not buy one at all. The open agents you install yourself have become good enough to belong in the comparison, and for a sovereignty-driven buyer they answer the data-path question more completely than any contract can.

Two are worth knowing by name. Strix is an autonomous application pentest agent under the Apache 2.0 licence: it runs locally in Docker, it takes whichever model you point it at, and it will run against a local model served by Ollama or LM Studio instead of a hosted API. CAI, from Alias Robotics, ships as a paper as well as a codebase: the write-up by Mayoral-Vilches and colleagues releases the framework under Creative Commons Zero and reports its agents averaging 11 times faster than humans across the tasks it was measured on, at roughly 156 times lower cost.

Read those two figures as the authors' own measurements on their own benchmark, not as an independent result. That is the first thing this branch of the shortlist asks of you. Nobody else is checking the claim.

The sovereignty case is genuinely stronger than anything a vendor can offer. Point one of these agents at a model running on your own hardware and no payload, no response body and no credential leaves your network. There is no ICT third-party arrangement to record, because there is no third party. The register question that ends most conversations about a US vendor never comes up.

What you do not get is the thing your auditor asked for. A self-hosted agent produces output, not a deliverable: no signed report, no evidence chain with a date on it, no named legal entity carrying liability for a finding that turns out to be wrong, no scoping guardrail stopping the agent at your perimeter. Someone on your side has to read the agent's output and tell a real finding from a confident-sounding one, and that is precisely the senior offensive engineer a 200-person company does not have on staff. If you take the shortcut of handing the agent an OpenAI or Anthropic API key rather than standing up local inference, the data path you were trying to control comes straight back, minus the contract and the sub-processor list a vendor would at least have given you.

So the split is clean. Self-hosted agents are the right answer for an internal red team with capacity, for pre-production testing where nothing needs to be filed, and for teams building their own harness. They are the wrong answer for the report that goes to an auditor, an insurer or a board, which is the reason most people reading this page started shortlisting at all.

September 2026: the model that sets your vendor's ceiling now sits behind someone else's approval queue

The comparison below splits vendors on open-weight against frontier-API inference, and until this month that was a data-path argument. In the first week of September OpenAI turned it into a capability and continuity argument as well.

OpenAI classified GPT-6 Astra at the Critical cybersecurity level of its Preparedness Framework, the first of its models to reach that tier. The threshold is defined by what the model does without a person steering each step: find and develop working zero-day exploits across hardened real-world systems, or take a high-level objective and run a novel end-to-end attack against a hardened target. The response was to split the release. General users get a safeguarded version with stronger refusal and monitoring controls, while the less restricted capability goes to approved testers and vetted defenders.

That gate has a name and two doors. Daybreak Blue serves defensive workflows on GPT-5.6 Sol, and Daybreak Red covers authorised vulnerability research, penetration testing and exploit development on GPT-5.6 Cyber, behind a separate approval. Neither tier had Astra on launch day, and OpenAI said it would arrive "at a later date". The same announcement put $1 billion of credits over six months behind defenders who cannot pay for the tooling: critical infrastructure operators, community banks, nonprofits and open-source maintainers.

Three things follow for a shortlist, and none of them appear on a vendor comparison page.

A frontier-API vendor carries a supplier you never contracted with, and that supplier can re-tier it. Ask which model the platform actually calls, which access programme it sits in today, and what happens to your service if that access changes. A vendor who will not name its inference provider is answering the residency question badly, and is now answering a continuity question badly too.

Access cuts the other way on capability, and it deserves an honest reading. A vendor inside Daybreak Red is testing with something most buyers cannot obtain, which is a genuine argument in its favour on hard targets. The point is not that frontier access is bad. The point is that it is granted rather than owned, so it belongs in the risk column and the capability column at the same time.

It is a failure mode your DORA exit plan does not currently name. Article 28(8) wants an exit strategy that is documented, tested and reviewed periodically, and the failures people write into those plans are insolvency, breach and termination. A vendor losing a discretionary model tier degrades the service you bought without breaching a single clause you could point at, so nothing else in your vendor-management process will catch it.

Open-weight inference does not make a platform better at finding bugs, and nobody should claim it does. What it changes is who can answer the question "what will this thing be able to do for us next quarter". On a frontier API that answer belongs to an access committee in another jurisdiction. On weights the vendor holds, it belongs to the vendor, and it can be written into a contract you sign.

September 2026: the Cyber Resilience Act starts asking for reports on 11 September, on a 24-hour clock

Every regulation on this page so far applies to you because of who you are: a financial entity in DORA scope, an essential or important entity under NIS2. The Cyber Resilience Act applies because of what you sell, which for a good share of the people reading a page like this one is the wider net.

From 11 September 2026, manufacturers of products with digital elements have to report actively exploited vulnerabilities and severe incidents affecting the security of those products. The Commission is blunt about the date and the staging: an early warning within 24 hours of becoming aware, a full notification within 72 hours, then a final report no later than 14 days after a corrective measure is available for a vulnerability, or within a month for a severe incident. Article 14 of the regulation carries the same three clocks. Reporting happens once, through the Single Reporting Platform ENISA is standing up, mandatory from 11 September 2026 and voluntary before that date.

Check the scope before deciding this is someone else's paragraph. It reaches manufacturers, importers and distributors of connected hardware and software placed on the EU market, plus providers of remote data processing solutions supporting connected products, and non-EU manufacturers selling into the EU are equally in scope, with fines for breaches of the essential requirements and the reporting duties reaching EUR 15 million or 2.5% of worldwide annual turnover. If you ship software into the EU from outside it, this is your regulation, not only your customers'.

Two things follow for a pentest shortlist, and neither appears on a comparison page.

A 24-hour clock that starts at "becoming aware" turns finding quality into a timing problem. One day is not enough to work out whether an agent's confident-sounding output is a real, exploitable vulnerability or a plausible sentence, and the clock does not pause while your engineers find out. So the validity question from the evidence section above, the share of delivered findings that carried a reproducible proof of concept, stops being a procurement nicety and becomes the input to a legal deadline. Ask each vendor how a finding arrives: with a working proof of concept and evidence of exploitation attached, or with a severity score and a recommendation to investigate. Those two deliverables behave very differently at hour 23.

Your findings pipeline is now pre-disclosure vulnerability data on a regulatory timetable. The CRA already asks manufacturers to "apply effective and regular tests and reviews of the security of the product with digital elements", alongside a software bill of materials, remediation without delay and a policy on coordinated vulnerability disclosure. It names no methodology, so a pentest is one way to satisfy that point rather than the required way. What matters for vendor choice is where the evidence and the unfixed findings sit while the clock runs. A platform holding your unpatched, actively exploited vulnerability through the days before a fix ships is holding the most sensitive file your company owns that week, and "which legal entity, in which country" arrives as the same question the DORA sections asked, from a second direction.

The dull step, and the one worth doing before 11 September: write down who inside your company decides that a finding counts as an actively exploited vulnerability, and which evidence that decision reads. The practical checklist around it is product inventory, detection capability, contractual notification from your component vendors, and one incident-response plan carrying the CRA, GDPR and NIS2 timelines together. If a pentest report is one of the evidence sources feeding that decision, the report format and the vendor's own notification path belong in the plan rather than in the procurement file.

September 2026: the one test on your DORA plan that no agent can sign

Everything above compares platforms. If you are shopping XBOW alternatives from inside DORA scope, it is worth separating out the one piece of your testing plan that no platform on this page can sell you, because vendors rarely draw the line and buyers keep discovering it late.

DORA splits testing in two. Article 24 and Article 25 are the ongoing testing programme every in-scope entity runs, and that is the budget line agentic pentest competes for. Article 26 is different: entities identified by their competent authority have to carry out "at least every 3 years advanced testing by means of TLPT", threat-led penetration testing, "performed on live production systems" supporting critical or important functions, and the authority issues an attestation at the end confirming the test was performed as required. The detail lives in Commission Delegated Regulation (EU) 2025/1190 of 13 February 2025, published in the Official Journal on 18 June 2025 and applicable from 8 July 2025. It reaches the larger and systemically important entities rather than everyone: systemically important credit institutions, payment institutions above EUR 150 billion in annual transaction volume, e-money institutions above EUR 40 billion outstanding, central counterparties, central securities depositories, significant trading venues and insurers, plus some crypto-asset service providers.

Now read who is allowed to run it. Article 27 asks for testers "of the highest suitability and reputability" who are "certified by an accreditation body in a Member State or adhere to formal codes of conduct or ethical frameworks", who provide independent assurance or an audit report on their own risk management, and who are "duly and fully covered by relevant professional indemnity insurances, including against risks of misconduct and negligence". Use your own staff instead and you need your competent authority's approval, plus an external threat intelligence provider either way.

The RTS turns that into a headcount. Article 7 of the delegated regulation asks an external provider for a lead at manager level with at least five years in penetration testing and red team testing, at least two additional testers with at least two years each, and at least five references from previous assignments, with the threat intelligence side needing a manager with five years and one more member with two, plus three references. Staff on the test cannot be doing blue team work for the same entity.

Three practical consequences for a shortlist.

An agent cannot be the tester, only a tool the tester uses. Accreditation, references and indemnity insurance attach to people and legal entities, not to software. So a platform that shortens your Article 24 programme from weeks to hours does nothing for the TLPT clock, and a vendor who says it covers "DORA testing end to end" should be asked which article. Two budget lines, not one: a red team engagement every three years, and the continuous testing where the vendors on this page actually compete.

Your pentest vendor may be dragged into a TLPT as a participant. Article 26 makes the entity "ensure the participation of such ICT third-party service providers in the TLPT" where they support the function in scope, with pooled testing allowed when going one by one would put service quality at risk. That is a clause to negotiate at signature. It is also where a provider whose engineers sit eight time zones away costs you calendar weeks that a European one does not.

The TIBER-EU procurement guidance is free homework you can reuse. The ECB updated the TIBER-EU framework on 11 February 2025 to align it with DORA and the TLPT RTS, making purple teaming mandatory, renaming the White Team to the Control Team, and expanding the guidance on how to assess the quality of a provider. On 21 November 2025 it issued an SSM implementation guide for significant institutions, down to appointing a single point of contact per test to keep it secret. The provider-quality questions in that set were written for red team procurement, and most of them work unchanged on an autonomous pentest vendor.

The dull step: on one page, write which article each line of your testing spend satisfies. Most of the confusion in this market comes from a vendor answering an Article 24 question with an Article 26 word, or the other way round, and the page ends that conversation in a sentence.

September 2026: the EU is grading sovereignty on four levels, and DORA puts a floor under the shortlist

Since April this article has answered "how European does an alternative have to be" with judgement. Two pieces of law are now answering it with text, and both are worth reading before your next vendor review.

On 3 June 2026 the Commission proposed the Cloud and AI Development Act. Article 16 would set a single EU-wide scale for how much sovereignty a public buyer has to demand, in four assurance levels. Level 1 asks for infrastructure, assets and customer data in the EU, plus a guarantee that a provider controlled outside the Union cannot be compelled to report vulnerabilities to a foreign authority. Level 2 adds EU-located personnel, certification under the EU cloud scheme, measures preventing third-country access, and source-code audits of foreign components. Level 3 requires the provider itself be owned and controlled in the EU, with derogations for recognised third countries. Level 4 removes the derogations and adds European cybersecurity certification at "high" assurance.

Read level 1 again with an offensive security vendor in mind. The thing your pentest provider holds is a list of your unfixed vulnerabilities, so "cannot be compelled to report vulnerabilities to a foreign authority" is not an abstract clause for this category, it is the clause.

Two honest limits. CADA is a proposal, not law, so do not tell procurement it binds them yet. And it is aimed at public authorities, with the possibility of reaching NIS2 essential entities through later secondary legislation. The separate "Union added value" procurement criterion is deliberately small, capped at around 15 of 120 evaluation points and described as ancillary rather than decisive. Anyone selling you "the EU now mandates European vendors" is overselling it. What CADA gives you today is vocabulary: four named levels you can paste into a questionnaire instead of arguing about the word sovereign.

The floor is already binding, and it sits in DORA. Article 31(12) says a financial entity may only keep using a third-country provider that has been designated critical if that provider has established a subsidiary in the Union within 12 months of the designation. The 19 firms designated in November 2025 are already EU-incorporated, so nobody has been cut off. The point is forward-looking: if the offensive testing you buy arrives through a US-headquartered provider that later gets designated, the continuity of that contract depends on a corporate decision you do not control, on a 12-month clock you do not set. That is a question for the exit-strategy section of your register, not a reason to panic.

Why CISOs search "XBOW alternative europe"

The query is high-intent. A buyer typing it has already concluded that XBOW is the technical reference and is now hunting for the EU-eligible equivalent. Three drivers stack on top of each other:

  1. DORA structural disqualification. DORA Article 28 + Article 30 supervisory expectations make US-headquartered ICT third parties hard to onboard for designated financial entities. Adding XBOW to a DORA-scope entity's third-party register triggers concentration-risk and third-party-risk reporting obligations that procurement does not want to carry.
  2. EU AI Act obligations, now on a later clock. High-risk deployments need documented data governance, human oversight, technical documentation per Articles 11-12, and a US frontier-model provider in the data path puts part of that documentation outside the buyer's control. The deadline is no longer 2 August 2026: Annex III systems have until 2 December 2027. Treat this as a contract clause to negotiate now rather than a gate that blocks a purchase this year.
  3. Schrems II and the unstable transfer regime. Privacy Shield was invalidated in 2020. The 2023 Data Privacy Framework partially restored a transfer mechanism, but Schrems III is widely expected. CISOs who got burnt once do not want to bet on it again.

The same query generates non-EU answers (Hadrian, Terra Security, Penligent), which an EU buyer must filter out a second time. This article does the filter.

XBOW: the reference, in one paragraph

XBOW runs a multi-agent autonomous offensive security platform. Multiple specialised agents (recon, planning, execution, validation) coordinate to map an attack surface, plan exploitation chains, and validate findings end-to-end. The architecture is the closest commercial match to what a senior red-team-led junior team produces, at machine speed. Detection rate, validated proofs of concept and time-to-finding are what made it the reference. The public benchmark it released to demonstrate that is now saturated, on XBOW's own assessment, which is covered above. It is still the right reference for what agentic pentest can do.

But XBOW operates on US-headquartered infrastructure with US LLM dependencies in the data path. For an EU regulated buyer in 2026, that is not a procurement-friendly pile.

Most "XBOW alternatives" lists compare company names. That produces the wrong shortlist, because XBOW does not cover everything a pentest programme needs. Its own product page is explicit about the input: you point it at a URL. The scope is web applications. Standalone API testing and mobile testing are roadmap items rather than shipped surfaces, and internal network, cloud and Active Directory are outside the product entirely.

So the useful question is not "who else does what XBOW does". It is "which part of my surface am I actually replacing".

  • Web application agentic pentest, the like-for-like set: Fleuret AI, Escape, Ethiack. Terra Security, RunSybil and MindFort do the same job from the US.
  • REST and GraphQL API depth: Escape and Fleuret AI. This is where a web-app-scoped agent is thinnest today.
  • External attack surface and exposure: Patrowl. Not a pentest, and often the gap that was actually causing the pain.
  • Internal network and Active Directory: Pentera and Horizon3. XBOW does not compete here.
  • Continuous testing with human validation on top: Ethiack, which pairs its Hackbot with a vetted hacker network.
  • One suite instead of best-of-breed: Aikido, when the buying logic is vendor consolidation.

A buyer who shortlists an internal-network validation tool as an "XBOW alternative" has bought the wrong category and will find out at the first readout, usually in front of a board.

How to compare XBOW alternatives on evidence, not on claims

Almost every benchmark in this category is published by a vendor about itself. The one independent 2026 reference point comes from a Stanford-led study that ran ten cybersecurity professionals and six existing AI agents against a live university network of roughly 8,000 hosts across 12 subnets. The study's own multi-agent framework, ARTEMIS, placed second overall with nine valid vulnerabilities and an 82% valid submission rate, ahead of nine of the ten human participants, at around $18 per hour against $60 per hour for the professionals.

Read the second number, not the first. An 82% valid submission rate means roughly one submission in five was not a real finding.

Speed you can verify in a demo. Validity you only discover after your engineers have spent a sprint on findings that were never exploitable.

That gives you a question that separates vendors quickly: what share of your delivered findings carried a working proof of concept that our team reproduced, on your last ten engagements. A vendor who validates before delivery can answer with a number. A vendor who ships raw agent output will answer with a detection rate instead.

The demand side is not much more disciplined. Accenture's investment note cites the World Economic Forum Global Cybersecurity Outlook 2026 finding that around two thirds of organisations expect AI to affect cybersecurity while only 37% have any process to assess AI tools before deploying them. If you are buying an autonomous agent to test your estate, you are also deploying an AI tool into it. Both reviews are yours.

Turnaround time became a comparable number

For most of 2025 you could not compare delivery speed across this category, because no vendor published one. That changed when XBOW launched Pentest On-Demand, a self-service pentest that returns complete results within five business days. Whatever you make of the depth behind it, a stated turnaround is a commitment a buyer can hold a vendor to, and it gives you a row to put in front of every other vendor on your list.

Use it carefully, because turnaround gets measured from three different starting lines and vendors pick the flattering one.

  • Scope agreed to report delivered. The only number that matches your calendar and your auditor's.
  • Scan started to results available. The number most vendors quote. It hides scoping, credentials and environment setup, which is usually where the calendar days actually go.
  • Report delivered to retest of the fixes. The number nobody quotes, and the one that decides whether your remediation window closes before the audit does.

Ask for all three in writing, with the definition attached to each. A traditional consulting pentest runs two to four weeks of testing plus a reporting lag. An autonomous vendor that cannot beat that on the first measure and match it on the third was selling speed rather than delivering it.

Side-by-side: the European alternatives

VendorHQScopeDORA-eligibleOpen-weight LLMAudit-PDF formatPrice bandICP fit
Fleuret AIFranceWebapp, REST/GraphQL API, external infraYesYes (gpt-oss-120b/20b, Kimi K2.5 on Scaleway France)DORA Article 24 + NIS2 mappings, Ed25519 signedPentest €4k / Advanced €8k / Continuous on quoteMid-market 300-5000, NIS2 / DORA-scope, sovereignty-buyer
EscapeFranceAPI, web app (DAST-derived)YesPartial disclosureCustom, framework-tailored snippetsPlatform pricing, not publicScaling SaaS, API-first, dev-team-owned security
Aikido SecurityBelgiumCode, cloud, runtime, AI pentestYesDisclosed in security pagesSOC 2, ISO 27001 mappingsMid-market SaaS-friendly50,000+ orgs, broad mid-market, code-cloud-runtime unified
PatrowlFranceContinuous exposure, attack surfaceYesDisclosedExposure management dashboardsEnterprise pricingMid-large enterprise, exposure-management buyer
EthiackPortugalWeb app, external surface, continuousYesNot disclosedContinuous findings feed, human-validatedEUR 4M seed, pricing on quoteMid-market wanting an agent plus a hacker network
PenteraIsraelAutomated security validation, internal + externalPartial (non-EU HQ)NoValidation reports~€46k/yrLarge-enterprise, mature security teams

A few reading notes on the table.

DORA-eligible is binary in column heading but actually a sliding scale. "Yes" means EU-headquartered legal entity with documentable EU operational chain. "Partial" means non-EU HQ but with EU-region offerings that may pass with supplementary measures. Buyers should still run the 7-question sovereignty checklist before signing.

Headquarters is the first filter, not the last. Every vendor in this table is EU or EU-adjacent, which is the whole point of the table. The names that arrived on comparison pages during 2026 (RunSybil, MindFort, Terra Security) are not, and land on your register the same way XBOW does.

Open-weight LLM matters for AI Act high-risk deployments and for documentable model governance. A "yes" means the vendor can produce model identifier, version, and training-data lineage on demand.

Audit-PDF format is the workflow-lock-in surface that distinguishes vendors who ship pentest as a deliverable from vendors who ship pentest as a workflow. See the compliance-workflow article for why this matters.

Architecture differences that matter at scale

Three architectural choices actually drive different outcomes.

Multi-agent vs single-agent. Public benchmark research shows multi-agent hierarchical architectures outperform single-agent approaches by 4.3× on validated-finding rate. Fleuret, XBOW, and Terra Security run multi-agent. Escape's engine is closer to an agentic-DAST hybrid (effective for API and web app, narrower for full grey-box). Aikido's pentest module is one component of a broader compliance suite.

Open-weight vs frontier-API LLM. Open-weight models (gpt-oss-120b, Kimi K2.5, Mistral) give the vendor full operational control of inference. Frontier-API LLMs (OpenAI, Anthropic, Google) constrain the vendor to a third-party data path. For pentest, where prompts include payloads, response bodies, and authentication artifacts, that constraint translates directly to sovereignty risk.

Coverage Graph vs flat scope tracking. A pentest that does not track exhaustiveness is just a vulnerability scanner with extra steps. Fleuret's Coverage Graph (hierarchical data structure tracking what was discovered, what was tested, what remains unexplored) is one approach. Patrowl's exposure-management dashboard is another. XBOW publishes equivalent coverage abstractions. The exact data structure matters less than whether the vendor can answer "what did you not test, and why".

Buying guide: which European tool fits which buyer

Regulated finance / DORA-scope mid-market. Sovereignty plus DORA Article 24 mapping plus auditor-ready signed reports are non-negotiable. Fleuret AI and Patrowl are the strongest matches. Escape works for the API perimeter alongside one of those.

Mid-market SaaS, dev-team-owned security. API and web app coverage with developer-friendly remediation matters most. Escape leads here. Aikido fits if a unified code-cloud-runtime suite is the buying logic.

Large enterprise, mature security teams. Pentera's cost is justifiable, the validation reports plug into existing red-team programmes. Patrowl scales for exposure management. XBOW is technically valid but adds the DORA / sovereignty drag.

Healthtech, fintech, retail mid-market with NIS2 obligations. NIS2 Annex I mappings, ANSSI ReCyF traceability, Ed25519-signed reports. Fleuret AI is purpose-built for this profile.

Agency or partner-led delivery. Aikido has the broadest channel programme. Fleuret AI sells direct and is building partner relationships with GRC platforms (separate motion, not the focus of this comparison).

What each vendor does best

Fleuret AI. Sovereign-by-default agentic pentest with the compliance-workflow surfaces (Jira, audit PDF, board export, weekly cadence) wired in by default. Best fit for EU mid-market 300-5000 emp inside DORA / NIS2 scope.

Escape. Strongest agentic engine for BOLA, IDOR, and access control on API surfaces. Developer-friendly remediation, framework-tailored code snippets, regression testing on continuous deployment.

Aikido Security. The "everything platform" play. 50,000+ organisations use the broader suite (SAST, DAST, IaC, container, secrets, plus AI pentest). Right answer when the buying logic is "consolidate vendors", not "pick the best pentest".

Patrowl. Continuous exposure management with attack-surface-monitoring depth, named in Gartner Market Guide for Preemptive Exposure Management 2026. Right answer when the primary problem is "we do not know what we are exposing".

Pentera. Mature automated security validation, strong on internal-network testing. Cost makes it a large-enterprise tool, not a mid-market answer.

What to verify before signing any of them

Run two checklists, both before the product demo:

  1. The 7-question sovereignty checklist for legal-review pre-clearance.
  2. The 7-question workflow-lock-in checklist for operational fit.

Vendors that fail more than two questions on either checklist are not serious candidates for an EU regulated mid-market buyer in 2026.

What DORA reporting does to the decision, in practice

For a financial entity in DORA scope, picking a pentest vendor is not only a security decision. Every contractual arrangement for ICT services lands in the register of information, filed for the 2026 cycle between 11 February and 31 March, with validation strict enough that submissions accepted in the first year were rejected in the second. Competent authorities pass those registers to the ESAs, who use them for the annual designation of critical ICT third-party service providers.

Three consequences worth carrying into the vendor call.

You will describe this vendor to a supervisor every year. Not once at signature. The register is an annual filing, and the entries have to survive validation each cycle.

A vendor reached through a consultancy is still your entry. If XBOW arrives inside an integrator's delivery, someone still has to work out whose contractual arrangement it is and how the sub-chain is described. Decide that before the statement of work, not during the filing window in February.

Data location has to be documentable, not asserted. "EU region available" is a sales answer. The register wants entity, country and the chain beneath it. A vendor who cannot name its sub-processors on a public page will cost you time every March.

If you are outside DORA scope, this section is the reason your regulated customers will ask about your pentest vendor during their own supplier reviews. NIS2 supply-chain duties push the same question one layer down the chain.

Under DORA, your shortlist is a document, not a preference

Most people reading a page like this one are shopping. If you are a financial entity in DORA scope, you are also doing homework you owe your supervisor, and it is worth knowing that before you decide the exercise is optional.

Article 28(8) requires that for ICT services supporting critical or important functions, financial entities put exit strategies in place, taking account of risks that may emerge at the provider. Those strategies have to let the entity exit the contractual arrangement without, in the regulation's own three tests, "disruption to their business activities", "limiting compliance with regulatory requirements" or detriment to the continuity and quality of service provided to clients. The plan has to be documented, has to identify alternative solutions, and has to be tested and reviewed periodically.

Read "identify alternative solutions" as the operative phrase. If an autonomous pentest platform is part of how you evidence resilience testing on a critical function, then naming a replacement you could actually move to is a filing obligation, not a sign that you are unhappy with the incumbent. A shortlist you assembled once during procurement and never touched again does not meet a periodic review.

Article 28(7) sits underneath it and is the clause people skip. You have to be able to terminate at all, on grounds that include a "significant breach by the ICT third-party service provider of applicable laws, regulations or contractual terms", circumstances found during monitoring that could alter how the function is performed, evidenced weaknesses in the provider's own ICT risk management (specifically around availability, authenticity, integrity and confidentiality of data), and the case where your competent authority can no longer effectively supervise you because of the arrangement. A contract that does not carry those exits is a finding waiting to be written, whoever the vendor is.

The practical effect on this page: treat the comparison above as the raw material for an exit plan, not as a buying guide you read once. Two vendors, named, with the scope each would cover and a rough sense of what it would take to move, is most of what the document needs.

What switching actually costs, and the four things to ask before you sign anyone

Because DORA asks whether you could move, it is worth knowing what moving costs. This is the part no comparison page covers, and it is where a shortlist stops being theoretical. Four questions, all of which should be answered in writing before signature rather than discovered at renewal.

Does your findings history leave with you. Two years of tested-and-fixed history is how you show an auditor that remediation happened, and how a new vendor knows what not to re-report. Ask for the export format, whether closed findings and proofs of concept come with it, and whether it survives after the contract ends. A platform that keeps the history is a platform you cannot leave cheaply, regardless of what the exit clause says.

Does the retest baseline transfer. If your evidence of resilience testing is "the incumbent retested and confirmed the fix", a new vendor starts from zero and your next report reads like a first engagement. Ask both vendors whether they will accept a third party's prior findings as a baseline. Some will, and it removes a cycle of noise from the transition.

What happens to committed cloud spend. This is the clause specific to 2026 and to the marketplace routes described above. Budget drawn down against an AWS or Microsoft commitment was allocated to that vendor's listing. It does not re-point at a European vendor who is not on the same marketplace, and it does not usually come back as cash. That is not a reason to avoid marketplace purchases, but it does mean the switching cost is real money rather than a migration weekend, and it belongs in the exit plan as a number.

Will your auditor accept the new report format. The deliverable your auditor signed off last cycle set an expectation. Changing vendors changes the document, the mapping to NIS2 Annex I or DORA Article 24 duties, and often the signature and integrity guarantees on the file itself. Ask to see a redacted sample report from every vendor on the shortlist and walk it past whoever owns the audit relationship, before the demo rather than after.

A note on the badges that come up in this conversation. XBOW reached AWS Security Competency status on 7 July 2026, a designation AWS grants for meeting its technical and quality requirements in application security. It is genuine engineering validation and it is often quoted in vendor reviews as though it settles more than it does. It does not name a contracting entity, a hosting region, or the share of delivered findings that carried a working proof of concept. Those three answers still have to come from the vendor.

Your next steps

  1. Name the surface first. Write down whether you are buying web app, API, external exposure or internal network testing before you look at a single vendor page. The shortlist falls out of that sentence.
  2. Ask for the validity number, not the detection rate. Share of delivered findings with a reproducible proof of concept, over the last ten engagements. Anyone who validates before delivery can answer.
  3. Drop the AI Act argument and re-arm with DORA and NIS2. The high-risk deadline moved to 2 December 2027. Procurement will notice if you cite the old date.
  4. Check how the vendor will appear in your register. Legal entity, country, sub-processors, and whether an integrator sits between you and the tool.
  5. Check the headquarters line before the feature table. Half of the 2026 alternatives list is US-headquartered. If eligibility is your reason for shortlisting an alternative at all, filter on that first and you will save a week of demos.
  6. Run both checklists before the demo. Sovereignty first, then workflow fit. A demo booked before the legal pre-clearance wastes the slot most of the time.
  7. Reconcile your marketplace subscriptions against your register. Anything bought against committed AWS or Microsoft spend never met your vendor questionnaire. Start with the tools that hold production credentials.
  8. Get turnaround defined in writing, on all three clocks. Scope to report, scan to results, report to retest. A vendor who will only quote the middle one is quoting the easy one.
  9. Write the exit plan while you are still choosing. Article 28(8) wants named alternatives, documented, tested and reviewed periodically. The cheapest time to produce that document is now, while you have three vendors answering questions and nothing has been signed.
  10. Price the exit before you price the contract. Findings-history export, retest baseline, committed cloud spend that will not transfer, and whether your auditor accepts a different report format. A shortlist that ignores those four is a shortlist you cannot act on.
  11. Decide whether you are buying a tool or a deliverable. If the output only has to satisfy your own team, a self-hosted agent under Apache 2.0 or Creative Commons Zero is a real option and costs you inference rather than licence fees. If the output has to satisfy an auditor, an insurer or a board, you are buying a signed report and someone to stand behind it, and no free agent supplies either.
  12. Check whether your delivery route is itself a designated critical provider. Nineteen firms were designated under DORA Article 31(9) in November 2025, Accenture, AWS EMEA and Microsoft Ireland among them. Buying an offensive tool through one of them does not inherit their supervision, and it makes the concentration-risk line on your register harder to defend.
  13. Name the person who decides a finding is actively exploited. From 11 September 2026 that call starts a 24-hour reporting clock for manufacturers of products with digital elements. Write down who makes it, what evidence they read, and where a pentest report sits in that chain.
  14. Ask whether findings arrive validated or scored. A proof of concept you can reproduce answers the 24-hour question. A severity score with a recommendation to investigate spends the window instead of closing it.
  15. Label each testing spend with the article it satisfies. Article 24 and 25 for the ongoing programme, Article 26 for the three-yearly TLPT. No platform on this page is an eligible TLPT tester, because accreditation, references and indemnity insurance attach to people, so budgeting one against the other leaves a gap your supervisor will find.
  16. Put TLPT cooperation in the contract before signature. If the vendor supports a critical or important function, you have to ensure its participation in the test. Agree it in writing, and ask which team and which country would answer.
  17. Stop scoring vendors on XBEN, and ask for the artefacts instead. XBOW's own benchmark repository states that industry-standard performance is now around 100% and that the set no longer separates attack frameworks. Ask for the run transcripts, the harness so your team can re-run the suite unattended, and the time and cost per target.
  18. Treat any coverage multiplier without a baseline as positioning. "Ten times more coverage" needs a stated denominator and a measurement method attached to it. A coverage map naming what was tested and what was cleared is the version your auditor can read.
  19. Ask which model the platform calls, and under which access programme. OpenAI now gates advanced cyber capability behind Daybreak tiers, with vulnerability research and exploit development behind a separate approval. A frontier-API vendor's ceiling and its continuity both depend on a decision you are not party to.
  20. Add "loses its model access tier" to the failure list in your exit plan. It degrades the service you bought without breaching any clause you could cite, so no other part of your vendor-management process will raise it.
  21. Ask every shortlisted vendor which CADA assurance level it could evidence today. The four levels in Article 16 of the proposed Cloud and AI Development Act run from EU-located data to EU ownership with no derogations. It is not binding on private buyers, and it is still the cleanest vocabulary available for a question that usually turns into an argument about the word sovereign.
  22. Add third-country designation to your exit triggers. Under DORA Article 31(12), a non-EU provider designated critical has 12 months to establish a subsidiary in the Union before financial entities have to stop using it. Whether your provider does that is a corporate decision, so the trigger belongs in your plan rather than in your assumptions.

Where Fleuret AI stands

Sovereign FR-headquartered, open-weight LLMs on Scaleway France, Vercel fra1, public sub-processors page, Ed25519-signed audit PDFs, Jira-integrated workflow, weekly cadence on the Continuous tier. Pricing visible on /pricing.

I am building Fleuret around the part of this that a supervisor actually reads: validated findings, EU data path, evidence an auditor accepts. See Fleuret in action.

Sources


Share this postShare on LinkedIn

Privacy Settings

This site uses third-party website tracking technologies to provide and continually improve our services, and to display information according to users' interests. I agree and may revoke or change my consent at any time with effect for the future.