Skip to main content

Continuous API security assessment: a practical guide

Yanis Grigy, CEO11 min read

A useful continuous API security assessment is not a pile of scans. It links a live inventory, pre-release tests and runtime signals to reduce real risk without breaking the delivery pace. If you cannot see your public, partner, internal, shadow and zombie APIs, you are mostly assessing what you already know.

Key takeaways

  • A continuous assessment starts with a multi-source inventory. Otherwise you are checking a scope that is wrong or already out of date.
  • Useful tests cover the specification, the code, the dynamic behaviour and production, at a frequency matched to how often things change.
  • Authorisation abuse, especially BOLA and BFLA, has to be tested with concrete scenarios, not only with generic rules.
  • Prioritisation has to combine exploitability, exposure, data sensitivity and compensating controls, or the noise wins.

A continuous assessment starts with the real inventory

Documented, undocumented, deprecated: count everything

You cannot continuously assess what you do not know you expose. A public API, a partner endpoint, a forgotten internal API, a shadow API created by a product team or a zombie API left running after a migration all belong in the same inventory. Otherwise you raise alerts on objects that are already visible and leave the riskiest part out of scope.

Cross three sources: OpenAPI specifications, API gateways and code repositories. OpenAPI gives the intended view, but it often lags. The gateway sees what actually flows, but it can miss business context and routes that are not wired yet. The code reveals the real endpoints, but not their exposure or usage. The gap is measurable: in its 2024 API security report, Cloudflare found 30.7% more API endpoints through machine-learning discovery on traffic than through what customers had declared, suggesting that nearly a third of APIs are shadow APIs. Inventories are rarely wrong in a spectacular way. They are wrong through small omissions that end up mattering.

The inventory must include data, auth and exposure

The minimum record fits on one page: owner, environment, schema, authentication type, sensitive data, Internet exposure, critical dependencies. Add one field many teams forget: the consumption mode. An API used by a B2B partner, a mobile client and an internal batch job does not carry the same risk, even if the route is identical.

A purely declarative inventory goes stale the moment a service evolves outside the process. A purely network-based inventory misses the why, and therefore the risk level. The right compromise is to accept imperfect coverage, as long as it stays tied to what is actually deployed and to its business use.

The right control scope covers the whole API lifecycle

Before merge and before deployment

The minimum baseline has four layers: specification validation, SAST and secret scanning, dynamic tests of API behaviour, and runtime monitoring. The OWASP API Security Top 10 is a framework to structure those controls, not a substitute for context. Use it so nothing gets forgotten, not to conclude an API is safe because it passes a checklist.

In a realistic CI/CD chain, a pull request triggers OpenAPI validation, a check of authentication and rate-limiting requirements, then tests on an ephemeral environment. If the schema changes, a sensitive route appears or secrets are detected in the repository, the merge is blocked. Static checks should run on every commit: they are cheap and catch simple regressions early.

In pre-production and production

Dynamic tests should run on every release, not on every commit, or the noise becomes unmanageable. In production you switch to guardrails: error monitoring, detection of behavioural drift, correlation with security events. The frequency follows the rate of change, not a wish to scan everything all the time. Tools such as OWASP ZAP, Burp Suite, Postman, 42Crunch or Kong each have their place on their own layer.

Keep an eye on real volume: APIs are no longer a marginal slice of the web. Cloudflare reports that "well over half of the dynamic traffic" on its network is API traffic rather than web pages. When API routes carry most of the useful traffic, runtime signals are not a bonus. They are the only way to see the gaps that pre-production does not reproduce.

Useful tests check abuse, not just 200 OK

Test object-level and function-level authorisation

The two risk families that come back most often are BOLA (Broken Object Level Authorization) and BFLA (Broken Function Level Authorization). BOLA is access to an object that does not belong to you. BFLA is calling a function your role should not reach. Generic rules sometimes flag the pattern, but only test scenarios show the real abuse.

A concrete example: user A is properly authenticated, changes the resource identifier in the request, and gets user B's data back. If the response comes back, the problem is not authentication, it is object-level authorisation. The same logic applies to an admin function exposed to a standard role. That class of flaw does not show up in a schema scan.

This is where an offensive test earns its place in the program. Fleuret runs agentic pentests on REST and GraphQL APIs that chain these authorisation scenarios across real identities, with results in hours and a replayable proof of concept for every finding. More on where automation and human testers each win in our comparison of automated and manual pentesting.

Test schemas, limits and business logic

Effective tests include at least parameter fuzzing, response schema validation, and tests of rate limiting and abusive pagination. On a search API, unbounded pagination can turn into data leakage at scale. On a financial API, a sign flip on an amount or a mistyped currency field can be enough to create an exploitable regression. Cover a few critical flows in depth rather than skimming every route. Shallow coverage reassures, but it often misses the bypass paths.

Production reveals the gaps scans cannot see

Watch traffic, errors and drift

Runtime shows what the specification does not say. Undocumented endpoints appear in traffic. Abnormal volumes point to scraping, batch abuse or a badly built client. New consumers, unusual 401, 403 or 429 responses and schema drift often show that reality has moved away from the contract.

The useful signals are concrete: a spike of 5xx errors on a sensitive route, a sudden rise in authorisation denials, a response that no longer matches the expected schema, or an over-privileged token appearing. Gateway, WAF, eBPF or service mesh data enriches that picture without replacing upstream tests. It is also the best way to tell a false positive from the start of real abuse.

These signals should trigger a targeted reassessment. A production incident must not stay an isolated ticket in the support tool. Link it to the logs, the API owner, the expected contract, the data sensitivity and the possible attack path. When three services report the same series of 403s after a role change, the answer is not more alerts: it is an authorisation drift that deserves an immediate regression test.

Prioritisation must follow exploitability and impact

What deserves an immediate fix

A simple four-criteria matrix is often enough: exploitability, Internet reachability, data sensitivity, and whether a compensating control exists. A BOLA on customer data exposed to the Internet is critical. An internal endpoint without auth, but segmented and barely exposed, is high, not necessarily critical. An inconsistent piece of documentation with no runtime impact is rather medium.

Raw vulnerability scores are not enough for APIs, especially for business logic flaws. Two findings with the same score can have very different impact depending on the data volume, the role involved and how easy the abuse is. Continuous assessment does not rank abstract CVEs, it ranks real abuse paths. In practice, always put first a flaw that lets someone extract data or act as another user, even if its technical score looks moderate.

What can be handled in batches

Not every gap needs immediate treatment. Schema anomalies on non-sensitive fields, forgotten test endpoints behind strict internal access, or documentation gaps with no runtime effect can be grouped into weekly batches. That matters for teams running twenty to fifty APIs without ten AppSec engineers available. The point is not to downplay the problem. It is to keep triage from saturating on items that change neither the risk nor the decision.

The common mistakes that sabotage a continuous program

Too many tools, not enough decisions

The same traps come up again and again: scanning only the specs, ignoring internal APIs, testing without a realistic set of identities, drowning the team in untriaged findings, and never checking fixes in production. Another classic is confusing scan coverage with risk reduction. They are not the same thing.

The organisational risk is simple: without an API owner and a remediation SLA, "continuous" turns into reporting. A team that moves from a quarterly scan to lightweight checks on every change often gets better risk reduction, because it fixes earlier and with less effort (the case against the annual cycle is laid out in our piece on continuous testing for SaaS). The right trade-off is neither "block all the time" nor "never block". Plan time-boxed exceptions, with an end date and a named owner, so delivery does not break.

Security kept apart from API teams

When security stays outside the API teams' workflow, findings pile up untreated. Organisations with excellent dashboards can still ship very few fixes, simply because the service owner never sees the context needed to act. The program works when the service owner sees the control as part of the delivery cycle, not as an external request. That takes short reviews, clear exit criteria and a shared vocabulary between security, platform and development. Pushing findings straight into the tools those teams already use (Fleuret sends them to Jira, Linear or Slack, then retests after the fix) removes one of the usual excuses.

A 90-day implementation plan is enough

Days 1 to 30

Start small. Pick 10 to 20 critical APIs, build the inventory, define the minimum control per API type, then choose 3 to 5 priority abuse scenarios. The deliverables of this phase are an API register, a control baseline and a first prioritisation grid. If you cannot name an owner per API, the rest will not hold. Record the data handled, the exposure level and the authentication type from the start, or you will reclassify the same services three times.

Days 31 to 60

Wire the tests into CI/CD and an ephemeral environment. Produce a prioritisation board with findings ranked by impact and exploitability. Add a remediation runbook for recurring cases: BOLA, BFLA, authentication errors, data exposure, missing rate limiting. By now you should see whether your controls reduce the noise or just move it. If alert counts climb but time to fix does not move, the program observes better, but it does not act better.

Days 61 to 90

Connect production: traffic metrics, errors, over-privileged tokens, undocumented routes. Track four simple KPIs: inventory coverage, share of APIs tested on every release, mean time to fix critical findings, and number of undocumented APIs detected. The goal is not a one-off audit, it is a learning loop that keeps running. From there, extend to partner APIs and older services, which are often the hardest to bring into line.

The target is not the perfect scan, it is the fast decision

Aim for a program that sees the real APIs, tests the most dangerous abuse and surfaces priorities people can act on. If you have to pick a first step, connect the inventory to pipeline controls, then add production observation on the 10 to 20 APIs that carry the most data or revenue. That is where risk reduction becomes visible.

A useful continuous assessment does not chase full coverage in the first month. It builds a discipline that learns, decides and fixes without needlessly slowing delivery.

Fleuret runs agentic pentests on web apps, REST and GraphQL APIs and external infrastructure, with replayable PoCs, findings in Jira, Linear or Slack, and a retest after the fix. See Fleuret in action.

Sources


Share this postShare on LinkedIn
TRY IT ON YOUR APP

Scan your own code, free.

The Free plan runs SAST, DAST, SCA and secret scanning for 2 users and 10 repos, no call needed. Several apps or an audit deadline? Book a demo instead.

Privacy Settings

This site uses third-party website tracking technologies to provide and continually improve our services, and to display information according to users' interests. I agree and may revoke or change my consent at any time with effect for the future.