Why AI Coding Agents Keep Writing Broken Access Control
Why AI Coding Agents Keep Writing Broken Access Control
October 1, 2026
0 mins readAI coding agents produce authorization logic that compiles, passes review, and enforces the wrong policy. Broken access control ranks first in the OWASP Top 10:2025, where 100% of applications tested showed some form of it, across 1,839,701 recorded occurrences, the highest count of any category on the list.
One part of that category is also the part that pattern-based scanning was never built to reach. When an agent omits an ownership check, the rule it broke belongs to the application rather than to a signature database. That leaves no known-bad pattern to match against.
What is broken access control?
Broken access control is a failure to enforce what an authenticated user is allowed to do or see. The user proves who they are, and the application then hands them data or actions belonging to someone else. Two named cases account for most of what teams encounter.
1. BOLA: broken object-level authorization
BOLA is the failure to verify that the requester is entitled to the specific object they requested. It ranks first in the OWASP API Security Top 10, which puts it at the top of both OWASP lists that apply to modern applications.
2. IDOR: insecure direct object reference
IDOR is the same failure viewed from the attacker's side. The application exposes an internal identifier such as a record ID, and substituting a different value returns data belonging to another account. A single bug is usually both BOLA and IDOR.
OWASP moved this category to the top of its list in 2021, and it has stayed there through the 2025 edition. These are simple flaws. They hold the position because they are easy to introduce, hard to detect automatically, and immediately valuable once found.
Why do AI coding agents get authorization wrong so often?
Software development went agentic, but security did not. Application security was built around a scan that matches known bad patterns and hands off results to people who know the product's rules. When an agent writes the endpoint, nobody in the loop adheres to those rules by default, and the resulting flaw has no pattern to match. Closing that gap takes more than another rule.
Because the authorization requirement was never stated, and the agent is not wrong about the task it was given. Consider a developer asking a coding agent to add an endpoint that returns an invoice by ID. The agent produces a route that verifies the caller is logged in, looks up the invoice using the identifier from the URL, returns a 404 when nothing matches, and sends the record back to the client.
Every one of those decisions is correct for the task as written. The endpoint authenticates, it handles the missing-record case, and it reads cleanly in review. It also lets any authenticated user read any invoice in the system by changing one value in the URL.
The correction is one additional condition on the lookup: match the invoice by its identifier, and by the organization the caller belongs to. Nothing else about the endpoint changes.
Why the fix is unknowable from the code alone
Knowing that the second version is right requires three facts that appear nowhere in the code being generated:
Invoices belong to organizations.
A user's access is scoped to the organizations they are a member of.
Cross-organization reads are never legitimate in this product.
In a different product, the third fact could be false, since some platforms let auditors, resellers, or parent accounts read across tenant boundaries by design. The rule is a property of the product, not of the language or the framework. It lives in the data model, in a decision made two years ago, and in the heads of the engineers who made it.
What a reviewer should check instead
A human developer on that team usually catches this for one reason: they know the product. An agent working from the prompt and the surrounding file has no access to the thing that makes the check necessary.
Two details are worth carrying into review. The first is the response code. Folding the ownership condition into the lookup means a request for another organization's invoice returns the same 404 as a request for an invoice that does not exist. That is the behavior you want, because a 403 would confirm that the record exists and allow an attacker to enumerate valid identifiers. The second is what comes back. Returning the stored record unmodified sends all its fields, so the same handler frequently carries a data exposure issue alongside the authorization one.
Can SAST tools find broken access control?
For vulnerability classes that have a pattern, detection is largely solved. Authorization is the exception, and it sits exactly where AI-generated code is weakest. Static analysis traces a describable shape, and object-level authorization does not present one.
Take an injection as the contrast: an injection vulnerability has a shape that can be written down: untrusted input reaches a dangerous sink without passing through sanitization. That shape becomes a rule, and an engine follows the data from source to sink and reports every matching path.
What static analysis does cover inside broken access control
Broken access control is the largest category in the OWASP Top 10, and not all of it is resistant to scanning. A01:2025 maps 40 CWEs, and several have exactly the traceable shape that static analysis handles well, including path traversal, open redirect, and server-side request forgery. Modern semantic engines go beyond signature matching by modeling data flow and code intent, and they can flag the structural absence of an authorization decorator or middleware on a route when the codebase follows a consistent convention.
Where the coverage boundary sits
The part that resists analysis is narrower. It is object-level authorization, mapped as CWE-639 (authorization bypass through user-controlled key), CWE-862 (missing authorization), and CWE-863 (incorrect authorization). Here, nothing dangerous is called, no untrusted value reaches anywhere it should not, and every line is idiomatic. The defect is the absence of a comparison that only the application's own rules require.
To flag it, an engine would have to know which fields represent ownership in this schema, which callers are permitted, and where the boundary between tenants sits. None of that is a question of the rule's quality. It is information that does not exist inside the file being scanned.
Deterministic engines and reasoning belong together rather than in competition. Engines return the same result every time for the classes they cover, which is what lets a team gate a pipeline on them. Object-level authorization lies outside what a rule can describe, so it requires a different instrument to address it.
How do you find flaws that have no pattern?
You can find flaws that don’t have a pattern by reasoning over a model of the application rather than over the file in front of you.
Finding the invoice bug requires four facts: what calls this endpoint, what the object being fetched belongs to, which users are entitled to it, and where the trust boundary between tenants sits. Those facts are recoverable from the codebase, the schema, and what is running in production, but they have to be assembled into something an analysis can read. Snyk calls that assembled model the application-context graph: architecture, data flows, data classifications, trust boundaries, and production reality, built once and updated as the code changes. Once it exists, the missing ownership check becomes visible because the analysis can compare what the endpoint enforces against what the application's own rules require.
Three conditions make the result trustworthy.
The reasoning has to be grounded in that model. Without it, frontier reasoning produces plausible findings about a codebase it has partly imagined, and it burns tokens scanning blindly.
The findings have to be confirmed by something that did not produce them. An agent that finds a vulnerability cannot be trusted to validate its own fix, and a system that grades its own output carries forward whatever its analysis missed.
The fix has to be proven, not only proposed. The one-line correction has to be validated against the application's rules and shown to hold as the code changes. A finding that never becomes a merged fix leaves the endpoint exactly as exposed as it was.
Reasoning over the assembled application, with deterministic engines running alongside it, is the direction Snyk is taking with Evo Agentic AppSec, first previewed in August. Runtime testing then confirms what an attacker can reach, which is where Evo Continuous Offensive Security fits today.
What should your team do about this now?
Five steps, none of which require new tooling to begin.
Inventory the endpoints that accept an object identifier from the request. Every route that takes an ID, filename, or key out of a URL or body is a candidate. This list is usually shorter than teams expect and longer than they hope.
Write down the ownership rules. For each resource type, record what it belongs to and who may read or modify it. Authorization flaws survive review because this information has never been written anywhere that a reviewer or an analyst can check.
Make authorization an explicit review item on agent-authored pull requests. Not a general instruction to review carefully, but a specific question: does this endpoint verify that the caller is entitled to the object, and against which field?
Add a cross-tenant test for each resource type. OWASP's prevention guidance is direct on this: developers and QA staff should include functional access control in their unit and integration tests. One test per resource, asserting that a valid session from one tenant gets a 404 on another tenant's object, turns the rule you wrote down in step two into one that stays enforced.
Add deep analysis across the whole codebase on a regular cadence. The control point has split into three: real time inside the agent's inner loop, fast deterministic scans on every change in CI, and deep contextual analysis across the whole codebase. Object-level authorization flaws surface in the third.
Which of your endpoints would fail that cross-tenant test today? Start with the inventory in step one. To see where Snyk is taking agentic application security, read A First Look at Evo Agentic AppSec.
Frequently asked questions
What is the difference between BOLA and IDOR?
They describe the same failure from different angles. BOLA names the missing control: the application never verified that the requester was entitled to the specific object. IDOR names the exposure that makes it exploitable: an internal identifier is visible to the client and can be changed. A single bug is usually both.
Can SAST tools detect broken access control?
Partly. Static analysis covers several members of the category well, including path traversal, open redirect, and server-side request forgery, and a semantic engine can flag a route missing an authorization decorator that the rest of the codebase applies. What it cannot determine is whether an authorization check is the correct one, because that depends on which field encodes ownership and which cross-tenant reads the product intends to allow.
Why is broken access control number one in the OWASP Top 10?
Because it combines high prevalence with high impact. In the 2025 edition, every application tested showed some form of it, and the category recorded the highest number of occurrences of any entry. A single instance often exposes every record of a given type rather than one.
Does AI-generated code contain more authorization flaws than human-written code?
The mechanism is clear: an agent generates code from a prompt and its surrounding context, and the facts that make an authorization check necessary appear in neither. Measurement is thinner than the mechanism, because most published studies of AI-generated code security test injection, cross-site scripting, and cryptographic misuse, which are the categories with automatable ground truth. Treat the direction as well-supported and specific per-category figures as still emerging.
How do you test for broken access control?
Three approaches, in increasing order of coverage. Manual testing is the most reliable single method: attempt to access another account's objects with a valid session. Automated functional tests keep the rule enforced as the code changes, with one cross-tenant test per resource type. Continuous dynamic testing exercises running APIs from the outside against a model of what each resource belongs to and who may reach it, rather than scanning the source for patterns.
BOOK A LIVE DEMO
Secure AI adoption at scale
Evo helps organizations safely adopt and scale AI by providing visibility, governance, and security across AI-driven development and AI applications.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.