The AI slowdown argument is about whether model capability gains are flattening. It's a fine debate, and it's also a distraction. The measurable pressure point isn't the frontier — it's the review layer underneath it. AI tools now produce code far faster than human reviewers can read it, and every unreviewed line is a line where a missing authorization check, an unsanitized input, or a leaked secret can sit unnoticed for months.
That's the vulnerability explosion: not a sudden spike in clever new attacks, but a widening gap between how much code gets written and how much of it gets genuinely examined. If you clicked this headline expecting a trend chart, here's the honest version — the reference material for this piece contains no reported-vulnerability statistics, so I won't invent any. What it does contain is a snapshot of 360 AI tools with recorded pricing and capability data, verified most recently on 2026-09-18. That snapshot tells you something useful: the tooling layer is crowded and well-documented. The review layer is not. This article is about closing that gap in practice.
Why does more code automatically mean more vulnerabilities?
It doesn't, strictly. More code means more opportunities for vulnerabilities, and the conversion rate from opportunity to exploit depends entirely on review. The mechanism is simple: a security bug has to be introduced by someone (or something) and then survive every check between writing and production. AI changes the first half of that equation dramatically and the second half barely at all.
Consider a worked example. A team asks an assistant to add a "download invoice" endpoint. The generated code might include a route handler that takes an invoice ID from the URL, queries the database, and streams the file back. Functional. It works in the demo. What it may not include is a check that the authenticated user actually owns that invoice. That's an IDOR — insecure direct object reference — and it's the single most common shape of bug in AI-generated CRUD code, because the model optimizes for "does this work" and ownership checks are invisible to that test.
Nothing about that bug is novel. What's new is the volume. When a team ships ten endpoints a sprint, a human reviewer catches most of them by pattern recognition. When the same team ships forty because generation is cheap, the same reviewer is skimming.
What does a missing authorization check actually look like in a diff?
This is the part most review guides skip, so let's be concrete. You're looking at a pull request. The diff adds a handler. Read it in this order:
- Find the identity source. Where does the code learn who the caller is? A session, a JWT claim, a middleware-injected user object. If you can't point at the line, stop — the endpoint may be unauthenticated.
- Find the resource lookup. The query that fetches the invoice, the document, the order. Note its filter clause.
- Check whether the two meet. The dangerous pattern is a lookup by primary key alone:
WHERE id =?. The safe pattern ties the resource to the caller:WHERE id =? AND owner_id =?. If the owner constraint lives in application code instead of the query, verify it runs before the data is returned — not after. - Check the failure path. A missing record and a forbidden record should both produce a denial. If a not-found returns 404 and a not-owned returns 200 with an empty body, you've leaked existence.
That's a four-step read that takes under a minute once you've done it a few dozen times. The reason it's worth formalizing is that it's exactly the check AI reviewers tend to skip. Ask a model "is this secure?" and it will often answer about injection and secrets, because those are the patterns it saw most in training. Ownership logic is contextual — it depends on your schema — and contextual checks are where automated review is weakest.
The tooling layer is documented. The process layer isn't.
Here's where the internal snapshot is genuinely useful. A database of 360 AI tools with pricing and capability recorded at verification time tells you the market is saturated and the comparison problem is real. If you're choosing a code assistant or a review tool, the vendor's own page is the only reliable source for current pricing — these change constantly, and any figure quoted in an article is stale the moment it publishes. What the snapshot does let you do is reason about categories rather than brands.
Three categories matter for this problem, and they fail differently:
- Generation assistants (Copilot, Cursor, and similar) optimize for throughput. They will happily produce the vulnerable handler above. Their value is speed; their cost is that speed is exactly what outruns your review.
- Static analysis and SAST tools catch injection and known-bad patterns well. They are weak on business-logic flaws like ownership checks, because those require knowing your data model.
- LLM-based review sits in between. It can reason about intent — "this handler fetches a resource but never checks the caller" — but it's non-deterministic and will miss things a rule would catch. Treat it as a second reader, not a gate.
The honest limit: no single category covers the gap. If your team adopts AI generation without adding review capacity, you have not sped up delivery. You have moved the bottleneck downstream, where it's more expensive to find.
How do you make review scale without hiring?
You can't fully. That's the uncomfortable answer, and any guide that promises otherwise is selling something. What you can do is make each review cheaper by narrowing what a human has to look at.
A pattern that holds up: let automated tooling handle the mechanical classes — dependency scanning, secret detection, injection patterns — and reserve human attention for the contextual classes. Ownership, state transitions, and error paths are where humans still beat tools, and they're also where AI-generated code concentrates its mistakes. That division of labor means the human reviewer reads less code, but reads the right code.
Second lever: make the generator's output smaller. A 400-line generated diff is unreviewable. A 40-line one gets read. Asking for a narrow change, or splitting a large generation into reviewable commits, costs a little time upfront and buys back the review capacity you just spent. This is unglamorous and it works.
Third: write the ownership check into your scaffolding, not your prompts. If your framework has a standard way to fetch a user-owned resource — a helper that requires the caller's identity as an argument — then generated code that skips it stands out in the diff. Making the safe path the obvious path is worth more than any review checklist.
Where this advice breaks down
Two honest caveats. First, none of this addresses the case where the vulnerability isn't in the code at all — misconfigured cloud permissions, exposed storage buckets, and leaked credentials in CI logs are infrastructure problems that a code review won't touch. Second, the review-capacity argument assumes your team has reviewers with security instincts. If it doesn't, adding process won't create that skill; it'll just produce faster sign-offs. Training or an external audit is the real fix, and it's slower and more expensive than any tool.
There's also a version of this where the "explosion" never materializes as incidents, because teams quietly slowed their merge cadence or added gates. That would be a good outcome and it would look, from the outside, like nothing happened. The absence of a disaster is not evidence the risk was fake.
Key Takeaways
- AI increases the volume of code faster than review capacity grows, widening the window for undetected flaws.
- Ownership and authorization bugs are the most common shape in AI-generated CRUD code and the hardest for tools to catch.
- Read diffs for identity source, resource lookup, and where the two meet — in that order.
- Split large generated changes into reviewable commits; an unreadable diff is an unreviewed diff.
- No tool category closes the gap alone; infrastructure misconfigurations sit outside code review entirely.
The takeaway that actually matters
The AI slowdown debate is about the supply of capability. The vulnerability question is about the demand side — how much code your team can responsibly absorb. Those are different curves, and the second one is the one you control.
If you take one thing from this: measure your review latency, not your generation speed. Track how long a pull request sits before a human actually reads it, and how large the average generated diff is. Those two numbers tell you whether you have an AI problem or a review problem. Most teams that feel fast are just shipping the second one.
Related reading on this site: how to stop AI from confidently shipping broken code, and which retrieved documents should actually reach the model — the same filtering instinct applies to what reaches your reviewer.
Sources
- AI Tool Database, Internal verified snapshot of 360 AI tools with pricing and capability data, 2026. Most recent verification recorded 2026-09-18; used here to characterize tooling categories rather than quote current prices.
Frequently Asked Questions
Is the "vulnerability explosion" a documented trend or a prediction?
It's a structural argument rather than a measured trend. The supplied reference material contains no reported-vulnerability statistics, so any specific figure would be invented. The claim rests on a mechanism: AI raises code volume faster than review capacity, and unreviewed code is where flaws survive. Whether that converts into more incidents depends on whether teams add review capacity or absorb the change elsewhere.
Can't I just ask an AI to review the code it wrote?
You can, and it helps, but treat it as a second reader rather than a gate. LLM-based review reasons well about intent — spotting a handler that fetches a resource without checking the caller — but it's non-deterministic and misses things a fixed rule would catch. Static analysis covers injection and secrets reliably but is weak on business-logic flaws that depend on your schema.
What's the single fastest change to make today?
Split large generated diffs into reviewable commits. A diff nobody can read in one sitting is a diff nobody reviews carefully, regardless of what tools you run. Pair that with a standard helper for fetching user-owned resources, so generated code that skips the ownership check stands out immediately in the diff. Both changes cost little and buy back review attention.