How to stop AI from confidently shipping broken code a pattern that actually works

Published: 2026-09-22 · Rewritten: 2026-09-23
A polished white pawn whose reflection reveals a cracked hollow interior, standing before a translucent glass wall.
Confidence and correctness are different things, and the reflection is where the difference shows. AI-generated illustration

Your AI wrote code that runs. That's not the same as code that works.

An AI coding assistant is a tool that generates code from a natural-language description of what you want. The failure mode that bites teams isn't syntax errors — those get caught immediately. It's code that compiles, passes a shallow test, looks plausible in review, and then quietly does the wrong thing in production.

The reason is structural, not mysterious. A language model is optimized to produce text that looks like a correct answer. It has no runtime. It can't execute your code, can't observe your database, and can't feel the difference between "this returns a list" and "this returns the list in the order the caller depends on." So it fills gaps with the most statistically likely completion. Confident, fluent, and sometimes wrong.

The pattern that actually works is to stop asking the model to be correct and start making correctness checkable. That means writing down what "done" means before generation, forcing the AI to argue against its own output, and gating the merge on evidence rather than vibes. Below is how to build that loop, plus a worked example of where it catches something a normal review wouldn't.

Why "it compiles" is the wrong success signal

Compilation proves type agreement. Tests prove that the cases you thought of behave as you expected. Neither proves the code does what the ticket asked for.

Here's the mechanism: when you prompt with something vague like "add pagination to the users endpoint," the model has to invent several decisions you never specified. Does the page number start at 0 or 1? Is the limit a max or an exact count? What happens when the caller asks for page 900 of a 12-page result set — empty list, or error? Each of those is a fork, and the model picks one silently. It won't flag the ambiguity because flagging it would look like an incomplete answer.

That's the confidence problem in one sentence: the model resolves ambiguity by guessing, and guessing doesn't announce itself. Your review then reads the code as written rather than the spec as intended, so the guess survives.

Step 1: Write the contract before you write the prompt

Before you ask for any code, write down the interface in a form the model can't quietly reinterpret. Not a paragraph — a list of explicit assertions.

This is the highest-leverage step and the one people skip. A contract turns an open-ended generation task into a constrained one. It also gives you something to test against later that didn't come from the model itself — which matters, because a test the model wrote to satisfy its own code will usually pass.

If you're using an assistant that takes a plain description and handles the prompt construction for you — AI-Mind works this way, taking a described task and a content type rather than a hand-written prompt — the contract still has to come from you. No tool knows your ordering guarantees.

Step 2: Make the AI argue against its own output

Once you have a draft, don't review it. Interrogate it. Paste the code back with the contract and ask a specific adversarial question:

Here is the contract and here is the implementation. List every input where this code's behavior is not determined by the contract, and every input where it contradicts the contract. Do not suggest fixes yet.

The "do not suggest fixes yet" clause matters. If you ask for problems and fixes together, the model tends to produce a short list of cosmetic issues and then a rewrite, which buries the real gap. Separating the two passes forces the critique to stand on its own.

This works because you've changed the task from "generate plausible code" to "find inconsistencies between two texts." The second task has a checkable answer, and models are noticeably better at it. You're not making the model smarter — you're giving it a job where being right is easier than being fluent.

Step 3: Gate the merge on evidence, not on review

A pull request that says "looks good to me" is not a gate. A gate is something that fails loudly and blocks. Three cheap ones cover most of the damage:

None of these are expensive. The boundary tests take minutes once the contract exists. The cost is the contract, and that cost is real — see the limits section below.

A worked example: pagination that passes review and fails the contract

This is a hypothetical walkthrough, not a reported run — the point is the shape of the failure, which is common enough to be worth recognizing.

Say the ticket reads: "Add pagination to the users endpoint. Support page and limit parameters." You hand that to an assistant. It returns something like a query that computes an offset from page * limit and returns the rows.

Read it as written and it's fine. Now apply the contract. Suppose your contract says page numbering starts at 1, limit defaults to 20 and caps at 100, and a request past the last page returns an empty list rather than an error.

Three gaps surface immediately. The offset formula treats page 1 as the second page, so every caller silently skips the first twenty users. The limit cap is absent, so a caller passing limit=100000 gets a full-table scan. And the out-of-range behavior was never specified, so the model picked one — probably empty, but you can't tell from the code whether that was a decision or an accident.

The first gap is the dangerous one. It doesn't throw, doesn't log, and returns a plausible-looking page of results. A reviewer scanning for correctness sees a working query. Only the contract catches it, because the contract is the only artifact that says page 1 means the first page.

Where this pattern breaks down

Honest limits, because the pattern isn't free:

It costs time up front. Writing a real contract for a non-trivial endpoint can take longer than the generation did. On a throwaway script or a prototype you'll delete next week, that trade is bad — skip it and accept the risk knowingly.

It doesn't catch wrong requirements. If the ticket itself is wrong, a perfect contract enforces the wrong thing with great precision. The pattern verifies implementation against intent, not intent against reality.

It degrades on large diffs. Adversarial review works best on a focused change with a clear contract. Hand a model a thousand-line diff and the critique goes shallow — it can't hold the whole thing in view any better than you can.

It assumes you can state the contract. On genuinely novel logic — a new ranking algorithm, an unfamiliar protocol — you may not know the boundary behavior until you see it fail. In that case the honest move is to treat the first version as a probe and write the contract from what you learn, not to pretend you had one all along.

The habit that matters more than the tooling

Every assistant in this space — Copilot, Claude, Cursor, Gemini, whatever you're on — shares the same property: it produces output optimized to look right. That's not a bug you can prompt away, and switching tools won't fix it. What fixes it is refusing to treat generation as the end of the task.

Write the contract first. Make the model attack its own work in a separate pass. Gate the merge on tests derived from the contract rather than the code. The specific tools matter far less than the order you do those three things in. If you're evaluating assistants, a maintained snapshot of what each one costs and does is more useful than a feature list — this site keeps a verified database of 360 AI tools with pricing and capability snapshots recorded at a known date, which is the kind of thing worth checking before you standardize on one.

Start with the next ticket you'd normally hand straight to an assistant. Write the contract first — inputs, outputs, boundaries, error cases — and see how many of the model's silent decisions become visible before a single line of code is generated.

Key Takeaways

Sources

Frequently Asked Questions

Does this pattern work with any AI coding assistant, or only some?

The pattern is tool-agnostic because it doesn't depend on model quality — it depends on the order of operations. Contract first, adversarial critique second, evidence-based merge gate third. Any assistant that can read a spec and read code can participate in that loop. What varies between tools is how well they handle long diffs and how reliably they follow the "don't fix yet" instruction. Test that on a small change before trusting it on a large one.

What if I genuinely can't write the contract before generating code?

Then you're exploring, not building, and you should label it that way. Treat the first generated version as a probe: run it, observe the boundary behavior, and write the contract from what you learn. The danger is treating a probe as a deliverable. If the code is going to ship, someone still has to be able to state the boundary rules in a sentence — otherwise the silent decisions are still live in production.

How is a boundary test different from a normal unit test?

A normal unit test often gets written by reading the implementation, which means it inherits the implementation's assumptions — including its bugs. A boundary test is written from the contract, before or independently of the code. It asks what should happen at empty input, at exactly the limit, and one past it. That independence is the whole point: a test derived from the code can't catch a decision the code made wrongly.

Identical parcels on a conveyor pass through a scanning arch while one is lifted aside for closer inspection.
Treat every AI output as a pull request from a stranger: most are fine, but each one gets checked. AI-generated illustration
A translucent glass bridge with hairline cracks across its middle, a small brass caliper measuring the fracture mid-span.
Some gaps AI simply cannot close yet, and knowing which ones is part of the workflow. AI-generated illustration

How this article was produced: it was generated by an automated content pipeline from the sources listed above. No human editor wrote or reviewed it, and we did not personally test the tools described. Facts and prices that appear here come from our own AI tool database, and its verification date is noted where relevant. Spotted an error? Tell us and we will correct or remove it.

Want to try this yourself? AI-Mind generates content from a plain description — no prompt engineering required.

Try AI-Mind