Your AI wrote code that runs. That's not the same as code that works.
An AI coding assistant is a tool that generates code from a natural-language description of what you want. The failure mode that bites teams isn't syntax errors — those get caught immediately. It's code that compiles, passes a shallow test, looks plausible in review, and then quietly does the wrong thing in production.
The reason is structural, not mysterious. A language model is optimized to produce text that looks like a correct answer. It has no runtime. It can't execute your code, can't observe your database, and can't feel the difference between "this returns a list" and "this returns the list in the order the caller depends on." So it fills gaps with the most statistically likely completion. Confident, fluent, and sometimes wrong.
The pattern that actually works is to stop asking the model to be correct and start making correctness checkable. That means writing down what "done" means before generation, forcing the AI to argue against its own output, and gating the merge on evidence rather than vibes. Below is how to build that loop, plus a worked example of where it catches something a normal review wouldn't.
Why "it compiles" is the wrong success signal
Compilation proves type agreement. Tests prove that the cases you thought of behave as you expected. Neither proves the code does what the ticket asked for.
Here's the mechanism: when you prompt with something vague like "add pagination to the users endpoint," the model has to invent several decisions you never specified. Does the page number start at 0 or 1? Is the limit a max or an exact count? What happens when the caller asks for page 900 of a 12-page result set — empty list, or error? Each of those is a fork, and the model picks one silently. It won't flag the ambiguity because flagging it would look like an incomplete answer.
That's the confidence problem in one sentence: the model resolves ambiguity by guessing, and guessing doesn't announce itself. Your review then reads the code as written rather than the spec as intended, so the guess survives.
Step 1: Write the contract before you write the prompt
Before you ask for any code, write down the interface in a form the model can't quietly reinterpret. Not a paragraph — a list of explicit assertions.
- Inputs and their valid ranges (including what's invalid).
- Output shape and ordering guarantees.
- Behavior at the boundaries: empty input, one item, exactly at the limit, one past the limit.
- Error semantics: which conditions throw, which return empty, which are caller errors.
This is the highest-leverage step and the one people skip. A contract turns an open-ended generation task into a constrained one. It also gives you something to test against later that didn't come from the model itself — which matters, because a test the model wrote to satisfy its own code will usually pass.
If you're using an assistant that takes a plain description and handles the prompt construction for you — AI-Mind works this way, taking a described task and a content type rather than a hand-written prompt — the contract still has to come from you. No tool knows your ordering guarantees.
Step 2: Make the AI argue against its own output
Once you have a draft, don't review it. Interrogate it. Paste the code back with the contract and ask a specific adversarial question:
Here is the contract and here is the implementation. List every input where this code's behavior is not determined by the contract, and every input where it contradicts the contract. Do not suggest fixes yet.
The "do not suggest fixes yet" clause matters. If you ask for problems and fixes together, the model tends to produce a short list of cosmetic issues and then a rewrite, which buries the real gap. Separating the two passes forces the critique to stand on its own.
This works because you've changed the task from "generate plausible code" to "find inconsistencies between two texts." The second task has a checkable answer, and models are noticeably better at it. You're not making the model smarter — you're giving it a job where being right is easier than being fluent.
Step 3: Gate the merge on evidence, not on review
A pull request that says "looks good to me" is not a gate. A gate is something that fails loudly and blocks. Three cheap ones cover most of the damage:
- Boundary tests written from the contract, not from the implementation. If a test was generated by reading the code, it encodes the code's bugs.
- A second model as reviewer. Run the diff past a different assistant with the same contract. Two models disagreeing is a strong signal; two models agreeing is weak evidence but not nothing.
- A named human owner per change. Someone has to be accountable for the decision the model made silently. If nobody can state the pagination base in a sentence, the guess is still live.
None of these are expensive. The boundary tests take minutes once the contract exists. The cost is the contract, and that cost is real — see the limits section below.
A worked example: pagination that passes review and fails the contract
This is a hypothetical walkthrough, not a reported run — the point is the shape of the failure, which is common enough to be worth recognizing.
Say the ticket reads: "Add pagination to the users endpoint. Support page and limit parameters." You hand that to an assistant. It returns something like a query that computes an offset from page * limit and returns the rows.
Read it as written and it's fine. Now apply the contract. Suppose your contract says page numbering starts at 1, limit defaults to 20 and caps at 100, and a request past the last page returns an empty list rather than an error.
Three gaps surface immediately. The offset formula treats page 1 as the second page, so every caller silently skips the first twenty users. The limit cap is absent, so a caller passing limit=100000 gets a full-table scan. And the out-of-range behavior was never specified, so the model picked one — probably empty, but you can't tell from the code whether that was a decision or an accident.
The first gap is the dangerous one. It doesn't throw, doesn't log, and returns a plausible-looking page of results. A reviewer scanning for correctness sees a working query. Only the contract catches it, because the contract is the only artifact that says page 1 means the first page.
Where this pattern breaks down
Honest limits, because the pattern isn't free:
It costs time up front. Writing a real contract for a non-trivial endpoint can take longer than the generation did. On a throwaway script or a prototype you'll delete next week, that trade is bad — skip it and accept the risk knowingly.
It doesn't catch wrong requirements. If the ticket itself is wrong, a perfect contract enforces the wrong thing with great precision. The pattern verifies implementation against intent, not intent against reality.
It degrades on large diffs. Adversarial review works best on a focused change with a clear contract. Hand a model a thousand-line diff and the critique goes shallow — it can't hold the whole thing in view any better than you can.
It assumes you can state the contract. On genuinely novel logic — a new ranking algorithm, an unfamiliar protocol — you may not know the boundary behavior until you see it fail. In that case the honest move is to treat the first version as a probe and write the contract from what you learn, not to pretend you had one all along.
The habit that matters more than the tooling
Every assistant in this space — Copilot, Claude, Cursor, Gemini, whatever you're on — shares the same property: it produces output optimized to look right. That's not a bug you can prompt away, and switching tools won't fix it. What fixes it is refusing to treat generation as the end of the task.
Write the contract first. Make the model attack its own work in a separate pass. Gate the merge on tests derived from the contract rather than the code. The specific tools matter far less than the order you do those three things in. If you're evaluating assistants, a maintained snapshot of what each one costs and does is more useful than a feature list — this site keeps a verified database of 360 AI tools with pricing and capability snapshots recorded at a known date, which is the kind of thing worth checking before you standardize on one.
Start with the next ticket you'd normally hand straight to an assistant. Write the contract first — inputs, outputs, boundaries, error cases — and see how many of the model's silent decisions become visible before a single line of code is generated.
Key Takeaways
- Compiling code proves type agreement, not that the code does what the ticket asked.
- AI resolves ambiguous specs by guessing silently; the guess never announces itself in review.
- Write an explicit contract — inputs, outputs, boundaries, errors — before prompting for code.
- Run adversarial review as a separate pass from fixing, or the critique goes shallow.
- Gate merges on tests derived from the contract, not from the generated implementation.
Sources
- AI Tool Database (internally verified snapshot), 2026. Maintained record of 360 AI tools with pricing and capability snapshots, most recently verified 2026-09-18.
Frequently Asked Questions
Does this pattern work with any AI coding assistant, or only some?
The pattern is tool-agnostic because it doesn't depend on model quality — it depends on the order of operations. Contract first, adversarial critique second, evidence-based merge gate third. Any assistant that can read a spec and read code can participate in that loop. What varies between tools is how well they handle long diffs and how reliably they follow the "don't fix yet" instruction. Test that on a small change before trusting it on a large one.
What if I genuinely can't write the contract before generating code?
Then you're exploring, not building, and you should label it that way. Treat the first generated version as a probe: run it, observe the boundary behavior, and write the contract from what you learn. The danger is treating a probe as a deliverable. If the code is going to ship, someone still has to be able to state the boundary rules in a sentence — otherwise the silent decisions are still live in production.
How is a boundary test different from a normal unit test?
A normal unit test often gets written by reading the implementation, which means it inherits the implementation's assumptions — including its bugs. A boundary test is written from the contract, before or independently of the code. It asks what should happen at empty input, at exactly the limit, and one past it. That independence is the whole point: a test derived from the code can't catch a decision the code made wrongly.