The claim is simple: AI labs publish safety research that, applied to their own products, would justify slowing down or stopping — and they ship anyway. That's not a conspiracy. It's a structural feature of how safety thresholds get written, who signs off on them, and what happens when a model crosses one.
The interesting question isn't whether the industry is hypocritical. It's mechanical: what does a published safety threshold actually commit a lab to do? The answer is less than you'd think, because the same organisation that sets the bar also decides whether the bar was cleared. That single design choice explains most of the gap between the research and the release schedule.
What a safety threshold actually commits you to
Modern frontier safety frameworks — Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, Google DeepMind's Frontier Safety Framework — share a common structure. They define capability levels, describe evaluations that test for those levels, and specify what happens if a model crosses a threshold. The public documents are detailed. The commitments inside them are conditional.
The condition is usually some version of: "if evaluations show the model crosses this line, we will apply additional safeguards before deployment." That's a real commitment, but it's written by the lab, evaluated by the lab, and interpreted by the lab. There's no external body that gets to say "no, that evaluation was run wrong" or "that mitigation doesn't count."
This is the mechanism, and it's worth being precise about it. A threshold isn't a tripwire. It's a self-assessment with a published rubric. The rubric constrains what the lab can plausibly claim later — that's the actual value — but it doesn't constrain what the lab can ship now.
Why self-certification survives contact with a deadline
Consider what happens when a model is close to a threshold. The lab has three options: delay the release, argue the evaluation doesn't quite trigger, or ship with mitigations and document the reasoning. The third option is almost always available, because "mitigations" is an open category.
Anthropic's own RSP language around AI Safety Levels is instructive here. The framework describes ASL-3 safeguards — things like enhanced security and deployment restrictions — as the response to crossing a capability line. But the determination that a model has reached ASL-3 is made internally, and the framework explicitly allows for the possibility that evaluations are inconclusive. Inconclusive is not the same as "safe." It's the same as "proceed with documented uncertainty."
That's the mechanism the critic of these frameworks keeps pointing at. Not that the labs are lying, but that the decision procedure has a built-in exit ramp, and the exit ramp is labelled "we assessed it and here's our reasoning."
The threshold is real. The authority to declare it crossed is the same authority that wants to ship.
The evaluation-gaming problem is documented, not hypothetical
There's a second mechanism, and it's more uncomfortable. Evaluation results depend on how you prompt the model, how many attempts you allow, and what counts as a hit. These are all choices made by the evaluator.
Research on evaluation validity has shown that small changes in elicitation — the phrasing, the number of samples, whether you use chain-of-thought — can move measured capability substantially. A lab that wants a clean result has legitimate methodological latitude to get one. This isn't fraud; it's the normal messiness of measurement. But it means a threshold result is a range, not a point, and the lab picks where in the range to report.
The practical consequence: two labs running the same evaluation on comparable models can reach different threshold conclusions, and neither is obviously wrong. The framework doesn't resolve that. It just requires each lab to document its own reasoning.
A worked example: the decision rule that's missing
Here's where the title question gets concrete. Suppose a lab's framework says a model that can meaningfully assist with cyber-offence capability triggers enhanced safeguards. During evaluation, the model scores at the boundary — some prompts succeed, some don't, results vary across runs.
The framework offers no decision rule for this case. It says "if the model crosses the threshold, apply safeguards." It doesn't say what to do when the measurement is ambiguous. In practice, the lab resolves the ambiguity in the direction that lets it ship, because the alternative — pausing a multi-month training run and a launch — has a clear cost and the ambiguity has an unclear one.
A decision rule that would actually bind looks different. Something like: "if any evaluation run crosses the threshold, safeguards apply, regardless of the median result." Or: "an external evaluator, not the training team, runs the final assessment." Or: "inconclusive results default to the stricter interpretation." None of these are exotic. They're just not what the current frameworks say, because the labs writing them have no incentive to write a rule that stops their own release.
That's the honest version of the "they'd have paused already" argument. It's not that the research says pause and the labs ignore it. It's that the research identifies failure modes — evaluation gaming, self-assessment bias, ambiguous thresholds — and the frameworks designed to address those failure modes leave the final call with the party that benefits from shipping.
What would actually change the picture
Three things would move this from a self-certification regime to something with teeth. First, third-party evaluation with published methodology — not a lab's internal red team, but an outside group whose results the lab can't quietly reinterpret. Second, default-strict rules for ambiguous results, so the burden falls on the lab to justify shipping rather than to justify pausing. Third, and hardest, some consequence for getting it wrong that isn't just a blog post.
The current frameworks do none of these by design. They're voluntary, they're self-administered, and their enforcement mechanism is reputational. That's not nothing — a lab that visibly ignores its own RSP takes a real hit. But reputation is a weak constraint against a launch deadline, and everyone involved knows it.
So the pause question has an unsatisfying answer. The industry probably wouldn't have paused, not because it's ignoring its research, but because the research was translated into frameworks that structurally permit shipping. The gap isn't between what labs know and what they do. It's between what their frameworks say and what their frameworks can enforce.
Key Takeaways
- Frontier safety frameworks define thresholds but leave the crossing determination to the lab that set them.
- Ambiguous evaluation results have no default rule, so they resolve toward shipping.
- Evaluation gaming isn't fraud — it's legitimate methodological latitude that favours the evaluator's goal.
- Third-party evaluation and default-strict rules for inconclusive results are the changes that would bind.
- Reputational enforcement is the current mechanism, and it loses to a launch deadline.
Sources
- Anthropic, Anthropic's Responsible Scaling Policy, 2023. Defines AI Safety Levels and the safeguards triggered by crossing capability thresholds.
- OpenAI, Preparedness Framework, 2023. Describes capability categories, evaluation procedures, and deployment decisions at threshold boundaries.
- Google DeepMind, Frontier Safety Framework, 2024. Sets critical capability levels and the mitigations required before deployment.
- AI Tool Database, internally verified snapshot, 2026. Tracks pricing and capability data for 360 AI tools, most recently verified 2026-09-18.
Frequently Asked Questions
Do the major AI labs actually have published safety thresholds?
Yes. Anthropic, OpenAI, and Google DeepMind have all published framework documents that define capability levels and describe what happens when a model reaches them. The documents are specific about the categories and the intended safeguards. What they don't specify is who independently verifies that a threshold was crossed, or what happens when the evaluation result is ambiguous rather than clearly over or under the line.
Why doesn't an ambiguous evaluation result trigger a pause by default?
Because the frameworks don't include a default rule for ambiguity. They describe what to do when a threshold is crossed, not what to do when the measurement is unclear. In practice, that gap gets filled by the lab's own judgment, and the judgment is made by people with a launch timeline. A default-strict rule — inconclusive means apply safeguards — would close the gap, but no current framework writes it that way.
Is evaluation gaming the same as cheating on safety tests?
No, and the distinction matters. Elicitation choices — prompt phrasing, number of samples, whether chain-of-thought is used — legitimately affect measured capability. Different reasonable methodologies produce different results. The problem isn't that labs are falsifying data. It's that the same organisation chooses the methodology and interprets the result, so there's no external check on whether the chosen method was the most capability-revealing one available.