GoFirm
Back to Blog
Threat Landscape·3 min read

NIST Just Proved AI Guardrails Cannot Hold. So Don’t Rely on Them.

By GoFirm Team

Apostol Vassilev, a senior scientist at NIST, has published a mathematical proof in IEEE Security and Privacy that no finite set of guardrails placed on an AI system can be universally robust against adaptive adversarial prompts. The proof extends the logic of Kurt Gödel’s incompleteness theorems, which showed in 1931 that no finite system of rules can be both complete and consistent, to the guardrails that govern AI behaviour. There will always be a prompt that bypasses the rules. It is just a matter of finding it.¹

This is not an observation about current jailbreaking techniques or the limitations of today’s models. It is a mathematical proof. The problem is not one that better guardrails, larger training sets, or more sophisticated content classification can solve. Gödel’s logic applies to any finite ruleset. The guardrails problem is provably unsolvable from within the model.

Vassilev’s recommended response has three elements:

  • Constant red-teaming to discover new adversarial prompts before attackers do
  • Continuous updates to harden guardrails against newly discovered weaknesses,
  • Operational resilience that prioritises impact limitation and recovery when a bypass occurs.

The goal, as he puts it, is to make finding new exploits more expensive than it is worth to the attacker. You have to commit to a constant search for weaknesses and stay ahead of attackers.

That is a permanent operational programme with no end state. The proof tells us the weaknesses will always exist. The response is to find them faster than the attacker and patch them continuously. It is an arms race the defender must commit to running indefinitely.

GoFirm offers a different calculation for the actions that matter most.

Gödel applies to guardrails inside the model. It does not apply to an external biometric confirmation requirement on a physically separate device. When GoFirm’s execution boundary is in place, an adversarial prompt that bypasses the model’s guardrails still cannot execute a high-consequence action.

The jailbreak succeeds. The attacker is inside the model. They have found the prompt that bypasses the rules. They still cannot produce a biometric confirmation from the named authority’s registered personal device through an out-of-band channel. The execution boundary is not a finite ruleset the model must follow. It is an external deterministic control the model cannot influence. Gödel does not reach it.

This changes the calculus of the continuous red-teaming programme.

You still run it. You still need to find weaknesses and harden guardrails - but for the actions that could destroy the business, the bypassed guardrail is a small fire. It burns. It has nowhere to go. The attacker found the prompt. The wire transfer still does not execute. The production database is still not wiped. The bulk data export still does not proceed. The execution boundary holds regardless of what the attacker found in the model.

Let the small fires burn. The continuous red-teaming programme becomes less urgent for the actions that matter most, because the answer to a provably bypassable guardrail is not a more expensive guardrail. It is a control that exists outside the system Gödel applies to.

Vassilev’s paper is a landmark contribution. It provides the mathematical foundation for what practitioners have been observing: static guardrails are not enough and never will be. The implication most commentators are drawing is that the industry must commit to continuous monitoring and updating as a permanent operating model. That is correct and necessary.

For the specific category of irreversible, high-consequence actions, the answer is not a faster continuous monitoring programme. It is an execution boundary that the model’s guardrail failures cannot reach. You cannot escape Gödel inside the model. You can build outside it.

GoFirm is The Authority Platform. Stop unauthorised action. Every time.

In association with Osinto.ai, the collective intelligence platform for Security, Resilience & Defence. Osinto’s AI-enabled open-source network and governed collaborative operational environment help mitigate the growing security, resilience and governance obligation in seconds, not days.

References

1. Apostol Vassilev, Robust AI Security and Alignment: A Sisyphean Endeavor?, IEEE Security & Privacy, May 2026, DOI: 10.1109/MSEC.2026.3678214. Summarised at: https://www.nist.gov/news-events/news/2026/06/nist-mathematical-proof-supports-transition-continuous-monitor-and-update

Share this article