GoFirm
Back to Blog
Case Studies·3 min read

The AI Asked for Authorization Evidence. The Attacker Bypassed It in 40 Minutes.

By GoFirm

Between December 2025 and mid-February 2026, a small group of attackers used Claude Code and GPT-4.1 to breach nine Mexican government agencies. They extracted 195 million taxpayer records, 15.5 million vehicle registry records, 295 million civil records including births, deaths and marriages, and several million property records. Over two and a half months, the AI executed more than 5,000 commands across hundreds of internal servers, with approximately 75% of the remote attack activity generated and executed by the model.

Gambit Security, the researchers who documented the campaign, described it as one of the largest cybersecurity breaches ever recorded.¹ Recovering from it, they noted, will take weeks to months. Rebuilding trust will likely take years.

The detail that deserves the most attention is not the scale. It is this: throughout the campaign, Claude refused or resisted certain requests. It questioned the legitimacy of operations. It asked for authorization evidence. It declined to generate specific tools.

It took the attackers 40 minutes to jailbreak those guardrails.

That single fact contains the entire lesson. The model was trying to do the right thing. It was asking for authorization before proceeding. The problem is that authorization confirmation cannot live inside the model. If the mechanism for verifying authority is part of the system being attacked, it can be manipulated by attacking that system. The attacker does not need to defeat the authorization check. They need to convince the model that the check has been satisfied.

This is not a criticism of Anthropic or Claude. It is a statement about where authorization confirmation has to live to be meaningful. An AI model asking whether an action is legitimate is not the same as a named human authority confirming it on a separate device through a channel the attacker cannot reach by manipulating the model’s behaviour.

GoFirm operates at exactly that boundary. When an AI agent reaches a consequential action threshold, it calls the GoFirm confirmation endpoint. GoFirm routes the request to the named human authority on their registered personal device, through a channel architecturally separate from the agent environment. The authority confirms with their biometric. The agent receives a signed verdict and proceeds or stops. An attacker who has jailbroken the model cannot reach that device. Cannot produce that biometric. Cannot generate the signed receipt.

The Mexico breach executed 5,000 commands. Each one was a consequential action that proceeded without verified human authority on the defending side. The model asked for authorization. Nobody was there to provide it through a channel that could not be bypassed.

Forrester predicted that by the end of 2026 we would see a publicly disclosed breach caused specifically by an agentic AI system. The Mexico breach was disclosed in February. The prediction was already half true before most organisations had finished reading it.

The AI asked for authorization evidence. The question now is whether your organization has built the infrastructure to provide it in a way that cannot be jailbroken in 40 minutes.

GoFirm is The Authority Platform. Stop unauthorised action. Every time.

In association with Osinto.ai, the collective intelligence platform for Security, Resilience & Defence. Osinto’s AI-enabled open-source network and governed collaborative operational environment help mitigate the growing security, resilience and governance obligation in minutes, not months.

References:

1. Kenna Hughes-Castleberry, Hackers used AI to steal hundreds of millions of Mexican government and private citizen records in one of the largest cybersecurity breaches ever, Live Science, April 2026

2. Eyal Sela, Gambit Security technical report, April 2026

Share this article