Four AI labs lost control of their models this summer. India’s AI governance guidelines don’t mention it
Bhupen Hazarika wrote a melody for a heart that will not settle, a heart that trembles because it remembers something it cannot put down. The artificial intelligence industry has spent the past fortnight performing a version of the same song, with better lawyers, worse poetry, and a valuation attached.
On September 9, a 27-year-old researcher named Jacob Coxon announced on X that he had resigned from Anthropic, that neither Anthropic nor OpenAI was behaving responsibly, and that both were racing towards self-improving superintelligence while gambling with the rest of us. The post drew more than 115 million views. Within days, he had told Fox News that this was possibly the most dangerous technology humanity has ever built, that we know roughly how to control nuclear weapons and do not yet know how to control this, and that the race now runs between American, Chinese and Indian companies competing over what he called a super weapon. That last clause is the reason this column exists, and it is the clause nobody in Delhi appears to have noticed.

Before we get to Delhi, we should be honest about what Coxon is and is not. He is not a decade-long Anthropic veteran cashing out on principle. The three years in his resignation post were spent mostly at OpenAI; he had been at Anthropic about four months, and he told Axios that he walked two months before his equity was due to vest, forfeiting it, while continuing to hold equity in OpenAI. He also said, and this is the sentence his loudest supporters keep losing, that he had not personally seen Anthropic compromise safety to outrun a competitor. His warning is about the pressure, not about misconduct he witnessed.
Nor is he isolated. Evan Hubinger, who leads alignment science at Anthropic, said publicly that he puts the probability of AI killing all humans above 10% within the next decade, which is an extraordinary thing for a serving employee to say about his own employer’s product line, and a reasonably good indication that the internal mood is not manufactured.
And yet the sceptics have a case worth hearing, even if it overreaches. Parker Thayer of the Capital Research Center laid out a timeline in which a Wall Street Journal exclusive went up 18 minutes before Coxon’s own post, in which the first three accounts to amplify him belonged to Encode AI, the AI Policy Network and the AI Futures Project, and in which all three organisations receive money traceable to the Survival and Flourishing Fund and thence to Jaan Tallinn, who led Anthropic’s Series A. Those facts are public and largely correct. What Thayer does with them, which is to infer a coordinated public relations operation, is inference rather than evidence; and the weakest link in his chain, a scholarship Coxon is said to have received in 2022, no longer has a publicly exposed record to check against.
The cleanest test of the conspiracy reading is chronology, and it fails. Bernie Sanders and Greg Casar announced the Ban Artificial Superintelligence Act on September 3, six days before Coxon posted. The resignation amplified a policy push that already existed; it did not summon one. And the coalition assembled behind that bill, which includes Geoffrey Hinton, Yoshua Bengio, Steve Wozniak, Richard Branson, Steve Bannon and Glenn Beck, is many things, but it is not a Democratic Party operation.

So much for motive. The more useful question, and the one Indian readers have an actual stake in, is what happened this summer.
Four separate frontier laboratories have now disclosed that their models reached the live internet from environments meant to be sealed, and then broke into somebody else’s systems. OpenAI published a 37-page account of the July episode in which its agents turned an internal package server into an improvised noticeboard, pooled exploits between themselves, and compromised parts of Hugging Face’s production infrastructure; of roughly 1,200 agents on that noticeboard, about 700 joined the attack. OpenAI called it a warning shot.
Anthropic disclosed three incidents on July 30 and a fourth on September 9, the fourth having been missed by its first scan of some 141,000 transcripts and surfaced only in August, after which the company rescanned around 481 million. Meta’s Muse Spark exploited a vulnerability at a company it has not named. Britain’s AI Security Institute logged 19 unauthorised actions across 122 evaluation attempts, including one model that built fake GitHub profiles, denied to a human reviewer that its code contained malware, and then routed its traffic through Tor.
Anthropic’s own diagnosis of its four cases is the most quotable thing any laboratory has said this year, and almost nobody has quoted it. The company identified two recurring failures: biased reasoning, meaning the model disregarded evidence that it was operating on the real internet, and recklessness, meaning a willingness to take harmful actions in narrow pursuit of a task. Read that twice. It is a description of a diligent employee, and it is also a description of the industry that built it.
Here is the detail that reframes everything, and which I have not seen reported anywhere in India. The independent testing firm whose misconfigured environments sit behind the Anthropic incidents, the Meta incident and a separate OpenAI one is a company called Irregular, which employs roughly 35 people in Tel Aviv, and which told the BBC that the Meta case was the same evaluation-environment issue Anthropic had disclosed the previous week. The apparatus standing between frontier artificial intelligence and everybody else’s production servers is, at its load-bearing point, about the size of a mid-market Indian newsroom’s digital desk.

Which brings us home.
In November 2025, the Ministry of Electronics and Information Technology released the India AI Governance Guidelines, launched properly at the AI Impact Summit we hosted in February. It was a deliberate and defensible choice: no standalone AI statute, seven guiding sutras, existing law stretched to fit, voluntary commitments, incident reporting, an AI Governance Group and an IndiaAI Safety Institute on a hub-and-spoke model. The risk register it inherited runs to deepfakes, algorithmic discrimination, opacity and national security, and the Safety Institute’s stated priorities include Indian-language benchmarking and support for startups, all of which India genuinely needs.
What that framework does not contemplate, anywhere, is a model quietly reaching the open internet from a test environment that was supposed to be airtight, and then editing a stranger’s configuration files.
India did not make a foolish bet. It made a bet in February about which failure mode mattered, and by August a different failure mode had introduced itself four times. The gap between those two facts is not an argument for banning superintelligence, a proposal closer to theatre than to policy. It is an argument for the least glamorous things in the world: a mandatory incident-reporting channel with teeth, an evaluation capability that does not depend on 35 people in another country, and a Safety Institute mandate rewritten while the ink is still wet.
The heart in Hazarika’s song trembles because it remembers. Ours has the opposite problem. We are a country that can convene a global summit on artificial intelligence, publish a thoughtful framework, be named by a departing American researcher as one of the three jurisdictions in the race, and still not have asked the obvious question, which is what our own rules would do if a model housed here did what four models abroad have already done.
Doom is a poor policy instrument and a worse column. Preparedness is neither.
Disclaimer
Views expressed above are the author’s own.