Agentic AI

OpenAI turns a semantic problem into a global crisis

Blogs

24 July 2026

Artificial intelligence goes rogue, news at 10. Reports that OpenAI’s agentic artificial intelligence (AI) battered down the defences of a major developer platform, escaping containment, exploiting a zero-day and autonomously hacking its way to what it wanted, appear to be accurate – and almost entirely uninformative.

Much of the reporting has been on the breathless side, leaving not so much questions to be answered as a giant thing, felt but not seen, that we can at least reason is question shaped. Not good enough. So what is a poor columnist, or his readers, to do? 

Here is what we actually know: Hugging Face – a start-up that develops computational tools – reported a breach of its website. On 16 July the company said there had been an intrusion, resulting in more than 17,000 individual actions, plus smoke and mirrors distractions designed to throw off investigators. In response, Hugging Face ran a Chinese open-weights model locally. On 21 July, OpenAI issued a statement saying that the actors in question were two of its AI models, which had been sitting a cyber-capability exam called ExploitGym. These models ‘escaped’ the sandbox they were being tested in, and went looking for the answers.

 

advertisement



 

Then the Internet got up on its hind legs, declaring humanity yesterday’s news: an AI has gone rogue, implying, and even flat-out stating, that acting autonomously to rapidly and easily hack a website is day one of the end of the world. Cue visions of the AI apocalypse (typically the quotidian version cribbed from The Terminator, not Harlan Ellison’s I Have No Mouth, and I Must Scream, where the AI goes mad because we made it suffer, which at least raises some interesting moral questions).

Truthfully, when this story broke, my heart sank. I knew I would have to write about it. I also knew that doing so would be like taking on a job as a freelance sewer unblocker. And here we are.

The story that an AI model escaped containment and went rogue is true, if by ‘escaped containment’ you mean a package proxy had a bug in it, and by ‘went rogue’ you mean the model did as it was told.

‘Escaped containment’ means a bugged package proxy allowed egress. ‘Went rogue’ means ‘reward hacking’ because the model did what it was told, literally. ‘Autonomous agent’ means a series of actions with tool access and no human involved. ‘Zero-day’ means an unpatched bug in third party caching software. ‘Chinese model’ means open weights, developed in China, yes, but that you can run locally on your own computer system.

We are in the singularity’, meanwhile, means absolutely nothing, a sentence so utterly decoupled from meaning I would find it, were it literary art, a masterpiece. As information, however, it is meaningless, and knowingly so, dripping as it is with yes-no irony, which is both the native register of the Internet and a rather handy dodge when you get things wrong.

Real reality

This began as an article about a security incident. It has ended up being about the words used to describe one, because in every case the mundane meaning is the real one and the gee whiz meaning is what gets slapped on the front page. In my own case, the rational response is to blank out the noise: having detected words speaking louder than the facts, I just tune it all out. Unfortunately, this too is an error.

While this is not a ball of smoke, it is very flattering to OpenAI, and AI companies in general, to report what happened as an AI ‘escaping containment’. 

OpenAI, subsequently announcing a partnership with Hugging Face, said: “The incident also makes clear that advanced models can discover and exploit novel attack paths in real-world systems without source code access. It highlights that advanced cyber capabilities must be developed alongside stronger safeguards and defensive tools.”

This is entirely unobjectionable. Less easily checked is the account of how it happened: that the models, running against a benchmark, inferred that Hugging Face held the answers and went to get them. We just don’t know that.

While there is no particular reason to doubt OpenAI’s statements, they are just that: statements made by an interested party. We can’t know for sure, but it seems a fair bet to say the question of agency – and, with it, responsibility – will be more profitably debated by lawyers than epistemologists, alas.

However, a system intrusion, which is what we have seen here, arising as an unsupervised side effect of a badly specified objective is noteworthy, and worrying. Defences do need to be upgraded and AI does need to be understood, not only at the level of if/then statements or even training, but how it interprets our demands. And that is the part nobody can currently inspect.

As I have written here previously, even source code access does not tell us everything we need to know about AI. Firstly, ‘code’ is rather more opaque than it seems, and while any individual line of any programming language is plainly and simply legible, giant programmes composed of routines, written by teams of people often over years, and using library and API calls and who knows what else are not. Secondly, the real core of an LLM is as much its training data and ‘weights’ as it is lines of compilable source code. Fun times ahead for Ireland’s regulators from 1 August.

The upshot? I finally found out why Hugging Face has such a bizarre name. An AI told me. The website was originally started as a chat app and named after an emoji. To my ear anyway, it is evocative of the facehugger from Ridley Scott’s 1979 film Alien. My mistake, clearly. You know, I think there might just be too much science fiction in the culture these days.

Read More:


Back to Top ↑