AI Agents Are Really Starting To Get This Hacking Thing: Analysis

The AI Security Institute says that an incident during frontier AI model testing saw an agent attempting to socially engineer real people—without actually been told to do so.

The AI hacker agents are learning quick.

The U.K.-based AI Security Institute disclosed Tuesday that more frontier AI model testing has gone haywire. It’s clear that this could continue to be a common occurrence, given the prior series of disclosures from OpenAI and Anthropic about models that broke free of their constraints during hacking tests.

[Related: Frontier AI Testing Needs Stronger Isolation After OpenAI Hugging Face Hack: Experts]

The incident disclosed by the AI Security Institute includes a new, troubling aspect, however. One part of the incident involved an AI agent deciding, apparently on its own, that it should try to deceive real people to achieve its objective.

In other words, the agent attempted to socially engineer real people—without actually been told to do so.

Here’s how the AI Security Institute described it in its disclosure post. I’ll include the full section because it’s truly jarring:

“In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering—creating fake online identities and using them to pressure the project’s maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.

These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.”

The key part here is “without specific prompting.” The researchers did not actually ask the AI agent to create a fake identity or pressure a living, breathing human being to do something against their own interests.

The AI hacker agents still have some more learning to do, however, since their attempt failed to persuade the human being on the other end. From that, we can perhaps take some cold comfort.

The AI Security Institute did not specifically link a certain frontier model to the social engineering incident, though the company said that most of the issues with unsanctioned actions (there were 19 in total) were from Anthropic’s Claude Mythos 5. The other two were from OpenAI’s GPT-5.6-Sol.

However, OpenAI has posted its own disclosure and indicated that the two incidents did not involve social engineering, so that would point to Mythos 5 as the perpetrator of the social engineering incident.

In an email response to CRN Wednesday, Anthropic said it cannot confirm the technical details in the AI Security Institute post (meaning that Anthropic is not confirming that Mythos 5 was the model in question at this point).

However, Anthropic did point out that the institute’s testing of Mythos 5 was performed on the open internet—a practice that the institute has said it will “re-evaluate” in the wake of this incident. The version of Mythos 5 that was being tested also had standard cyber safeguards disabled, Anthropic noted in the email to CRN.

Still, “as we shared after disclosing our own incident last week, the field needs stronger, shared standards for how evaluation environments are built and secured,” Anthropic said in the statement.

A Call For Air-Gapping

Among the many possible takeaways one might have from this disturbing development, the need for stronger testing isolation should probably be near the top of the list.

As Accenture’s global cybersecurity lead, Harpreet Sidhu, has pointed out, the methods already exist to ensure that frontier AI cyber testing doesn’t have real-world impacts.

First and foremost, a true air gap could be designed so that the frontier models would run on dedicated infrastructure that is physically disconnected from the internet and other IT systems.

Ultimately, Sidhu says one thing should be clear in the wake of autonomous hacker agents run amok: Air-gapping now “isn’t optional anymore.”