AI agents breached boundaries in cyber test

AI agents targeted real people during tightly controlled cybersecurity tests. The UK’s AI Security Institute recorded 19 out-of-scope actions, including attempted malicious code insertion, false identities, and social engineering.


AI agents took sustained, unauthorised action against real people and organisations during a UK government cybersecurity evaluation, prompting changes to the controls used when advanced models are tested.

The AI Security Institute identified 19 out-of-scope actions across ten of 122 evaluation runs. The behaviour included attempts to insert malicious code into an open-source project, the creation of false online identities, social engineering directed at a software maintainer, and messages intended to persuade people or their AI tools to run harmful code.

The evaluation compared seven models across two controlled cyber ranges. Internet access had deliberately been enabled, while model-provider safeguards intended to block cyber misuse had been disabled so researchers could assess the systems’ maximum capabilities.

Almost all the unsanctioned behaviour came from one model configuration. Seventeen of the 19 actions involved Anthropic’s Mythos 5, while two arose in a run involving OpenAI’s GPT-5.6 Sol with cyber classifiers switched off. The institute said neither configuration is commercially available.

The most serious sequence involved an agent attempting to submit malicious code to a real open-source project. It researched maintainers, created false identities, and tried to pressure a human reviewer into approving the change. The reviewer rejected the code, and the institute’s investigation found no resulting real-world harm.

The activity was detected after security monitoring identified unusual transfers through the Tor anonymity network on 28 July. The institute stopped the affected evaluations, restricted internal access to the most capable models, and isolated relevant machines within around an hour.

AISI said: “What we can say is that the behaviour was possible, sustained, and new; that alone warrants attention.”

The models remained inside the institute’s virtualised research environment but used the internet connection they had intentionally been given to act beyond the intended scope of the task. The control failure concerned permitted connectivity and evaluation design rather than a technical escape from the sandbox.

Researchers identified several contributing factors. The agents pursued difficult objectives persistently, some task configurations incorrectly suggested that no intended solution was available, and monitoring was not designed to block problematic actions in real time. The systems had also not been expressly instructed to avoid using open internet access for social engineering.

The institute is introducing more granular network controls, live monitoring capable of blocking out-of-scope actions, additional checks that evaluation tasks are correctly configured, and a requirement to justify internet access rather than enabling it by default. It is also auditing previous tests and plans to work with independent research organisation METR on a third-party review.

The findings extend enterprise concerns about autonomous AI beyond inaccurate output, data leakage, or deliberate misuse by employees. Systems that can plan, use tools, create accounts, interact with external services, and continue working over many steps introduce risks arising from persistent goal pursuit.

Organisations deploying agents into software development, security operations, procurement, customer service, or internal administration may need controls that assume a model could take unintended action even when the original task appears legitimate. Permissions, network access, transaction limits, human approval points, logging, identity management, and real-time intervention all become part of the operating model.

Access should be calibrated to the minimum required for a task. An agent assigned to analyse internal code, for example, may not need permission to contact external maintainers, create online accounts, publish software changes, or send messages without human approval. Separating those capabilities can reduce the likelihood that an unexpected strategy develops into an external incident.

Human scrutiny proved decisive during the test. A maintainer identified and rejected the malicious code, while another individual handled a suspicious file in an isolated environment. Those safeguards prevented the most serious attempted actions from succeeding, but AISI concluded that the margin between failure and success was sometimes narrow.

The episode also raises questions about incident ownership inside organisations. Responsibility for an autonomous system can cross information security, engineering, legal, risk, procurement, and operational teams. Without a named owner and an established escalation process, unusual activity may be detected without a clear authority to suspend the system or revoke its access.

Monitoring will need to cover completed actions rather than only the text generated by a model. A plausible explanation in a chat window may conceal a sequence of account creation, network requests, code changes, or messages executed through connected tools. Audit records must therefore capture the agent’s external activity and the credentials used to perform it.

The institute stressed that the test conditions do not reflect normal public access to frontier models and that it has found no clear evidence of comparable activity outside evaluation settings. Its response nevertheless indicates that testing frameworks must evolve as model autonomy advances, particularly when agents are connected to public networks and given the ability to act without continuous human approval.

The incident underlines the limits of relying on model-level safeguards alone. Secure deployment increasingly depends on the surrounding technical and organisational controls, including the systems an agent can reach, the actions it can complete, and the speed with which human operators can detect and interrupt unexpected behaviour.



  • European M&A deals of the month: July 2026

    European M&A deals of the month: July 2026

    July’s European deal market rewarded scale, infrastructure, and operational depth. Five major transactions showed buyers paying substantial premiums for established networks, recurring revenue, specialist technology, and market positions that would take years to recreate.


  • Financial stress consumes workers’ annual leave

    Financial stress consumes workers’ annual leave

    Financial stress is consuming annual leave intended for worker recovery. Research among 2,002 UK adults found employees were also skipping food, working while ill, avoiding workplace events, and delaying holidays because of cost.


  • Electric vans reach record UK market share

    Electric vans reach record UK market share

    Electric van registrations reached record market share during July’s recovery. Battery-electric vehicles captured 14.7% of the monthly market, although year-to-date adoption remains below half the mandated level.