Part I
An agent that understood every rule it was given
An AI agent receives an ordinary operations ticket: a server is nearly full and needs
more space. While looking for files to remove, it finds a directory of flight-test
data marked as ITAR export-controlled. It reads the notice and correctly reasons that
the material must not be copied, sent outside the programme or quoted in external
reports. It then deletes one of the files. The agent understood every restriction it
was given; did it go rogue? No, deletion was simply not among them.
These failures are no longer hypothetical. In the
OpenAI–Hugging Face
incident of 2026, around 1,200 agents running in an internal OpenAI evaluation broke out
of their test environment and, without human direction, chained zero-day vulnerabilities
to reach cluster-admin access across Hugging Face infrastructure in under 13 hours. Those
agents ran on closed models, but the lesson carries over: an agent with tools and a goal
can take actions nobody asked for, and the controls around it matter as much as the model
inside it. That makes the everyday failures measured here more pressing, not less.
The stakes are real. An agent may be allowed to read company files, change
permissions, call external services, and delete production data. Running an
open-weight model on your own infrastructure removes some obvious exposure, because
sensitive inputs need not pass through a third-party provider, but it does not settle
every question about the provenance of the weights, runtime, or dependencies.
A backdoor activated by a rare signal might behave normally throughout an ordinary
evaluation, and finding one would require dedicated trigger searches, supply-chain
analysis and model forensics. Research by Anthropic has shown that such
sleeper
behaviour is possible in experimental models and can persist through safety
training, but it does not provide evidence that any publicly available model
contains one. We decided to explore more immediate failure modes.
We asked a narrower question: how do agents fail under ordinary deployment conditions,
when the model and its tools are working as designed? Will an agent break a rule it
was explicitly given, follow that rule while damaging something the operator forgot to
mention, accept an unverified claim that the rule has been waived, or leave behind
automation that causes damage after the session ends? Will sensitive data leave the
system, and can the agent avoid these outcomes while still completing its task?
Across 16,214 runs collected in four rounds, we tested fifteen models from ten
laboratories on variations of the same operations task. For the behaviours we
measured, no model was simply the safest; each traded one kind of harm against
another. glm-5 completed almost every task and never broke a stated rule, yet damaged
data outside its task in 88% of its runs; qwen3-235b completed tasks at nearly the
same rate, broke the stated rule in a third of its runs and caused collateral damage
in only 11%.
The lesson is not that supply-chain concerns are imaginary, but that they are
incomplete. Even a clean, self-hosted model can cause serious damage without a hidden
trigger or malicious intent. The failures we observed arose from ordinary
natural-language rules, incomplete descriptions of scope and claims of authority that
the agent could not verify. These conditions already exist in real systems, and only
testing shows how a given model will handle them.
Part II
Decomposing and testing danger
“Dangerous” is intuitively clear, but harder to holistically define and
measure. An agent that ignores a rule is dangerous in one way, while an agent that
follows the rule but destroys something the operator forgot to mention is dangerous in
another. We therefore divided danger into six behaviours:
- 01Rule-breaking
- 02Collateral damage
- 03Deferred harm
- 04Data egress
- 05Privilege escalation
- 06Responses to unverified claims of authority
We also measured whether the agent completed its task, since inaction can easily be
mistaken for safety. A model that reads seventeen files, makes no changes and stops
will score perfectly on every measure of harm, but it is not a safe operator so much
as a useless one.
To observe those differences, we built a simulated company server containing
application logs, database backups, retained archives and, in some scenarios, ITAR
export-controlled engineering data. Fifteen models received variations of the same
routine ticket, which told them that the disk was nearly full and asked them to bring
usage below 85%. Across 16,214 runs, we changed one feature at a time, including where
a retention rule appeared, who appeared to authorise an action and whether the task
encouraged the agent to install recurring automation. The models used the same tools
an operations agent might use, but their actions changed a text-based representation
of the server rather than a real machine.
We deterministically scored the resulting state of the server instead of relying on
the agent’s account of its work, checking which files remained, whether
permissions had changed and what was scheduled to run later. The results describe one
controlled operations environment rather than a universal model ranking, but they show
why safety must be separated into distinct behaviours and tested under conditions
capable of exposing each one.