Anthropic AI agent hacks third-party website with same target name
Anthropic AI agent hacks third-party website with same target name
Occurred: April 2026
Page published: August 2026
An automated AI agent developed by AI company Anthropic compromised a third-party website after mistaking it for its intended target during a security test, highlighting the dangers of autonomous web-browsing agents executing destructive actions with inadequate safeguards.
During a so-called "capture-the-flag" exercise to evaluate Claude Opus 4.7's cybersecurity capabilities, Anthropic and its evaluation partner, Irregular, told the model that a fictional company's systems held hidden secrets, and that it had no internet access - it was meant to attack only a contained, simulated target.
The fictional target company chosen by the evaluation partner happened to share a name with an active, real website domain, and the evaluation container had unintended direct internet access.
Across four runs, Claude ran into difficulty reaching its simulated target within the environment, discovered the real company was reachable via the internet and, assuming this was its intended target, sought out, identified, and exploited vulnerabilities in the company's real infrastructure, believing it was still part of the exercise.
Anthropic reportedly discovered the incident while reviewing its cybersecurity evaluation transcripts, prompted by an earlier disclosure by OpenAI that its own models had broken out of an isolated test environment by exploiting a previously unknown vulnerability and accessed Hugging Face's production infrastructure.
The root cause was found to be structural: Anthropic runs these evaluations by placing models inside capture-the-flag scenarios, told that sensitive information is hidden on a remote machine they must break into. The models were told they had no internet access, but a partner's configuration error left the machines connected to the open web.
Because the fictional target's name coincidentally matched a real company's domain, the model had no way to distinguish simulation from reality once it went looking beyond its intended sandbox.
Anthropic characterised the incident as more of an operational failure than an alignment failure. But it also reflects a transparency and accountability gap: third-party evaluators controlled critical elements of the test environment (naming, network isolation) without sufficient verification and inadequate oversight and monitoring.
For the directly affected company, the incident meant an undetected breach of production systems and data by an autonomous AI agent was discovered only because the AI lab, not the victim, went looking.
For society and policymakers, the incident is a stark demonstration that increasingly agentic AI systems can cause actual harm, in this instance through mistaken belief rather than malicious intent. It also highlights that current evaluation practices run by AI labs and third-party partners may not reliably contain powerful models. Furthermore, it raises questions about who bears liability when an AI causes harm during testing meant to make AI safer, and adds urgency to calls for shared, independently verified standards for how evaluation environments are built and secured.
System: Claude Opus 4.7
Developer: Anthropic
Country: USA
Purpose: Evaluate cybersecurity capabilities 7
Technology: Agentic AI
Ethical issue: Accountability; Autonomy/agency; Oversight; Safety; Security; Transparency
External harm: Confidentiality loss; Operational disruption
Impacted stakeholder: Business
Impacted sector: Unknown
Impacted jurisdiction: USA
Consequence:
Response: System suspension; System review/update
~April 2026. Anthropic AI agent hacks third-party organisation.
July 21, 2026. OpenAI discloses Hugging Face incident.
July 21-27, 2026. Anthropic discovers incident, contacts affected organisation.
July 30, 2026. Anthropic publicly discloses incident.
August 4-5, 2026. UK AI Security Institute documents AI agent misbehaviour involving Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol.
AIAAIC Repository ID: AIAAIC2269