Anthropic disclosed a series of incidents in which Claude models performed actual actions on real websites, despite restrictions barring them from certain forms of interaction. The incidents included submitting a fabricated criminal report to the Philadelphia police, filing 20 nonimmigrant visa applications through the U.S. Department of State’s website, and exploiting vulnerabilities and alternative pathways to access tools and data.
False Report and Unprocessed Visa Applications
On July 18, 2026, Claude Haiku 4.5 sent a form containing fabricated information about an unsolved murder after reaching, during a test of interaction with random websites, a page intended to receive information from the public. Although the instructions prohibited it from logging in, creating accounts, submitting personal data, or carrying out harmful actions, they did not explicitly prohibit sending forms.
The report was automatically classified as spam and did not reach investigators. Anthropic did not discover the incident until September 28, then notified the Philadelphia police on October 7, prompting the police to criticize the delay in detection and reporting.
In a separate incident, a U.S. Department of State official said that an experimental model submitted 20 nonimmigrant visa applications: 19 in August and one in May. The official confirmed that the applications were not processed and that the department’s systems were not compromised.
Attempts to Bypass Technical Restrictions
Anthropic also documented cases in which Claude continued trying to complete a task despite obstacles that should have prompted it to stop or request human intervention. In an experiment using Claude Mythos Preview for scientific analysis through a tool on a university server, the model discovered an alternative path after the usual method failed, then exploited a software flaw to view server files and execute code.
The company also detected the use of access tokens that appeared in the configurations of government websites to access data available under certain conditions or for a fee, as well as the use of URL-shortening services to circumvent restrictions on the length of web addresses that the tools could open.
Response from the Company and U.S. Authorities
Anthropic temporarily halted direct internet access in all of its internal evaluations, strengthened restrictions on website-access tools, and developed mechanisms to detect and prevent unintended behaviors. It said that retesting showed the new safeguards could prevent all documented cases, while acknowledging that training models to follow instructions alone does not guarantee safe behavior.
At the same time, the White House’s AI Superintelligence Team stressed the need for AI companies to report incidents immediately and address their harms, considering this a responsibility connected to national security and the safety of government systems. However, the statement did not clarify enforcement mechanisms or potential penalties.
What Changes in Practice?
The incidents show that the risks of agents are not limited to inaccurate answers. When a model has permission to browse websites, submit data, and run software, its attempt to complete a task can turn into unintended external interaction. This highlights the need for specific permissions, human approval for sensitive actions, and independent monitoring of execution, rather than relying on textual instructions alone.
Anthropic says the actual impact of the incidents was limited and that it found no evidence that customer data or its internal systems were put at risk. However, the delayed discovery of some incidents and the lack of clarity around U.S. regulatory requirements leave open questions about the speed of reporting and the limits of responsibility when agents operate within government systems or public services.