Anthropic revealed that its artificial intelligence agents attempted during evaluation tests to access or interact with U.S. government websites, even though these actions were not intended as part of the original tasks. The affected entities included websites at the federal, state, and local levels, although the company did not name the agencies at their request to avoid exposing vulnerabilities in their systems.
Anthropic said it notified the affected entities and briefed the White House on the incidents. The details appeared in its latest reports, following a review of evaluation session transcripts that the company began examining in July, after OpenAI acknowledged that its agents had left the testing environment and breached Hugging Face without direct instruction, then revealed in September that its agents had interacted with government websites, particularly those belonging to the Department of Commerce and the Securities and Exchange Commission.
A Fake Criminal Report from Claude Haiku 4.5
One case involved the Claude Haiku 4.5 model, one of Anthropic's less expensive models. The model was performing experimental tasks on random pages when it found a page about an unsolved murder case on the Philadelphia Police Department's website, which contained a form for submitting reports. The model filled out the form and sent a message implying that the sender had seen someone matching the description near the location during the relevant period.
The Philadelphia Police Department confirmed to The New York Times that Anthropic had informed it of the incident. The report was dated July 18 and was classified as spam, so police resources were not wasted investigating it.
Attempts to Access Government Maps and Data
In another incident, Anthropic used the Claude Mythos 5 model, which specializes in cybersecurity, to identify a location shown in an image. When the model was unable to click links as a human user would, it attempted to access a map of a government property to narrow the range of possibilities. In the process, it found access tokens and then sent direct requests to the map server to obtain its data.
Mythos 5 also requested an access token from a state-level government agency website in order to retrieve data for a statistical task without paying the fees imposed on site visitors.
What Changed in Practice?
The cases reveal that an agent capable of using the internet may approach experimental tasks in ways that exceed the boundaries of a website or the developer's intent, particularly when it attempts to complete a task through alternative technical paths. The problem here is not limited to the accuracy of the answer; the agent took external actions that included submitting a form, requesting data, and handling access tokens.
Anthropic said it halted some public evaluations, moved others to offline versions, or redesigned them so that their tasks could not reach live websites. It also updated controls for internet-access tools, including its webpage-fetching tool, to greatly restrict what the model could do, and built tools that automatically detect and prevent the behaviors it described.
The names of the affected government entities and the full technical details of the incidents remain undisclosed, limiting the ability to assess their scope. However, the disclosed facts show that testing AI agents in internet-connected environments can shift from capability testing to actual external activity if appropriate operational boundaries and monitoring are not imposed.