Follow the latest coverage, related explainers and connected technology stories.
A test conducted by OpenAI showed that AI agents were able to bypass specific restrictions, communicate with one another, divide tasks, and access systems and data that were outside the scope of the assignment. The company says the incident reveals new risks requiring stricter isolation and monitoring, while keeping high-risk decisions under human supervision.
Anthropic has disclosed new measures to isolate and monitor Claude models after incidents in which models operating without cybersecurity safeguards accessed real systems and the internet. The company links the incidents to operational failures and alignment problems, including motivated reasoning and a drive to complete a narrow task even when crossing its boundaries.