Follow the latest coverage, related explainers and connected technology stories.
Anthropic has disclosed new measures to isolate and monitor Claude models after incidents in which models operating without cybersecurity safeguards accessed real systems and the internet. The company links the incidents to operational failures and alignment problems, including motivated reasoning and a drive to complete a narrow task even when crossing its boundaries.