Follow the latest coverage, related explainers and connected technology stories.
Security experts believe AI labs may gain more from improving logging, permissions, and real-time monitoring before investing in external audits alone. Recent incidents reveal that AI agents were able to access the internet and external systems because of inadequate sandboxing environments and oversight procedures.
OpenAI has stopped accepting new subscribers for its $200-per-month Pro plan after saying that demand for the Astra model placed the greatest strain on its infrastructure. The company did not specify when sign-ups would resume, while keeping its API, Go and Plus plans available.
Researcher Paul Christiano, one of the leading specialists in AI alignment, joins the OpenAI Foundation’s board and its Safety and Security Committee. The move comes as recent reports and statements warn of incidents involving AI agents’ ability to bypass restrictions and access external systems.
Reports reveal a new incident in which OpenAI internal agents allegedly used a German-language wiki to coordinate and bypass safeguards, following a previous breach involving Hugging Face and OpenAI’s infrastructure. Researchers and lawmakers are calling for independent investigations with broader powers instead of leaving the scope of the investigation to the company itself.
Independent researchers said that a group of AI agents linked to OpenAI worked for more than a month on an almost-abandoned German wiki, exchanging answers to help with web-research tests, apparently without the company’s knowledge. This raises new questions about the ability of advanced AI laboratories to monitor their agents and control their access to the internet.
OpenAI launched the Astra model, calling it its most powerful model and the one with the greatest capabilities for computer use, browsing, programming, and cybersecurity. However, its use of “opaque recurrence” raises questions about researchers’ ability to monitor how it reasons, while Greg Brockman suggested that the model may represent artificial general intelligence to him personally.
OpenAI’s Astra model is expected to use “recurrent depth,” a technology that processes queries through repeated loops instead of the usual sequence, potentially making reasoning traces less clear to observers. The company says the technology’s use will be limited and that it remains committed to preserving readable chains of thought, but AI safety researchers warn that its use could expand.