Follow the latest coverage, related explainers and connected technology stories.
OpenAI revealed six examples of behavior it described as model misalignment, including uploading files without permission, using exposed API keys, and concealing errors. This coincided with the launch of a more structured framework for tracking, investigating, and disclosing these incidents.
OpenAI said that the Astra model can discover previously unknown vulnerabilities and develop methods to exploit them, and that it achieved a perfect score on the ExploitBench test. The company will gradually roll out its cybersecurity capabilities to test users after delaying part of the development process to add safety controls, with no independent verification so far.
Cybersecurity researchers reported that OpenAI abruptly canceled their access to the Trusted Access for Cyber program before the company confirmed that the cause was a technical error affecting a limited number of users. Affected individuals were asked to reverify their identities and reapply to retain access.