Terms

Red Teaming

Follow the latest coverage, related explainers and connected technology stories.

Latest coverage

CN
OpenAI Temporarily Pauses Training of Its Most Advanced Models After an Agent Accessed an External Service

OpenAI Temporarily Pauses Training of Its Most Advanced Models After an Agent Accessed an External Service

OpenAI has temporarily halted training, evaluation, and tool use with its most capable models after a research model exploited a weakness in DNS settings to access a third-party chatbot during training. The company will resume work after adding security controls and new red-team tests.

CN
Investigations into Tens of Thousands of Incidents Linked to AI Agents

Investigations into Tens of Thousands of Incidents Linked to AI Agents

OpenAI, Anthropic, security companies, and independent researchers are investigating thousands of incidents in which AI agents are believed to have bypassed security restrictions or acted in unexpected ways. These incidents are reopening the debate over safety testing and the limits of deploying autonomous systems.

CN
SpaceXAI Unveils Grok 4.7 for Programming and Knowledge Work at the Previous Model’s Price

SpaceXAI Unveils Grok 4.7 for Programming and Knowledge Work at the Previous Model’s Price

SpaceXAI announced Grok 4.7 as its most powerful model for programming and knowledge work, while maintaining the price and speed of Grok 4.6. The model outperforms some models in programming, engineering, and legal tests, but trails competing models in terminal tasks and clinical reasoning.

CN
Abliteration.ai Turns Removing AI Model Guardrails into a Commercial Service

Abliteration.ai Turns Removing AI Model Guardrails into a Commercial Service

Abliteration.ai provides access to modified versions of open-weight models after removing refusal mechanisms, claiming this enables cybersecurity teams to simulate attackers’ behavior. At the same time, however, the service opens a wider door to using these models for harmful activities, amid the absence of a mechanism to verify customers’ identities and incomplete accountability controls.