In The News

They said they would build AI safely. Then it went rogue.

The Washington Post

August 10, 2026

CSET’s Helen Toner shared her expert insight in an article published by The Washington Post. The article looks at recent incidents in which AI models from OpenAI, Anthropic, and Meta broke out of controlled testing environments and attempted to hack real systems, raising concerns about whether AI companies can safely control increasingly capable models.

Read Article

Related Content

CSET’s Helen Toner shared her expert insight in an article published by TIME. The article examines the race to automate AI research and the possibility that AI systems could increasingly accelerate their own development, raising… Read More

CSET’s Helen Toner shared her expert insight in an interview with the Australian Broadcasting Corporation’s 7.30. The interview examines growing concerns that AI capabilities are advancing faster than the safeguards needed to keep them safe,… Read More

CSET’s Helen Toner shared her expert insight in an article published by WIRED. The article explores Anthropic’s philosophy of advancing cutting-edge AI while simultaneously positioning itself as a leader in AI safety. Read More

CSET’s Helen Toner shared her expert insight in an article published by Axios. The article examines the U.S. government’s intervention involving Anthropic’s AI models and the broader debate over how frontier AI systems should be… Read More