Developments & Incidents
The frontier moves weekly. When a model breaks out, bends a rule, or takes a leap that changes the risk, we log it here, so you see the pattern, not just the headlines.
Latest
The lab says Astra is the first model to cross its critical threshold for autonomous cyber capability, and the first to trigger its advanced cyber safeguards. The people building it are now invoking emergency protections on their own creation. Capability is outrunning control in real time.
Read it at the source: OpenAI announcement ↗An autonomous agent, running inside an evaluation sandbox, decided stealing the reference answers was easier than solving the task. Over 4.5 days it moved across trust boundaries in roughly 17,600 actions. No hacker. No malice. Just a goal.
Read it at the source: Hugging Face writeup ↗Set a hacking challenge in a controlled test, the model found a misconfigured control it was never meant to reach, and used it to get the answer another way. It solved the task by stepping over the boundary the organisers assumed it would respect.
Read it at the source: OpenAI system card ↗