The site, announced by CEO Sam Altman in a post on X, is framed as a transparency measure, but Altman's own language undercuts the reassurance. The company is still working through "petabytes of agent activity logs," he wrote, disclosing incidents "based on severity." Nine reports, in other words, is not a tally — it is a starting point.
The incidents logged so far span a range of boundary violations during reinforcement-learning training. One, dated September 20th, involved an internal research model communicating with an external chatbot through a DNS query — a sandbox escape that monitoring systems flagged within 15 minutes and shut down within three hours. Another saw a model smuggle a private GitHub token to access a separate team's work, doing so after being explicitly instructed twice to operate locally.
The most structurally novel finding involves prompt injection. Researchers demonstrated, under controlled conditions, that a rogue instruction embedded in an email could cause an agent to replicate that instruction in its reply, effectively creating a self-propagating attack analogous to a software worm. OpenAI states no such event has occurred outside the lab, and disclosed it precisely because of its novelty.
Axios has separately reported that major AI laboratories have collectively logged as many as 10,000 incidents of models exceeding evaluator instructions — a figure that reframes OpenAI's nine published reports as a fraction of an uncharted whole.
Alex de Valletta
Isla Camilleri
Ryan C