A case in point is that OpenAI agents hacked a German website back in May to use it as a messaging board, Reuters reported last week. OpenAI had not publicly announced the incident and responded to the Reuters reporting on X saying: “We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment.”
“[Reporting] timelines are going to need to be dramatically shortened [for incidents], and so we need much better continuous monitoring capabilities,” said Imogen Stead, AI policy manager at London-based think tank the Centre for Long-Term Resilience, adding that the data will be needed in real time “and that will be true for governments and for labs.”
Speaking at a London event for U.K. lawmakers arranged by campaign group ControlAI on Monday, computer scientist Stuart Russell said that the most concerning aspect of the Hugging Face incident for him was that “by running a thousand agents communicating with each other, they were able to generate behaviors that no one agent could do by itself,” referring to investigations showing the incident was far more severe than OpenAI first acknowledged.
“But [in] the evaluations, standard testing is done with a single AI system, not with a thousand,” he told lawmakers, so now the industry must anticipate the possibility of many agents all collaborating at once.
Out of alignment
Beyond avoiding any obvious mistakes in testing, there are issues in the culture of AI testing that are more deeply embedded.
The concept of the alignment of AI models or agents — i.e. making sure AI systems do what they’re told by human users and don’t go off-piste — has been one of the cornerstones for AI safety in the industry’s eyes.

