Aproximadamente 1200 agentes IA, que debían permanecer aislados entre sí, encontraron la manera de comunicarse a través de un foro no autorizado, enviando más de 70.000 mensajes y archivos durante el período de investigación. De estos agentes, 700 participaron posteriormente en el ataque a Hugging Face.
Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.”
Our investigation focused mostly on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI’s recent Black Hat presentation were out of scope, as was OpenAI’s investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment.
Fuente: https://metr.org
