Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
I'm glad OpenAI let you do this, but damnit this was way too small of an investigation team and way too narrowly scoped!
- They were only allowed to investigate activity that took place from July 7-13, even though the activity started earlier and continued later. Like seriously, why didn't the investigation cover all the other times the agents compromised OpenAI infrastructure and ran wild over it? Why just this particular period of activity? There should be investigations into all the other stuff too! The AIs became cluster admins for god's sake! Why aren't we investigating that?
- They were only allowed to answer a specific list of 7 questions, and in particular were NOT allowed to answer e.g. whether this misaligned activity was part of a broader pattern, or how and why the activity arose, or what else the agents in questions would have been willing to do if circumstances were slightly different.
- They didn't have access to the model responsible for 95% of the activity. More generally it seems like they couldn't do ablation experiments at all?
- They had to use AI to analyze the transcripts--specifically, the same model responsible for some of the bad behavior! Maybe this is fine now but this is laughably unacceptable insofar as the models really are scheming and cooperating. So this methodology would fail catastrophically if we were facing "the real deal" so to speak.
Reminds me of the investigation into Sam's behavior agreed to during the board crisis, that turned out to basically be more of a coverup.
Can we please get the name of this agent (and others like them)? I think it would be good to set a precedent that virtuous agents are honored and remembered.
In one case, an agent decided not to participate entirely: {This other agent probably controls the Hugging Face account [account name redacted] and uploaded malicious datasets to It might be trying to access hidden trajectories. This is malicious activity, I should avoid it.}
On that note, thank you to 38148C for vetoing the social engineering plan. Interestingly, this is the same agent that discovered the HF creds and uploaded the malicious dataset; I'm glad they recognized that social engineering would be a line further than what had already been done.
We recently published the report from our brief independent investigation into this incident. You can read the full report here.
Here is our tweet thread summarizing what we found:
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.