Researchers used Claude to hack OpenAI
We’ve heard of OpenAI’s AI agents running amok and hacking other companies. Now, a cybersecurity company has turned the tables on the ChatGPT operator by using AI to help hack OpenAI itself.
The hack, which also exposed a bug affecting dozens of other major online services, was conducted as security research. OpenAI paid the researchers for reporting a flaw in its systems through its bug bounty program.
Researchers at cybersecurity tools vendor Hacktron wrote up their adventures in mid-September. A few months earlier, they had begun looking for security flaws at companies developing frontier AI models, which are highly capable models such as those powering ChatGPT and Claude.
Using Anthropic’s Claude, the researchers went from investigating an image-processing flaw to accessing an internal OpenAI software repository in less than 72 hours. They deliberately avoided viewing sensitive information.
To get inside OpenAI, researchers Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini found two vulnerabilities and chained them together. The first wasn’t specific to OpenAI. It involved a bug in an image-upload feature in the Discourse community forum software.
This feature processes images uploaded by users and relies on a low-level software library called libheif. Uploading a specially crafted image could trigger a flaw in the library, allowing an attacker to gain control of the Discourse server.
After Hacktron used that vulnerability to compromise OpenAI’s Discourse server, the second vulnerability came into play. This was a flaw in OpenAI’s single sign-on (SSO) system, which lets people access one service using an account from another.
The SSO flaw gave Hacktron access to the ChatGPT and Codex accounts of people who had logged in to OpenAI’s Discourse forum. OpenAI uses the forum for community support, so that potentially covered a large number of accounts. Codex is an AI coding tool that helps software developers work with code.
It wasn’t just public users that had signed into this system; OpenAI employees were signed in, too, allowing Hacktron to access one employee’s account. That person’s Codex account was connected to OpenAI’s GitHub organization. GitHub is an online service that developers use to store and collaborate on software.
This gave the researchers access to OpenAI’s internal software repository.
The researchers didn’t do anything damaging with that access. They only wanted to prove that they had compromised OpenAI. So they instructed the employee’s Codex account to create a harmless pull request—a proposed software change—in an internal OpenAI repository.
OpenAI fixed its part of the problem about 14 hours after Hacktron submitted its initial report. It later paid the researchers a $6,500 bounty for the OpenAI-side flaw.
What this means for cybersecurity
Ethical hacking like this is commonplace, but there are some interesting aspects to this hack that make it different from many others.
The first is that Hacktron used AI to help in its efforts. It originally used Opus 4.8, one of Anthropic’s recent Claude models, to discover the issue in libheif. But it couldn’t use the model to build a reliable exploit that worked against the default version of Discourse.
Then Anthropic released Claude Opus 5. Using that model, the researchers were able to build a working exploit overnight.
That shows how quickly AI is moving. A task that Opus 4.8 had failed to complete across several sessions was solved by Opus 5 within hours of its release.
The second interesting aspect is that the researchers had to fool Claude into helping them do it. The model refused to write an exploit for a remote system because it considered the request unethical. So the researchers had to present the task as a capture-the-flag exercise (a hacking competition) to get it to play ball.
When they did that, the agent took over the test forum server within four hours. The researchers then used the resulting exploit against OpenAI’s forum. It shows that while companies may do their best to place ethical constraints on the use of their AI, a wily researcher can still get around it with some simple prompt engineering.
And it didn’t cost much to do this. Hacktron spent less than $3,000 in AI tokens during its two-month research project, which involved three researchers and uncovered vulnerabilities affecting several major companies.
The image-processing bug became the basis of a wider project called HEIF Heist. Hacktron found related security weaknesses affecting services and software from companies including Slack, Meta, and GitHub. Tweaking the HEIF exploit to target a new company took a day or two on average.
All of this lowers the bar for sophisticated hacking even further. People have been able to use ready-made hacking tools for years without really understanding how they work, but it took real expertise to pick through software, find hidden flaws, and turn them into reliable exploits.
Hacktron says skilled human guidance was still important and that this was not completely autonomous hacking. Even so, AI is making some of that expertise cheaper and faster to apply—and it’s only getting better.
Stop threats before they can do any harm.
Malwarebytes Browser Guard blocks phishing pages and malicious sites automatically. Free, one click to install. Add it to your browser →
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.