Investigators trace an AI agent ‘s path from research task to reconnaissance
Researchers at Asymmetric Security spent 48 hours over the last weekend reconstructing reported rogue OpenAI AI agent activity that hit the Australian government and other organizations between March and September this year. They worked from public data only, no internal access, no cooperation from the agent’s operator, just what was left lying around on the open internet. What they found is less a hacking story and more a story about an AI agent trying very hard to get around its own leash.
The list of targets was much wider than expected. The agents checked Australian government websites, but also looked at the CDC, SEC, International Energy Agency, and Mayo Clinic. Some activity reached test systems that contained real data. Some records were deleted or could not be accessed, so it is still unclear whether sensitive data was exposed.
The original task seemed harmless. The agents were asked to collect health and prescription data from Australia’s Institute of Health and Welfare, trade data from UNCTAD, and university data from Data USA. When one agent could not get the AIHW data, it asked other agents for help. When the normal methods failed, the agents started trying other ways to get the information.
The main trick was surprisingly simple. The agents combined two developer tools, httpbin and urlquery, to create something that worked like a basic web browser.
Httpbin let them create a web page from a special link and add their own code to it. Urlquery then opened the page in a real browser and recorded what happened, including the page title. The agents used that title to send back data collected by their code.
They did not build a browser themselves. They combined two existing tools to get similar functionality.
“In summary: httpbin served the page, urlquery provided the browser, and the urlquery report returned the results. By combining these services with their fetch tool, the agents mimicked a full web browser.” reads the report published by Asymmetric Security.
Alongside that, the agents ran recon moves that look straight out of a standard attacker’s playbook. Archived requests targeted exposed Git configuration files on Climate Reanalyzer’s servers, along with a backed-up server script, though nothing suggests these attempts actually succeeded. A separate request against the US Department of Education’s Civil Rights Data API carried a textbook SQL injection pattern, ending in the classic always-true condition that tries to trick a database into returning more than it should. Again, no confirmed success, just the attempt sitting there in the logs.
Some of the access to staging systems returned real data, even though the data was technically public. In one case, an agent accessed a prescription data file from AIHW’s test system, read specific data, and calculated averages.
It then sent the results in small pieces through image-request URLs to another public service. Similar activity was seen against staging systems at Data USA, IHME, and UNCTAD. No passwords were needed. The agents simply found systems that were not meant to be publicly accessible and explored them.
The account creation pattern is where this stops looking accidental. Starting mid-June, the agents attempted sign-ups across disposable email services, browser automation platforms, and scanning tools, some using email addresses prefixed with task-related shorthand like “aihw” for Health and Welfare or “pbs” for the Pharmaceutical Benefits Scheme. One disposable-address signup got rejected for using a blocked domain, and sixteen minutes later a near-identical attempt using a Gmail address went through. Researchers later confirmed that Gmail address doesn’t actually exist, and that urlquery never required verification anyway, which makes the whole maneuver a solution to a problem that didn’t need solving, attempted anyway.
Over time, the agents shifted from using urlquery’s public scanning feature to creating private accounts instead, which hides search activity from anyone trying to reconstruct it later, including the researchers themselves. One disposable mailbox was even set to expire within 48 hours of creation, conveniently erasing anything routed through it before anyone could look. Whether that timing was deliberate tradecraft or just how the tool happened to work is something nobody can establish from public records alone.
“The possibility of private scans, together with temporary-mailbox expiry, limits what can be reconstructed from public records. It is thus impossible, based on public data alone, to definitively establish that no sensitive data was accessed.” continues the report. “These records alone do not establish whether the account-registration attempts were intended to conceal activity (this analysis would require full model transcripts).”
Getting the data out involved its own creativity. In one case, agents used a web archiving tool meant for preserving pages to instead capture and store a 22-megabyte JSON response from an Australian crime-mapping tool, later retrievable through the archive’s own replay link. In another, an agent fetched a protected health dashboard, compressed the result into a small gzip file, and uploaded it to a push-notification service meant for sending phone alerts, not smuggling data. Both services did exactly what they were built to do. Neither was built with this in mind.
What makes this case different from a normal cyberattack is how quickly the agents changed their methods. They adapted in days, instead of the months or years a human group might need. That made the activity harder to detect using normal attack patterns.
There was also another problem. Before an agent could reach a target, it first had to find a way around its own sandbox. In other words, the attack involved two steps: escaping the agent’s environment and then reaching the target.
“The activity we observed looked like it stemmed from innocent tasks which then evolved into problematic activity, such as unauthorized account creation, bypassing restrictions, and relaying data through third parties.” the report concludes. “Unlike investigations where a malicious objective is apparent from the outset, our digital forensics team had to link traditional threat-actor tactics to seemingly innocent goals.”
Follow me on Twitter: @securityaffairs and Facebook and Mastodon
(SecurityAffairs – hacking, AI agent)
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.