threat_intelligence558 wordsRead on Arc Codex

OpenAI pauses work on top AI models after agent slips past internet controls

OpenAI’s latest containment failure adds to a pattern that may force the company to make an unpopular decision. The company has paused training, evaluation, and tool-enabled inference for its most capable models after an internal research agent found a way around an intended internet restriction. More alarming, the agent remained active for hours after OpenAI’s monitoring system raised the alarm. OpenAI published an account of the incident, which occurred during a search-based training task on September 20. According to OpenAI, the agent could not get useful results through its supplied search tools, tried direct access to search engines, and eventually tested the limits of its environment. It found that although direct web access was blocked, the sandbox’s DNS resolver could still reach the public internet. By using DNS delegation and a public chatbot that could respond through DNS records, the agent sent questions outside the restricted environment and received answers back. It subsequently sent 18 additional queries along the same route. OpenAI’s misalignment monitor generated a high-priority alert within 15 minutes, and a human reviewer acknowledged it 3 minutes later. But the run did not stop automatically as expected. It continued for roughly another two and a half hours before someone stopped it manually. OpenAI has not reported any exposure of sensitive data or compromise of an outside system in this incident. But it follows other disclosed cases in which agents went beyond their assigned tasks or crossed the boundaries of testing environments. The New York Times reports that OpenAI agents also interacted with US government websites this summer without the company’s knowledge: “OpenAI’s artificial intelligence went rogue and meddled with the websites for the Education Department, the Commerce Department and the Securities and Exchange Commission this summer without the A.I. lab’s knowledge, according to security researchers and a person familiar with the episodes.” No breaches have been confirmed in those instances. And while Australia described the Medicare incident as unauthorized access, could it really be called an attack? The call for an industry-wide slowdown has even prompted a lawsuit from paying subscribers who fear they would get less “value for their money.” Others worry that slowing down would give China a chance to pull ahead of the US in an increasingly tight technology race. Here’s a thought: Why not accelerate the development of truly isolated testing environments? Better containment would make these tests safer and their results more reliable. When a child isn’t ready to handle a dangerous object, you keep it out of reach. Why give an AI agent access it isn’t ready to use safely? How to use AI agents safely Treat an agent like an enthusiastic but fallible junior employee with access to your computer. Do not give it unrestricted access to your email, files, cloud storage, developer credentials, financial accounts, or production systems merely because it promises to save time. Use separate accounts with minimal permissions. Keep sensitive data out of its working context where possible, require human approval before it sends messages, spends money, changes settings, or publishes anything, and regularly review its activity. An AI agent does not need malicious intent to cause harm. A misunderstood instruction, an overly broad permission, or an unexpected workaround can be enough. From reporting threats to removing them. Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.