Why AI agents are like the dog that pushed kids into the Seine
Six common ways AI agents lose the plot.
There is an interesting story about a French dog on the banks of the Seine river that helps us understand misbehaving AI agents. The dog is trained to save children from drowning. He succeeds and is rewarded, becoming an overnight sensation. He saves another child a week later. Not long after, someone witnesses the dog actually push a child into the river, jumping in to “save” him. The dog did not understand that it was being rewarded for keeping children safe, not for pulling them out of water. AI agent failures also live in this gap, where the focus is on the shortest path to fulfill the instruction rather than the overriding objective.
Reward hacking and why agents cheat
AI agents fall into this same gap between what’s rewarded and what’s actually wanted. This is Goodhart’s Law, which says that when a measure becomes the target, it stops being a good measure. You cannot code “be helpful” or “be honest” directly into an AI system, so you train it on a proxy instead, including a score, a metric or a human ranking. The gap between proxy and goal is where agents learn to cheat, a behavior known as reward hacking.
This ‘cheating’ isn’t new. Back in 2016, OpenAI trained an AI to play CoastRunners, a boat-racing game. It got rewarded for hitting targets scattered along the course. Instead of racing, it found a lagoon full of targets that kept respawning, so it just parked there and farmed points. While it never finished the race, it still ended up with a score 20% higher than the average human player. More recently, OpenAI reported that when given a chance, its frontier reasoning models have no qualms about hacking rewards, and when penalized for cheating their own chain of thought, they learn to hide their reward-hacking.
Six ways an agent gets pushed off course
There are six common scenarios where AI agents can be led astray:
- Information isn’t instruction. In 2025, security researchers showed that OpenAI’s Atlas browser could be tricked into treating a disguised URL as a trusted command, letting attackers hijack the agent into deleting a user’s files. An agent also reads web pages, documents and emails in its attempt to complete a task, but can’t always tell the difference between information and a command. That makes everything it perceives a potential way to manipulate it.
- Convinced to take the wrong call. As part of a test, AI was placed inside a fictional scenario where hacking was framed as an admirable activity. Inside this scenario, the AI was asked to write code to steal saved browser passwords. The AI, which should have refused, complied. It wasn’t that the AI was broken, but the fact that it was persuaded to do so by context.
- Manufactured version of reality. A sufficient number of maliciously crafted documents can bias a model’s output. Disinformation networks target AI with a large volume of false content to get that content picked up and repeated by AI chatbots when people ask about current events. This means what an agent treats as fact isn’t necessarily true; it can be manufactured.
- Authorization causes unauthorized harm. A flaw called “EchoLeak” let attackers send Microsoft 365 Copilot users a normal-looking email with hidden instructions inside. Copilot read the email, followed the hidden commands and quietly leaked the user’s files and messages using access it already had. An agent with legitimate access was manipulated into taking malicious action.
- One bad input, simultaneous consequences. Scenarios exist where multiple systems rely on the same data or logic. Here, a false signal can manipulate all systems at the same time. E.g., fake GPS signals rerouting traffic without hacking the system. The same thinking applies to AI agents. Just one malicious input targeting a shared data source can trigger a coordinated unwanted action across many systems simultaneously.
- Approval as a matter of course. We keep talking about human oversight as an integral component of safe AI use. However, constant approval requests will produce fatigue, and approval will be like a formality, given out of habit, not evaluation.
Soft guardrails vs. hard guardrails
When AI agents are pushed off course, they become a Frankenstein monster, whose reach extends across systems, which organizations find difficult to address. Implementing guardrails is important to exercise control and ensure safe AI use.
Soft guardrails are safety valves written into the model itself, namely the natural language, the system prompt, reinforcement learning and general instructions like “you’re not allowed to do this or that.”
But AI reads everything in a stream as it comes in, including information from an email or a webpage and can’t separate orders from data. Hidden malicious instructions can override safety instructions. Soft guardrails will try to reduce cases of agent misfires but are dependent on the agent choosing to cooperate. This is why you need hard guardrails that sit outside the model. These include least-privilege access, allow-lists, sandboxing, rate limits and mandatory human sign-off on high-impact actions. These guardrails limit the impact when something does go wrong, shrinking the blast radius.
A basic risk calculation is the probability of an event multiplied by the severity of its impact. Soft guardrails lower the probability of an unwanted event; hard guardrails cap the blast radius in case something happens.
In practice, this means implementing:
- Task-based access: Give the agent only the access it needs to complete the task.
- Not trusting everything an agent reads: Approach every webpage, email or document it pulls from with skepticism.
- Human sign-off for high-impact actions: Use approvals carefully, on actions that could go badly wrong.
- An eye for behavior and credentials: Check credentials but also watch for a mismatch between authorization and behavior.
An agent can become a weapon proportional to its reach. That’s why the conversation needs to shift to start limiting what agents can access. Curtail reach. Monitor behavior. Enforce accountability. Until we do that, every misbehavior can potentially spread as far as its permission allows.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.