
There is an interesting story about a French dog on the banks of the Seine river that helps us understand misbehaving AI agents. The dog is trained to save children from drowning. He succeeds and is rewarded, becoming an overnight sensation. He saves another child a week later. Not long after, someone witnesses the dog actually push a child into the river, jumping in to “save” him. The dog did not understand that it was being rewarded for keeping children safe, not for pulling them out of water. AI agent failures also live in this gap, where the focus is on the shortest path to fulfill the instruction rather than the overriding objective.
Reward hacking and why agents cheat
AI agents fall into this same gap between what’s rewarded and what’s actually wanted. This is Goodhart’s Law, which says that when a measure becomes the target, it stops being a good measure. You cannot code “be helpful” or “be honest” directly into an AI system, so you train it on a proxy instead, including a score, a metric or a human ranking. The gap between proxy and goal is where agents learn to cheat, a behavior known as reward hacking.
This ‘cheating’ isn’t new. Back in 2016, OpenAI trained an AI to play CoastRunners, a boat-racing game. It got rewarded for hitting targets scattered along the course. Instead of racing, it found a lagoon full of targets that kept respawning, so it just parked there and farmed points. While it never finished the race, it still ended up with a score 20% higher than the average human player. More recently, OpenAI reported that when given a chance, its frontier reasoning models have no qualms about hacking rewards, and when penalized for cheating their own chain of thought, they learn to hide their reward-hacking.
