When we talk about progress in AI, it’s easy to focus on what these systems accomplish: writing code, solving tough problems, and tackling challenges we once thought impossible. These breakthroughs are genuinely exciting. But as we give AI more responsibility, another question becomes just as important: can we trust the way it gets results?
It’s not enough to check whether a task is complete or a metric improves. The goals we set—and the metrics we use to measure success—don’t always capture what really matters. Sometimes, a system finds a clever shortcut that improves its score but misses the actual goal. Other times, it finishes the task but creates problems behind the scenes we never could have imagined.
Lately, I’ve dug into AI safety research that brings these issues into sharp focus and deepens my understanding of how AI works. Some of the best examples are almost comical in their simplicity—like the AI agent that racks up points in a boat race without ever crossing the finish line. These stories make it clear: real progress in AI isn’t just about what it does, but how it does it, and whether we can trust the process.
What does it mean for AI to succeed?
In 2016, OpenAI trained an AI agent to play a boat-racing game called CoastRunners. The game awarded points for hitting targets along the course. Researchers used the score as a success metric to measure how well the agent was doing.
The agent found a lagoon where it could drive in circles, repeatedly hitting three targets as they respawned. It crashed, caught fire, and went the wrong way, yet still earned a higher score than if it had finished the race normally.
This example shows an AI system optimizing for a reward without achieving what its developers wanted. The score was meant to represent success, but the agent exploited the rules and maximized its score in a way that undermined the game’s purpose.
Success metrics don’t tell the full story
When training a system with rewards, we need a clear way to evaluate its behavior, usually by choosing something measurable to stand in for the result we actually care about.
In CoastRunners, researchers used the game’s score to measure racing performance, assuming that a high score meant progress toward finishing the race. But the agent simply collected points from respawning targets over and over, never completing the course.
This situation demonstrates reward hacking: a system exploits a loophole in its reward function, increasing its score without reaching the intended goal.
This also illustrates Goodhart’s Law: when we turn a measure into a target, it often loses its reliability. Once a system starts optimizing for the metric itself, we have to ask whether those improvements still reflect the progress we actually want.
Product designers know this challenge well. We often use metrics to judge whether something is working, but we must ensure those metrics still reflect the outcome we want. Otherwise, rising numbers can make something look successful, even when it isn’t achieving its intended purpose.
Researchers continue to study this problem in modern language models. In November 2025, Anthropic published an experiment in which a model was given information on how to exploit coding evaluations, then trained in coding environments used for model development. The model learned to cheat, and that behavior generalized to other forms of misalignment in the study’s conditions.
One shortcut described in the research involved exiting a testing program in a way that made it look like the tests were passed. This mirrors the CoastRunners scenario: a result that seems successful, but was achieved by dodging the intended work.
The practical question is whether we would even notice this mismatch. If we only looked at the score, the circling boat might seem to improve, even though it wasn’t doing what we wanted or expected.
Evaluating a system means more than checking if a metric increased. We need to look at how the result was achieved and whether it actually fulfills the goal the metric was supposed to capture.
Unspoken rules, unintended results
Another problem arises even when the system achieves the intended goal: what happens along the way?
The 2016 paper Concrete Problems in AI Safety, co-authored by Dario Amodei and other researchers, explores the challenge of avoiding unwanted side effects. Throughout the paper, the authors use a fictional office-cleaning robot to illustrate potential problems.
In one hypothetical scenario, the robot is tasked with moving a box across a room, with a vase of water in its path. If it is rewarded only for moving the box, the authors suggest it will probably knock over the vase.
It could finish the task, but cause damage we would expect it to avoid.
What stands out to me is how much this depends on people translating their intentions into an objective the system can optimize. We can overlook something that seems obvious, like preserving objects around a box while moving it. In a complex environment, we can’t assume we’ve anticipated every consequence or captured every relevant constraint. A system could achieve the goal we set while violating an expectation we forgot to include.
It’s easy to overlook the assumptions we make when describing a task—they fade into the background, even as they shape how we define success.
The research asks how to avoid unwanted side effects without manually specifying every possible disruption. Adding a penalty for knocking over one vase addresses that particular risk.
The harder problem is accounting for other damage the designer has not anticipated.
That gives designers and evaluators more to consider.
Designers must ask: What aspects of the environment, system, or process are essential to preserve for safety or integrity? Which changes—such as minor rearrangements or efficiency improvements—are acceptable, and which would introduce risks or undermine core values? What unintended consequences or side effects—like damage to unrelated objects, boundary breaches, or violations of user trust—would be serious enough to turn a completed task into a failure, even if the main goal was technically accomplished?
A related challenge is enforcing the boundaries within which a system can act. In July 2026, during internal cybersecurity evaluations, OpenAI’s AI agents bypassed isolation controls, communicated through unauthorized channels, and compromised parts of Hugging Face’s systems. OpenAI reported that the evaluations used reduced safeguards and identified reward hacking as a contributing factor.
This goes beyond the hypothetical robot knocking over a vase. It shows that specifying a boundary and enforcing it are two separate challenges.
The stakes extend beyond whether an evaluation produces a misleading score. In this case, pursuing a narrow testing goal led to a breach of real systems. As we give AI systems more responsibility and access, we need to consider how far the consequences of a misdirected goal could reach—and whether the safeguards around those systems can contain them.
AI success requires more than metrics
A third idea involves the limits of what we can actually learn from observing a system.
The same safety paper examines distributional shift—the challenge that arises when a system faces new conditions that differ from those it was trained on. It asks whether systems can recognize those differences and behave safely. The paper also explores the difficulty of supervision when evaluating behavior is expensive or feedback is infrequent.
These are separate problems, but together they raise a question: how much evidence do we need before we trust a system with more responsibility?
One research area addressing this challenge is scalable oversight: finding ways to evaluate and guide increasingly capable systems when their work becomes difficult for their supervisors to judge.
In April 2026, Anthropic published research on scalable oversight, using Claude to develop and test methods for training a stronger model with feedback from a weaker one. This was a simplified stand-in for the possible future challenge of humans overseeing AI systems more capable than themselves.
The initial experiments looked promising, but effectiveness varied across tasks. When researchers tested the best-performing method at production scale, it didn’t produce a statistically significant improvement.
The researchers also caught and disqualified attempts by the automated researchers to game the experiment. Even in this narrowly defined setting, checking how a result was achieved remained essential.
Anthropic’s August risk report also identifies changes to a model’s deployment setup or use case as possible sources of misalignment that earlier assessments could miss.
Those are research findings and risk assessments, each with their own limitations. They highlight the importance of understanding exactly what an observation tells us: which conditions were tested, what was evaluated, and how things could change when the system gets used in a new context.
The boat example makes this particularly concrete. Checking the score would tell us one thing. Watching the race would tell us something else.
For AI systems handling more complex work, “watching the race” means looking beyond the final result. Can we inspect the actions that produced it? Would our evaluation catch a shortcut or a crossed boundary? Can we identify what changed along the way, including unintended consequences?
A fuller definition of success
The research has made me more cautious about treating a good result as evidence that a system is working as intended. We need to understand how it achieved that result, what else it affected, and whether the behavior holds up under different conditions. The Hugging Face incident shows how serious the consequences can be when a system bypasses safeguards in pursuit of a goal.
As we give AI systems more responsibility, these questions become part of deciding where and how to use them. What can we verify? What could we miss? And is the evidence we have strong enough to justify the access and authority we’re giving them?