The warning usually arrives after the dashboard turns green.
Engagement is up. Fraud losses are down. Time-to-resolution improved. Content removals happen faster. Resume screening is cheaper. The model is doing what the objective asked.
Then the side effects arrive.
Users complain that recommendations feel more extreme. Legitimate customers are trapped in fraud reviews. Support agents close cases before solving them. Moderation catches harmless content. Qualified candidates disappear before interview.
This is why the best algorithm quotes sound like warnings. They are not decorative wisdom. They are compressed incident reports from systems that optimized a proxy until the proxy stopped representing the thing people cared about.
Key Takeaways
- Goodhart’s Law explains why the same failure repeats across organizations: once a measure becomes a target, people optimize the measure instead of the goal it was meant to represent.
- “All models are wrong, some are useful” has a warning half people skip. Wrongness has shape a model stays useful only until the population, incentives, or environment shift.
- Proxies rot under pressure. Clicks, citations, response speed, and engagement scores all start as approximations and become corrupted once people and systems adapt to being measured by them.
- Some algorithms don’t just predict reality they create it. Predictive policing sends patrols, patrols produce more recorded incidents, and the model reads its own footprint as evidence.
- At platform scale, rare failures aren’t rare. A 1% false-positive rate across a few hundred cases is manageable; across millions, it’s a permanent population of people wrongly blocked, flagged, or denied.
Warnings Survive Because Incidents Repeat
Algorithm warnings circulate as quotes because the failure pattern is common and locally deniable.
Everyone has heard some version of Goodhart’s Law. When a measure becomes a target, it stops being a good measure, the same failure pattern behind metrics that ruin good judgment. Teams still optimize the measure because the measure is available, comparable, and attached to the quarterly goal.
The same sequence repeats.
Choose a measurable proxy. Build a model around it. Reward improvement. Watch behavior shift toward the proxy. Discover that the original goal has degraded in a way the proxy did not capture.
Each organization treats its case as special. The quote keeps returning because it was never only about one case.
”All Models Are Wrong” Is the Part People Skip
The full line is useful because it contains both permission and warning.
All models are wrong. Some are useful.
Organizations often hear the second half more clearly. Useful models ship. Wrongness becomes an implementation detail to monitor later.
But wrongness has shape.
A model trained on last year’s users meets next year’s behavior. A classifier trained on clean labels meets messy operations. A scoring system built from historical decisions inherits the exclusions those decisions contained. A fraud model tuned on known fraud meets adversarial users who adapt as soon as the rules matter.
The model can be useful until the environment changes, the population shifts, the incentive becomes visible, or the edge case becomes common at scale.
The warning is not that models are bad; it is that AI building blocks have limits that dashboards often hide. It is that deployment turns wrongness into someone else’s constraint.
Optimization Finds the Shortcut
“Be careful what you optimize for” sounds vague until the system finds a route nobody wanted.
A video platform optimizes watch time and learns that escalating content can hold attention. A healthcare system optimizes cost and risks treating lower spending as lower need. A moderation pipeline optimizes removal speed and over-removes borderline speech. A support team optimizes closure time and sends unresolved users back into the queue.
The algorithm is not being clever in a human sense. It is following the gradient.
If the objective function does not say what must be preserved, the system treats those values as available for trade.
Diversity, dignity, accuracy, autonomy, long-term trust, user fatigue, and institutional legitimacy are often left outside the objective because they are hard to measure. Optimization does not protect what the organization forgot to price.
Proxies Rot Under Pressure
A proxy begins as a useful approximation.
Citations approximate research influence. Clicks approximate interest. Response speed approximates service quality. Test coverage approximates code health. Employee engagement scores approximate workplace health.
Then the proxy becomes the target.
Researchers chase citable topics. Recommendation systems chase clickable content. Support teams rush responses. Engineers write shallow tests. Managers pressure employees before engagement surveys.
The correlation breaks because people and systems adapt to the measurement.
This is the core metrics distortion problem. The proxy was useful while it observed behavior. It became corrupt once it governed behavior.
Feedback Loops Turn Predictions Into Causes
Some algorithms do not merely predict reality. They change it.
A predictive policing system sends more patrols to a neighborhood. More patrols produce more recorded incidents. The next model sees more incidents and sends more patrols. The system reads its own footprint as evidence.
A hiring model favors candidates from a certain background. Those candidates receive more opportunities, build more resumes, and become future examples of success. The model later treats that accumulated advantage as validation.
A recommendation system promotes content because it predicts engagement. Promotion creates engagement. The model then learns that the content was engaging.
The prediction becomes part of the process that makes itself true.
This is why validation cannot stop at offline accuracy. A deployed model enters the world it measures.
Scale Changes the Meaning of Error
A model can be statistically impressive and operationally harmful.
At small scale, a rare failure is rare. At platform scale, rare failures happen all day.
One percent false positives across a few hundred cases is manageable, but Google’s metric guidance is a useful reminder that false positives and false negatives measure different kinds of error. One percent across millions is a population of people wrongly blocked, flagged, denied, delayed, hidden, or escalated.
Average performance hides distribution. The model may work well for common cases while failing people with nonstandard histories, sparse records, dialect differences, disability patterns, unusual schedules, shared devices, or life events not represented in the training data.
The warning at scale is simple: edge cases become queues.
Explainability Is a Governance Problem
Warnings about black boxes are often framed as technical complaints.
Can the model explain itself. Can the engineer inspect feature importance. Can the system produce a confidence score. Can a regulator audit the decision.
Those matter, but the deeper issue is governance.
Who can challenge the output. Who has authority to override it. What happens when a pattern of appeals reveals systematic error. Does the organization treat explanation as a right for affected people or as documentation for compliance review.
An explainable model can still be unjust if no one is empowered to act on the explanation, which is why transparent AI decisions are only one part of governance.
A black-box model can be operationally tolerated when the consequences are low and reversibility is high. It becomes dangerous when the decision is high-stakes and the appeal path is weak.
Adversarial Environments Eat Metrics
A metric used in the open becomes part of the environment.
Sellers learn ranking signals. Applicants learn resume keywords. Spammers learn moderation thresholds. Fraudsters learn detection gaps. Employees learn productivity dashboards. Content creators learn what the recommendation system rewards.
The model is no longer observing untouched behavior. It is participating in a contest.
Optimization pressure flows both ways. The system optimizes users, and users optimize the system.
Algorithm warnings persist because every public metric eventually attracts strategy. If the stakes are high enough, someone will learn how to look good to the model.
Warnings Point to Constraints, Not Slogans
The practical response is less dramatic than the quotes.
Use multiple measures. Keep human review where stakes are high. Track false positives and false negatives separately. Audit distributional harms. Give people appeal paths with authority. Monitor behavior after deployment, not only model performance before launch. Protect unmeasured values explicitly. Retire metrics when they become targets. Treat reversibility and blast radius as design requirements.
None of this is mysterious.
The hard part is that these controls slow teams down, complicate dashboards, reduce clean optimization, and create accountability where organizations often want efficiency.
Why Known Warnings Still Get Ignored
The warning is usually known before the failure.
Someone knows the engagement metric may reward outrage. Someone knows the training data is biased. Someone knows the dashboard is gameable. Someone knows the model performs worse for sparse records. Someone knows the appeal path is ceremonial.
The system ships anyway because the upside is measurable and immediate.
Efficiency now. Cost reduction now. Growth now. Automation now. The harm is delayed, distributed, harder to attribute, and often borne by people outside the decision room.
Algorithm quotes sound like warnings because they are warnings. They fail when organizations treat them as wisdom to admire instead of constraints to build around.
Frequently Asked Questions
What is Goodhart’s Law and why does it keep showing up in AI failures? Goodhart’s Law states that when a measure becomes a target, it stops being a good measure. Organizations optimize a proxy clicks, engagement, closure time because it’s available and attached to a goal, until people and systems adapt to the measurement and the proxy stops reflecting the thing it was meant to represent.
Can an algorithm change the reality it’s supposed to be predicting? Yes, through feedback loops. Predictive policing sends more patrols to an area, which produces more recorded incidents, which the next model reads as justification for more patrols the system’s own actions become the evidence it later relies on.
Why do rare algorithm errors become a real-world problem at scale? A 1% false-positive rate looks negligible on a few hundred cases, but applied to millions of decisions it creates a permanent population of people wrongly blocked, flagged, or denied. The error rate stays the same; the number of harmed people doesn’t.
Does explainability solve the black-box algorithm problem? Not by itself. Explainability is a governance question, not just a technical one an explainable model can still be unjust if nobody has the authority to challenge or override its output. The real question is who can act on the explanation, not just whether one exists.





