Skip to main content
AI Inside Organizations

How Metrics Destroy Good Judgment in Organizations

You stopped thinking when the dashboard turned green.

Why do metrics destroy good judgment? Organizations that optimize for measurable outcomes lose the ability to recognize when the measurements themselves are wrong.

How Metrics Destroy Good Judgment in Organizations

The dashboard is green and the customer is angry.

Support response time improved. Ticket closure rate improved. Escalations are down. The weekly report says the support operation is healthier than last quarter.

Then account managers start hearing the same complaints from customers whose tickets were technically closed.

The metric did not lie. It measured what it was told to measure. Agents responded faster, closed more tickets, and escalated fewer cases. The judgment failure happened earlier, when the organization decided those numbers could stand in for whether customers were actually helped.

Metrics ruin judgment when people stop treating them as evidence and start treating them as reality, a pattern that becomes sharper when algorithms optimize the metric.

Key Takeaways

  • The substitution happens quietly. “Are customers satisfied” turns into “what is our NPS” and once that swap is complete, the metric is evaluating judgment instead of the other way around.
  • Goodhart’s Law isn’t trivia it’s a live mechanism. Reward 80% test coverage and engineers write shallow tests to hit the number; the dashboard improves while confidence in the system doesn’t.
  • Metrics attached to promotions, bonuses, or rankings corrupt faster. By the third quarter of tracking a number tied to incentives, it may show compliance theater more than real behavior.
  • The McNamara fallacy: available numbers win over important ones. Revenue is easy to defend; “the organization became more brittle” is not so the unmeasured reality waits and returns later as churn, incidents, or attrition.
  • Good judgment treats metrics as suspicious evidence, not conclusions. If response time improves but complaints rise, the metric is only telling part of the story the disagreement between signals is the real finding.

The Substitution Happens Quietly

A metric usually enters as a helper.

Customer satisfaction is hard to inspect, so the company tracks NPS. Code quality is hard to summarize, so engineering tracks test coverage. Productivity is hard to compare, so teams track story points. Hiring quality is hard to know early, so recruiting tracks time-to-fill.

At first, the number sits beside judgment.

Then reporting pressure arrives. Numbers travel upward better than context. They fit slides, quarterly reviews, scorecards, and executive comparisons. The metric becomes easier to discuss than the underlying thing.

The question changes.

Are customers satisfied becomes what is our NPS. Is the code maintainable becomes are we above coverage target. Are we making useful progress becomes how many points closed.

Once that substitution is complete, judgment is no longer evaluating the metric. The metric is evaluating judgment.

Goodhart’s Law Is Not a Trivia Item

“When a measure becomes a target, it ceases to be a good measure.”

Organizations quote the line and then behave as if their version is immune.

A team is told to reach 80 percent test coverage. Coverage rises. Engineers write shallow tests, exclude complex files, and protect the number during refactors. The dashboard improves while confidence in the system does not.

A sales team is rewarded for new meetings booked. Meetings rise. Qualification falls. Account executives spend more time with prospects who were never going to buy.

A product team is measured on monthly active users. Notifications increase. Re-engagement improves. User trust erodes slowly enough to miss the first few reports.

The metric changes behavior because it changes consequences, especially once optimization meets human stakes. That is the point of setting a target. The damage begins when the organization forgets that behavior can satisfy the metric while betraying the goal.

Campbell’s Law Shows Up in Performance Reviews

Metrics corrupt faster when they are attached to social decisions.

Promotion. Bonus. Budget. Ranking. Headcount. Public praise. Executive attention.

Once a number affects those outcomes, people learn to manage the number. They may still care about the underlying work, but the number now carries a price.

This is why performance metrics become less informative over time. The first quarter shows behavior. The next quarter shows behavior plus adaptation. By the third quarter, it may show compliance theater.

A team that knows velocity matters estimates differently. A manager who knows engagement scores matter schedules listening sessions before the survey. A department that knows cost reduction matters delays necessary spending until after the reporting period.

The metric starts as measurement and becomes negotiation.

The McNamara Pattern

Organizations measure what is available, then slowly promote availability into importance.

Revenue is available. Brand trust is vague. Feature count is available. Product coherence is harder. Meeting attendance is available. Decision quality is not. Lines of code are available. System comprehension is not.

The available number wins because it can be defended.

A leader can say revenue increased. It is harder to say the organization became more brittle. A product manager can say ten features shipped. It is harder to say the user experience now feels incoherent. An engineering manager can say throughput improved. It is harder to say the team no longer understands the architecture.

The unmeasured reality does not vanish. It waits.

It returns as churn, rework, incidents, attrition, customer distrust, and strategic drift.

Metrics Shorten Time

Short-term effects are easier to measure than long-term consequences.

A pricing change boosts quarterly revenue. The dashboard shows success before customer trust has time to degrade. A growth feature increases activation before low-quality users create support load. A shortcut accelerates launch before technical debt slows every future release.

Metrics pull attention toward the period they can see.

Teams become rationally impatient. They choose the action that improves the reporting window and leave the cost outside the frame. By the time the long-term cost enters the dashboard, the people who made the decision may have moved roles, been promoted, or gained enough narrative distance to call the damage a new problem.

This is how a green quarter becomes a red year.

Quantification Wins Arguments

Numbers feel cleaner than judgment.

A manager saying “this customer segment is losing trust” sounds subjective. A chart saying retention dropped 3.2 percent sounds real. A senior engineer saying “the architecture is becoming hard to reason about” sounds interpretive. A chart saying deployment frequency improved sounds objective.

In hierarchy, objectivity travels.

That gives metrics political advantage even when they are partial or misleading. The person with judgment must explain context. The person with the metric can point.

Over time, people learn to bring numbers to arguments that numbers cannot fully answer. The organization becomes more numerate and less perceptive.

The Illegible Work Gets Starved

Important work often leaves weak traces.

Preventing a failure. Mentoring someone before they break something. Simplifying an architecture. Repairing trust with a customer. Slowing a bad decision. Noticing that two metrics are pointing in opposite directions. Keeping a team calm during ambiguity.

These things matter because they shape future capacity. They are hard to measure because their output is often absence: no incident, no escalation, no resignation, no customer blowup, no bad launch.

Metrics privilege work that produces visible units.

When organizations reward only visible units, people reduce investment in the work that prevents future damage.

Then the future damage arrives and gets measured as a new problem.

Metrics Become Accountability Theater

A scorecard can make accountability look precise while responsibility remains vague.

The number missed target. The team is red. The owner needs an action plan.

That may be useful if the owner controls the system producing the number. It becomes theater when the owner controls only a small part of it.

A support manager owns satisfaction scores but does not control product defects. A product manager owns adoption but does not control sales promises. An engineering lead owns delivery but does not control headcount, scope, or executive commitments.

The metric identifies someone to question. It does not necessarily identify the cause.

Accountability becomes the act of assigning a person to a number.

Precision Hides Uncertainty

A metric with decimal places can make a weak model look settled, even though Google’s ML glossary notes that high accuracy can still mean no useful predictive power in the wrong setting.

Employee engagement is 7.4. Customer sentiment is down 2.1 percent. Productivity improved 6 percent. Model confidence is 0.82.

The number may depend on sampling, survey design, missing data, changed definitions, seasonal effects, instrumentation bugs, or behavior created by the metric itself.

Precision makes uncertainty harder to admit.

People argue over whether the number moved enough instead of whether the number should govern the decision. The conversation becomes technically precise and strategically lazy.

Good Judgment Uses Metrics Against Themselves

Metrics are useful when they are treated as suspicious evidence.

A good operator asks what the number excludes, who might be gaming it, what changed in the collection process, whether the movement matches lived experience, and what would be true if the metric were improving while the outcome was worsening.

Good judgment looks for disagreement between signals.

If response time improves and customer complaints rise, the metric is telling only part of the story. If velocity rises and incidents rise, delivery may be borrowing from reliability. If engagement rises and retention falls, attention may be turning into fatigue.

The metric is not the conclusion. It is a prompt for investigation.

What to Measure Without Losing Judgment

Measure the thing, the proxy, and the damage the proxy can cause.

If support speed matters, also measure repeat contact and customer resolution quality; if sentiment matters, remember that sentiment analysis compresses feedback before it clarifies anything. If feature velocity matters, also track defects, operational load, and maintainability. If engagement matters, also track retention, user trust, and unwanted behavior encouraged by the engagement loop.

Keep some decision spaces qualitative on purpose.

Ask people close to the work what the metrics are missing. Treat anecdotes as early-warning sensors, not statistical proof. Rotate metrics before they become games. Retire metrics that have turned into theater. Protect leaders who override a green dashboard because the underlying system is unhealthy.

The humility is the hard part.

A metric can improve while judgment should say stop. Organizations that cannot do that are not data-driven. They are metric-driven, and the dashboard is driving without looking out the window.

Frequently Asked Questions

What is Goodhart’s Law and why does it matter for business metrics? Goodhart’s Law states that when a measure becomes a target, it stops being a good measure. Once a team knows a specific number is being tracked and rewarded, they optimize for that number directly like writing shallow tests to hit a coverage target which can improve the dashboard while the actual underlying goal doesn’t improve or even gets worse.

Why do metrics tied to bonuses or promotions become less reliable over time? This is sometimes called Campbell’s Law. Once a number affects a person’s compensation or career, people learn to manage the number itself rather than just the underlying work. The first period of tracking shows real behavior; later periods increasingly show adaptation and compliance theater instead.

What is the McNamara fallacy? It’s the tendency to treat easily measured things as automatically important, and to discount what resists measurement. Revenue is easy to report and defend; organizational brittleness or eroding trust is not so leaders lean on the available number even when it doesn’t capture what actually matters.

How can a team use metrics without losing good judgment? Track the proxy alongside the outcome it’s supposed to represent and the damage it could cause for example, response speed alongside repeat contact rate and resolution quality. Treat metrics as evidence to interrogate rather than conclusions, and pay attention when two signals disagree, since that disagreement is often where the real problem shows up.

Internal

External