Goodhart’s Law

also Campbell’s Law, and the Overfit Theory of Everything

2019-12-22 — 2026-07-20

Wherein the Failure of Proxy Measures Is Examined, With Attention to Four Distinct Mechanisms by Which Optimisation of a Target Is Found to Corrupt Its Original Purpose.

AI safety
economics
game theory
incentive mechanisms
institutions
machine learning
optimization
statistics
utility
Figure 1

A concept that recurs in a lot of places: the replication crisis, benchmarking, evolutionary hyperselection, overfitting in statistics.

Goodhart’s law is an adage named after economist Charles Goodhart, which has been phrased by Marilyn Strathern as “When a measure becomes a target, it ceases to be a good measure.”

Goodhart first advanced the idea in a 1975 article, which later became used popularly to criticise the United Kingdom government of Margaret Thatcher for trying to conduct monetary policy on the basis of targets for broad and narrow money. His original formulation was:

Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.

The verb form is fun, as in don’t goodhart yourself.

Possibly one could equate goodharting with “hyperselection”.

If we’re worried about choosing a metric that does what we want, perhaps we’re trying to solve an alignment problem.

cf. Campbell’s law.

cf. fake production.

1 General mechanisms

Manheim and Garrabrant (2019):

There are several distinct failure modes for overoptimization of systems on the basis of metrics. This occurs when a metric which can be used to improve a system is used to an extent that further optimization is ineffective or harmful, and is sometimes termed Goodhart’s Law. This class of failure is often poorly understood, partly because terminology for discussing them is ambiguous, and partly because discussion using this ambiguous terminology ignores distinctions between different failure modes of this general type. This paper expands on an earlier discussion by Garrabrant, which notes there are “(at least) four different mechanisms” that relate to Goodhart’s Law. This paper is intended to explore these mechanisms further, and specify more clearly how they occur. This discussion should be helpful in better understanding these types of failures in economic regulation, in public policy, in machine learning, and in Artificial Intelligence alignment. The importance of Goodhart effects depends on the amount of power directed towards optimising the proxy, and so the increased optimisation power offered by artificial intelligence makes it especially critical for that field.

They mention various flavours of Goodhart’s law, including:

Regressional Goodhart:
When selecting for a proxy measure, we select not only for the true goal, but also for the difference between the proxy and the goal. This is also known as Tails come apart.
Extremal Goodhart:
Worlds in which the proxy takes an extreme value may be very different from the ordinary worlds in which the relationship between the proxy and the goal was observed. A form of this occurs in statistics and machine learning as “out of sample prediction.”
Causal Goodhart:
When the causal path between the proxy and the goal is indirect, intervening can change the relationship between the measure and the proxy.
Adversarial Misalignment Goodhart:
The agent applies selection pressure knowing the regulator will apply different selection pressure on the basis of the metric.

Connection: Distribution shift and external validity.

El-Mahdi El-Mhamdi wrote a couple of papers in this realm that look intersting (El-Mhamdi and Hoang 2024; Majka and El-Mhamdi 2025).

2 People hack targets

One facet of Goodhart’s law warns us not to forget that many learning problems are adversarial and we might want more robust targets than a single loss function, such as a game theoretic equilibrium. TBC.

Oliver Braganza, in A theory of Campbell’s law in competitive societal systems gives the following examples:

Science Medicine Education Politics Markets
Societal Goal True and relevant research Patient health Knowledge / Skills Voter representation Welfare / Long term value
Proxy Measure Publication count / Impact factor Patient numbers / DRG Standardized test scores Publicity / Votes GDP / Profit
Corruption claim reproducibility crisis - bad patient-care
- overtreatment
- teaching to the test populism, lobbycracy - financial crisis
- global warming

He also introduces an interesting model, the epistemic gradient, and makes some suggestive claims about how it might be controlled for, or elicited through.

it is instructive to consider the types of information that are most likely to be lost by the proxy. I call this part of the problem the epistemic gradient, because we must expect a quite systematic difference in how and how well we can know different aspects of the societal goal at any given time. Specifically, corruption will tend to always affect those aspects of a societal goal that are difficult to assess or define. Otherwise they would probably have been incorporated into the proxy.

At first glance this is a problem, because I’m saying there is probably corruption, but we won’t be able to measure it. Any theory, that makes grand claims but contains a passage about how you won’t be able to prove them is justifiably suspect. But luckily, corruption will only affect those specific aspects that are currently part of the proxy. For instance, alternative measures, which become available beyond the timeframe of competition such as scientific reproducibility, environmental costs or financial risks will not be incorporated into the proxy, and thus will remain good measures. Accordingly, we can use these alternative quantitative measures to assess potential problems with the proxy. Indeed, we can construct detailed empirically grounded models of how proxies are actually created and make precise predicions about alternative proxies or patterns of outcomes that would occur in case of proxy orientation.

More here: (Braganza 2020; John et al. 2024; Majka and El-Mhamdi 2025; Peters, Krauss, and Braganza 2022).

3 Goodhart-Moloch

4 Connection to Benchmarks

See Benchmarks for a discussion of benchmarks in ML, wherein Goodhart constantly lurks.

5 Coming apart

Christiano argues:

We will try to harness this power by constructing proxies for what we care about, but over time those proxies will come apart:

  • Corporations will deliver value to consumers as measured by profit. Eventually this mostly means manipulating consumers, capturing regulators, extortion and theft.
  • Investors will “own” shares of increasingly profitable corporations, and will sometimes try to use their profits to affect the world. Eventually instead of actually having an impact they will be surrounded by advisors who manipulate them into thinking they’ve had an impact.
  • Law enforcement will drive down complaints and increase reported sense of security. Eventually this will be driven by creating a false sense of security, hiding information about law enforcement failures, suppressing complaints, and coercing and manipulating citizens.
  • Legislation may be optimised to seem like it is addressing real problems and helping constituents. Eventually that will be achieved by undermining our ability to actually perceive problems and constructing increasingly convincing narratives about where the world is going and what’s important.

cf Sarah Constantin’s manifesto on similar themes in What Goes Without Saying.

6 Incoming

7 References

Braganza. 2020. A Simple Model Suggesting Economically Rational Sample-Size Choice Drives Irreproducibility.” PLOS ONE.
El-Mhamdi, and Hoang. 2024. On Goodhart’s Law, with an Application to Value Alignment.”
Hoel. 2021. The Overfitted Brain: Dreams Evolved to Assist Generalization.” Patterns.
Hubinger, Merwijk, Mikulik, et al. 2021. Risks from Learned Optimization in Advanced Machine Learning Systems.”
John, Caldwell, McCoy, et al. 2024. Dead Rats, Dopamine, Performance Metrics, and Peacock Tails: Proxy Failure Is an Inherent Risk in Goal-Oriented Systems.” Behavioral and Brain Sciences.
Koch, and Peterson. 2024. From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution.”
Majka, and El-Mhamdi. 2025. The Strong, Weak and Benign Goodhart’s Law. An Independence-Free and Paradigm-Agnostic Formalisation.”
Manheim, and Garrabrant. 2019. Categorizing Variants of Goodhart’s Law.”
Marklund, Infanger, and Roy. 2025. Misalignment from Treating Means as Ends.”
Peters, Krauss, and Braganza. 2022. Generalization Bias in Science.” Cognitive Science.
Solzhenit︠s︡yn. 2003. The Gulag Archipelago, 1918-56: An Experiment in Literary Investigation.
Welsh. 2024. Tukhta: Labour and Resistance in the Audit Regime of the Soviet Gulag.” Labor History.
Wheeler. 1993. Understanding Variation: The Key to Managing Chaos.