Special Report

Policy-based evidence, not evidence-based policy

“Evidence-based policy” is the phrase every government department and large organisation already claims to practise. Economist John List's argument, most fully laid out in a 2024 Nature paper, is that the phrase hides a quiet ordering problem.

Most of that evidence was generated by a small pilot that was never designed to survive being scaled up, so the policy inherits a result the evidence was never built to support.

List's proposed fix flips the order: deliberately design the evidence-generating study around the constraints of the full-scale rollout from the start. What comes out the other end is policy-based evidence, evidence actually built for the policy you intend to run, at the size you intend to run it.

The story in four parts

“Evidence-based policy” hides an ordering problem

Most of that evidence came from a small pilot that was never designed to survive being scaled up.

A pilot's result and a scaled result are different things

The same effect, tested small, often shrinks once it's rolled out to a whole population, List's “voltage effect.”

Scaling dilutes exactly what made the pilot work

You can't hire enough of your pilot's best staff to run the same programme nationally.

List's flip: design the evidence for the policy you intend to run

Test with realistic staff, the full population, and at real scale, from the very first study.

The evidence at a glance

The shrinkage has a name

List's own 2024 Nature paper calls it the voltage effect: a pilot's real, measured result reliably falls once the same programme runs at full population scale.

Governments trial far less than companies do

List and Heffetz describe governments as unusually resistant to randomised policy trials, even as companies routinely run ad A/B tests as a matter of course.

A pilot's own best staff can't scale with it

A national rollout can't hire enough of the pilot's own best-performing staff to run the programme at the same quality everywhere.

A story can outlast decades without a control group

Scared Straight-style programmes kept getting funded for decades on the strength of a compelling story, not a controlled comparison.

The pilot that never had to survive scale

A typical policy pilot runs with an enthusiastic delivery team, a small and often self-selected group of participants, close hands-on oversight from the researchers who designed it, and a short enough timeframe that the novelty of being studied hasn't worn off. None of that is dishonest. It's also nothing like the conditions the policy will actually run under once it's serving a whole population through an ordinary public agency with ordinary staff and ordinary funding.

This site already covers the general version of that gap under the “voltage effect”: an intervention's measured effect shrinking, sometimes to a fraction of its pilot size, once it's rolled out for real. See Reading the Research: the voltage drop at scale and the field session Why your best pilot is the one most likely to lose its voltage for the full mechanism. What's specific to government policy is what happens next. A small pilot's scattered, limited-external-validity result still gets written into permanent legislation or a national programme, because by the time anyone asks “will this actually work at this size, with this staff, for this whole population,” the funding and political commitment are already spent.

The site's own cautionary case for exactly this pattern is “Scared Straight” programmes: a plausible-sounding intervention scaled nationally, and kept being funded for decades, on the strength of a story rather than a control group. List's argument in this piece is the more general version of that same failure: it isn't only that pilots can be badly designed. Even a properly randomised pilot can produce evidence that simply wasn't built to answer the question “what happens at scale,” because nobody designed it to answer that question in the first place.

List's flip: design the evidence for the policy you intend to scale

List's proposal is to reverse the usual sequence. Instead of running a convenient small study and then hoping, or lobbying, for it to scale, the recommendation is to build the scaling constraints into the evidence-generating study from day one. Test with the staff quality the full rollout will actually have access to, not the pilot's specially recruited team. Test across the full range of the population the policy will eventually cover, not just the most reachable or most motivated slice. And design the study to detect the general-equilibrium effects (prices, capacity, spillovers onto people not directly treated) that only show up once a programme is operating at real scale.

That's the substance behind the phrase “policy-based evidence”: evidence whose design already answers to the shape of the eventual policy. “Evidence-based policy” works the other way around: the policy gets built afterwards on top of a study that was only ever answering a narrower, more convenient question. List frames this explicitly as a call to social scientists to change how they design studies meant to inform public policy, not a claim that all existing evidence should be distrusted.

List, J. A. (2024). “Optimally Generate Policy-Based Evidence Before Scaling.” Nature, 626, 491–499

Why this is rare, not just difficult

Designing a study this way costs more, takes longer, and is politically harder to greenlight than a small, fast, encouraging pilot. That's a large part of why it doesn't happen by default.

Political economy

Randomising a Policy Feels Different From Randomising an Ad

Why is a government far more reluctant to run a fair test than a company is?

List and co-author Ori Heffetz argue that governments are unusually resistant to randomised policy trials, even when doing so would settle a genuinely contested question. Randomly withholding a policy from some citizens reads as unfair in a way that a company's A/B test on a marketing email never does, even though the statistical logic is identical.

What it teaches: The barrier to scaled, well-designed evidence in government is often not technical. It's that the randomisation step itself needs its own political and ethical case made before the study can even start.

Heffetz, O., & List, J. A. (2021). “Who's Afraid of Evidence-Based Policymaking?” Project Syndicate
Voltage effect

The Staff You Can't Hire At Scale

What happens when a pilot's best people aren't the people running it nationally?

One of the recurring drivers behind the voltage effect, already covered on this site as the general pattern, is specific and structural. A small pilot is often staffed by unusually motivated or skilled people, whether that's teachers, caseworkers, or delivery staff, because a small programme only needs a handful of them. Scaling the same programme to a full population means hiring far more staff than the pool of similarly exceptional people can supply, so the average quality delivering the programme falls exactly as the programme grows.

What it teaches: A pilot's result is partly a result about the specific people who ran it. Evidence designed to survive scaling has to test with realistically staffed delivery, not the pilot's best available team, or it's measuring a programme that can never actually be run at the size it's meant for.

List, J. A. (2022). The Voltage Effect: How to Make Good Ideas Great and Great Ideas Scale. Currency. See also Reading the Research: the voltage drop at scale.