A negative result usually gets filed under one label: it didn't work. That single label hides six real cases where an intervention didn't just fail, it produced the exact opposite of its intended effect, and each one traces to a different, specific, well-documented mechanism.
None of these six mechanisms are new to this site. Naming which one an initiative is actually risking, before it ships, is what makes it preventable instead of a surprise in the data afterwards.
The story in four parts
A fine, a ban, a scare tactic, a comparison letter, a bounty, and a reminder email each produced the opposite of what it was built to do.
Crowding out, substitution, a wrong theory of the audience, a hidden subgroup, a gamed proxy, and an uncounted second number: never the same story twice.
Childcare, campus policy, criminal justice, energy use, colonial administration, and fundraising all produced the identical shape of failure.
A team that can name which of the six it's risking can check for it before shipping, not just discover it in the data afterwards.
The evidence at a glance
Gneezy and Rustichini's 2000 field study of 10 Israeli daycare centres found a modest lateness fine roughly doubled late pickups, and the increase held even after the fine was later removed.
A Campbell Collaboration review of nine randomised and quasi-randomised trials found “Scared Straight”-style programmes raised reoffending by 1 to 28% against untreated control groups.
A Danish charity's field experiment sent a second reminder to a subset of 17,000 past donors: one-time giving rose, and so did the rate of donors unsubscribing for good.
Six mechanisms, one shared thread underneath every one of them: a program's own causal theory, its measurement discipline, or both, is what would have caught each reversal below before it shipped.
The connection, in short
Same failure, a different name: the causal idea itself was wrong, and no amount of careful delivery could fix that.
None of this gets caught without calling the test what it is.
Not testing wouldn't have prevented these. It would have hidden them.
Six cases, six different rungs of the same causal ladder.
Start with the mechanism most people already sense intuitively, even without a name for it. A fine can turn a moral obligation into a fee anyone can simply pay off, and once it's a fee, the guilt that used to keep most people on time has nothing left to attach to.
Why did adding a fine for lateness make parents show up later, not sooner?
Ten Israeli daycare centres, frustrated by parents arriving late to collect their children, introduced a modest fine for lateness. Late pickups didn't fall. They roughly doubled, and stayed elevated even after the fine was later removed.
What it teaches: A fine can “crowd out” an existing moral motivation and replace it with a transactional one. Before the fine, lateness felt like imposing on the teacher; once a price was attached, it became something parents could simply buy, and the guilt that had kept most parents on time disappeared even once the price itself went away.
Crowding out replaces a motivation. Substitution replaces a behaviour, rerouting the same underlying want to whatever option is next closest, which can be worse on exactly the dimension the policy was trying to improve.
Why did banning bottled water increase the plastic shipped to campus?
The University of Vermont banned the sale of bottled water on campus in 2013, intending to cut plastic waste and nudge people towards water fountains and reusable bottles. A study tracking beverage sales before and after found sugary-drink purchases rose substantially, and the total volume of plastic bottles shipped to campus, now mostly full of soda and juice instead of water, increased rather than fell.
What it teaches: Removing one option in a choice set doesn't remove the underlying want. It reroutes it to whatever's next closest, which can be worse on exactly the dimension you were trying to improve. A policy aimed at one outcome needs to be measured against every outcome it could plausibly move, not just the one it was designed for.
Crowding out and substitution both assume the audience wants roughly the same thing before and after the intervention, just routed differently. The next mechanism is more unsettling: the intervention can teach the opposite lesson from the one intended.
Why did showing teenagers prison make them more likely to reoffend, not less?
Modelled on an idea that seems obviously sound: show at-risk teenagers the harsh reality of prison to deter future crime. These programmes remain popular and are still occasionally revived, including on reality TV. A systematic review of nine randomised and quasi-randomised trials found participants were, on average, more likely to reoffend afterwards than a matched comparison group who received no intervention at all.
What it teaches: Intuitive plausibility is not evidence. “Scaring” someone away from crime is a reasonable-sounding theory of behaviour change, but the randomised evidence points the other way. A programme can keep being funded and re-run for decades on the strength of a plausible story alone, without ever being properly tested against a control group.
Scared Straight's reversal showed up across the whole treated group. The next case hides its reversal inside just one slice of the population, invisible in the average and only visible once that population is actually split.
Why did telling frugal households they used less energy make them use more?
A field experiment sent households a letter comparing their energy use to their neighbours' average. Among households already using more than average, usage dropped. The intervention worked exactly as intended. But among households already using less than average, usage rose: apparently reassured that they had “room” to relax, they used more energy than before receiving any letter at all.
What it teaches: The same real, well-replicated mechanism (descriptive social norms, see Social Norm) can push a subgroup in exactly the wrong direction if an intervention isn't tested on the full range of people it will actually reach, not just the ones it's designed to help. The fix, found in the same research programme: adding an injunctive signal (a simple smiley or frowny face) eliminated the boomerang while keeping the reduction for high-usage households.
The energy-bill letter backfired on a subgroup nobody split out in advance. The next case backfired because the reward was attached to a measurable stand-in for the goal, not the goal itself.
Why did paying for rat tails end up breeding more rats, not fewer?
In 1902, French colonial administrators in Hanoi paid a bounty for every rat tail turned in, hoping to control a plague-linked rat population. Rat-catchers quickly began trapping rats, cutting off their tails, and releasing them back into the sewers to breed, and health inspectors soon found rat farms springing up specifically to supply tails for profit.
What it teaches: This one is documented colonial-administrative history, reconstructed from municipal archives, not a designed experiment, worth flagging plainly, the same way this site flags any illustrative case that isn't a controlled study. Even so, it's a genuine, well-sourced example of a perverse incentive: rewarding a proxy for a problem (tails) rather than the actual outcome (fewer rats) let people optimise for the reward in a way that made the real problem worse.
Rewarding a proxy is a measurement mistake made before an intervention runs. The last case is a measurement mistake made after it: the number everyone was tracking improved, and a different number nobody was watching quietly got worse.
Why did a second reminder raise donations and drive donors away for good?
A Danish charity sent reminder emails to roughly 17,000 past donors, and a subset also received a second reminder a week later. The extra reminder worked exactly as designed. One-time donations rose substantially. But donors who received it unsubscribed from future contact at a significantly higher rate, and a follow-up test found announcing monthly reminders (versus a single future one) drove even more people to opt out for good.
What it teaches: A backfire doesn't always show up in the metric being optimised. Judged purely on short-term donations, the reminder was an unambiguous win. The real cost only shows up in a different number, months later, once you count the donors who can never be reminded again. Measuring the metric you changed isn't the same as measuring everything you changed. This mechanism now has its own Principles entry: Reminder Fatigue.
Six different mechanisms, one shared structural gap. Every case above passed whatever check its own team was running at the time. None of them had a check built for the specific way their own intervention could reverse, because nobody had named that mechanism as a risk before shipping it.
The six cases above are specific and varied. The four checks below are general enough to run on any intervention before it ships, whatever mechanism it happens to risk.