Special Report

The test was real. The conclusion wasn't.

Ask why a famous study or a famous experiment fell apart and most people guess fraud or sloppiness. The more common failure is quieter, and comes in three flavours. A study run honestly that never had enough data to see what it claimed. A real result from one place, one sample, one afternoon, treated afterwards as a law of nature. And, especially common in product and marketing testing, a change that was never built on a real theory of why it should work in the first place, just intuition and a metric that moved, with no real evidence that metric ever reaches the outcome that actually matters. This is a field guide to thirteen real cases where one of those three failures mattered, what the correct habit would have caught, and when, if ever, someone checked.

The story in five parts

A p-value under 0.05 isn't proof of anything

It says a pattern that size is unlikely under pure chance, nothing about how big, important, or true it is.

Small samples can't reliably see what they're looking for

Low statistical power means a real effect gets missed, and a “significant” one is usually exaggerated.

One place is not everywhere

A result from 24 volunteers, or one factory in the 1920s, describes that group, not humanity.

Guessing without a theory runs out of road fast

Tactics copied without a mechanism produce a short list of things to try, then nothing.

A rising metric doesn't prove the real outcome moved

Clicks, opens, and time-on-page can climb while the thing they're supposed to predict stays flat, or falls.

The evidence at a glance

8 of 9 experiments “worked”

Bem's 2011 precognition studies found significant results in 8 of 9 experiments. A pre-registered, independent three-lab replication found nothing.

2.4 million responses, still wrong

The 1936 Literary Digest poll drew 2.4 million responses and predicted the wrong winner by a 19-point margin, skewed by who actually replied.

A famous effect, cut in half

The Marshmallow Test replicated with ten times the original sample found the effect roughly half the size, and it vanished after controlling for income.

36 students funded a statewide policy

The Mozart Effect came from 36 college students and a temporary spatial-task boost. Georgia funded a classical CD for every newborn on the strength of it.

Netflix's own model, checked against reality

Tested against 200 real A/B tests, Netflix's surrogate model agreed with the true result only 65 to 79% of the time.

Three terms, used constantly, understood loosely

None of the cases below required fraud. The first ten are explained by a plain misunderstanding of one of these three ideas, so it's worth being precise about what each one actually means before getting to the examples.

Concept

Statistical power

If the effect is real, will a study this size even be able to see it?

Statistical power is the probability a study will detect a real effect, given how many people it tested and how big that effect truly is. A study with 20 or 30 people per group, common among the cases below, can easily have less than a 50 percent chance of detecting a real, moderate effect, even when one genuinely exists. That cuts both ways: a real effect can be missed entirely, and when a small study does turn up something “significant,” the true effect is usually smaller than reported, an inflation researchers call the winner's curse.

What it teaches: a small sample doesn't just risk missing a real effect. When it finds something anyway, that something is usually exaggerated, which is exactly why so many small, striking early results shrink once a larger study repeats the test.

The mechanism
Small sample (20–30 people)
A real effect exists
Not enough power to see it clearly
If “significant” anyway, the size is inflated

Small studies don't just miss real effects. When they do find something, it's usually bigger on paper than in reality.

Concept

Statistical significance

What does “p less than 0.05” actually rule out?

A p-value answers one narrow question: if there were truly no effect at all, how likely would data this striking be, purely from random noise? A p-value under 0.05 means a result that size would turn up by chance less than 5 percent of the time under that one assumption, nothing more. It says nothing about how large the effect is, how likely the underlying claim is to be true, or whether it will replicate. Test several outcomes and report only the one that clears the bar, or quietly drop a few “messy” data points, and a significant p-value becomes easy to manufacture from pure noise, a practice now called p-hacking.

What it teaches: “statistically significant” is a far narrower claim than it sounds. It means a pattern is unlikely to be pure chance under one specific assumption, not that the finding is important, large, or even real.

The mechanism
Random noise, no real effect
Try several splits or analyses
One clears p < 0.05 by chance
Published and cited as a real finding

Enough tries, and noise eventually clears the bar. That's what pre-registration exists to prevent.

Concept

Confidence intervals

A result can be “significant” and still be a coin flip's width from meaning nothing.

A confidence interval gives a plausible range for the true effect, not just a single number. A study reporting “improved outcomes by 10 points (95% CI: 1 to 19)” is a real, significant result, but that range spans a barely-there effect and a genuinely large one; the study can't tell those two stories apart. Small studies tend to produce wide intervals, precisely because there isn't enough data to pin the true value down, and a single headline number often gets repeated without anyone mentioning how uncertain it actually was.

What it teaches: two studies can both clear the significance bar while one is a precise, trustworthy estimate and the other is barely distinguishable from finding nothing at all. The interval, not just the number in the headline, is the honest version of a result.

The mechanism
One headline number: “+10 points”
Real range: 95% CI, 1 to 19
Spans barely-there and genuinely large

The range is the honest version of the result. A single point estimate hides how uncertain it really is.

Part one: small samples, misread significance

Four other famous replication failures, power posing, ego depletion, the facial feedback hypothesis, and the original elderly-priming study, already have their own case studies on this site's Reading the Research page. The five below are a different set, chosen to show the same failure mode from different angles: flexible analysis, a fragile statistical model, a hidden bias in the method itself, and a sample so large it fooled everyone into skipping the one check that mattered.

The pattern across all five
Small sample or flexible analysis
Noise looks like a real pattern
Published, cited, believed for years
Fails once someone finally re-checks

The gap between looking real and being real is exactly the gap statistical power and honest analysis are supposed to close.

Significance

The Nine Experiments That Found ESP

How did a respected psychology journal publish nine experiments appearing to prove precognition?

In 2011, Cornell psychologist Daryl Bem published nine experiments in a leading psychology journal, eight of which reported statistically significant evidence that people could sense random future events before they happened. Bem later acknowledged he had not pre-registered his hypotheses or methods, and had used the flexible, exploratory analysis common at the time, trying several ways of splitting and scoring the data and reporting whichever cleared the significance threshold.

What it teaches: a determined analyst can find “significant” effects in almost any noisy dataset if the hypothesis, sample, and analysis aren't locked in before looking at the results. A pre-registered attempt to replicate one of the nine effects, run independently by three labs, found no evidence of it at all. The episode became one of the direct triggers for psychology's push towards pre-registration.

Bem, D. J. (2011). “Feeling the Future: Experimental Evidence for Anomalous Retroactive Influences on Cognition and Affect.” Journal of Personality and Social Psychology, 100(3), 407–425. Replication: Ritchie, S. J., Wiseman, R., & French, C. C. (2012). “Failing the Future.” PLOS ONE, 7(3), e33423
Significance

Himmicanes

Do people really underprepare for a storm just because it has a woman's name?

A 2014 study analysed 92 U.S. hurricanes from 1950 to 2012 and reported that storms with more feminine names caused significantly more deaths, its proposed explanation being that people take a “Hurricane Sandy” less seriously than a “Hurricane Charlie.” The finding was picked up by hundreds of news outlets within days. Other statisticians who re-ran the analysis found the result depended heavily on a handful of unusually deadly storms and on specific choices inside the statistical model; changing the model or excluding a few outlier years made the effect disappear or reverse.

What it teaches: with a small number of independent events, 92 storms is a small sample for this kind of modelling, and many defensible ways to build the statistical model, researchers can end up choosing, often without meaning to, the exact setup that happens to produce a striking result. This is sometimes called researcher degrees of freedom, a close cousin of p-hacking.

Jung, K., Shavitt, S., Viswanathan, M., & Hilbe, J. M. (2014). “Female Hurricanes Are Deadlier Than Male Hurricanes.” PNAS, 111(24), 8782–8787. Critiques published in the same volume by Bakkensen & Larson, and Christensen & Christensen
Significance

The Hot Hand, Reversed

What if the famous study that “debunked” a real effect had a bug in its own maths?

A landmark 1985 study concluded that the “hot hand” in basketball, the belief that a player who just made several shots is more likely to make the next one, was a myth: players' sequences of hits and misses looked statistically indistinguishable from a coin flip. That conclusion stood as settled fact for three decades. In 2015, economists Joshua Miller and Adam Sanjurjo identified a subtle bias in the original method: the natural way of measuring “chance of a hit right after a streak” in a short, finite sequence is itself biased towards finding no streak effect, even in truly random data. Correcting for the bias and re-running the original numbers turned up real evidence of a hot hand after all.

What it teaches: the mistake isn't always a small sample or a flexible choice. Sometimes the statistical method has a hidden bias that manufactures a false negative rather than a false positive, and it can take decades to notice, since a null result rarely gets the scrutiny a striking one does.

Gilovich, T., Vallone, R., & Tversky, A. (1985). “The Hot Hand in Basketball.” Cognitive Psychology, 17(3), 295–314. Correction: Miller, J. B., & Sanjurjo, A. (2018). “Surprised by the Gambler's and Hot Hand Fallacies? A Truth in the Law of Small Numbers.” Econometrica, 86(6), 2019–2047
Significance

The Dunning-Kruger Effect's Maths Problem

Does “the least skilled overrate themselves most” show up even in random numbers?

The Dunning-Kruger effect, that the least competent people rate their own ability the most inaccurately high, is one of the most widely cited findings in psychology. Its standard test method sorts people into quartiles by their actual score, then compares each quartile's average self-rating to its actual score. Statisticians later showed this exact method produces a Dunning-Kruger-shaped pattern even when the underlying ability and self-rating numbers are pure, unrelated random noise, because sorting people by the same score used to measure their error creates a spurious correlation, and a regression-to-the-mean effect, on its own. Re-analysed with methods that avoid this, the effect shrinks to far smaller than the original claim.

What it teaches: a real-looking pattern can be partly, or even mostly, a property of the method used to look for it rather than the thing being measured. It's worth asking not just whether a result is significant, but whether the analysis technique itself could produce that exact shape out of noise.

Kruger, J., & Dunning, D. (1999). “Unskilled and Unaware of It.” Journal of Personality and Social Psychology, 77(6), 1121–1134. Critique: Gignac, G. E., & Zajenkowski, M. (2020). “The Dunning-Kruger Effect Is (Mostly) a Statistical Artefact.” Intelligence, 80, 101449
Confidence

The Poll With 2.4 Million Responses, Still Wrong

Why did a poll of 2.4 million people get a landslide election backwards?

Ahead of the 1936 U.S. presidential election, the magazine Literary Digest mailed 10 million straw-poll ballots and received 2.4 million back, an enormous sample by any standard, and confidently predicted Alf Landon would beat Franklin Roosevelt 57 to 43 percent. Roosevelt won in a landslide, 62 to 37. The mailing list, built from telephone numbers, car registrations, and club memberships, skewed towards wealthier households in the middle of the Great Depression, and a modern statistical reanalysis found an even larger driver: the roughly 76 percent of people who never returned a ballot leaned towards Roosevelt far more than the 24 percent who did.

What it teaches: a sample size in the millions cannot fix a sample that isn't representative of the population it claims to describe. It's the oldest lesson in survey statistics, and still the first thing worth checking about any “huge dataset” claim: huge compared to what, and who's missing from it?

Reanalysis: Lohr, S. L., & Brick, J. M. (2017). “Roosevelt Predicted to Win: Revisiting the 1936 Literary Digest Poll.” Statistics, Politics and Policy, 8(1), 65–84

The four already covered elsewhere on this site belong to the same family: power posing, ego depletion, the facial feedback hypothesis, and the original elderly-priming study all trace back to a sample too small, or a measurement too unblinded, to support the claim built on top of it.

Part two: one study, one place, extended everywhere

These five aren't statistically underpowered in the same way. Several are well-documented, honestly reported results. The failure is what happened after: a finding measured in one narrow sample, one factory, one afternoon, got generalised into a claim about people, or policy, far beyond anything the original study could support.

The pattern across all five
One narrow sample or setting
A real, honest effect there
Extended into a claim about everyone
Breaks down outside the original context

Real, inside a specific group, isn't the same as universal. The extension, not the original result, is where each of these went wrong.

Overgeneralised

The Marshmallow Test's Missing Context

Does delaying a marshmallow at age four really predict your whole life?

Starting in the late 1960s, fewer than 90 children, all enrolled at a preschool on Stanford's own campus and mostly from financially comfortable, highly educated families, were offered one marshmallow now or two later if they waited. Children who waited longer showed modestly better outcomes years afterwards, and the finding became one of psychology's most famous parables about willpower and self-control. A 2018 study using a much larger, more economically and racially diverse sample of roughly 900 children found the same basic pattern, but at about half the size, and it disappeared almost entirely once researchers accounted for family income and the child's early cognitive ability.

What it teaches: an effect measured in one narrow, unrepresentative group can be genuine within that group and still misrepresent the underlying cause once tested on a broader population. “Delaying gratification builds a better life” turned out to overlap heavily with “growing up with financial security lets a child feel safe enough to wait, and also predicts a better life for other reasons.”

Mischel, W., & Ebbesen, E. B. (1970). “Attention in Delay of Gratification.” Journal of Personality and Social Psychology, 16(2), 329–337. Replication: Watts, T. W., Duncan, G. J., & Quan, H. (2018). “Revisiting the Marshmallow Test.” Psychological Science, 29(7), 1159–1177
Overgeneralised

The Prison Experiment That Ran Once

Can 24 college students, in one basement, for six days, tell you what's inside human nature?

In 1971, Philip Zimbardo randomly assigned 24 male college volunteers to play guards or prisoners in a mock prison in a Stanford basement. Guards, the story goes, spontaneously grew cruel, and the study became one of the most cited demonstrations in psychology that situations, not personalities, corrupt behaviour, taught for decades in introductory textbooks worldwide. There was no control group, no pre-registered hypothesis, and no real statistical analysis, only a dramatic narrative from one uncontrolled run. A 2018 investigation using newly available recordings and interviews found guards had been explicitly coached by the research team on how to be “tough,” directly undercutting the claim that cruelty emerged spontaneously from the situation alone.

What it teaches: a vivid, well-told single case can feel like overwhelming evidence while containing none of the safeguards, a comparison group, blinding, a real sample, that separate a genuine finding from a compelling story. It ran once, on 24 self-selected volunteers, with essentially no data analysis, and still shaped how millions of people were taught to think about human nature.

Haney, C., Banks, C., & Zimbardo, P. (1973). “Study of Prisoners and Guards in a Simulated Prison.” Naval Research Reviews, 9, 1–17. Investigation: Le Texier, T. (2019). “Debunking the Stanford Prison Experiment.” American Psychologist, 74(7), 823–839
Overgeneralised

The Effect Named After Data Nobody Had Checked

What happens when someone finally runs the numbers behind a term used in every management textbook?

Between 1924 and 1932, researchers at Western Electric's Hawthorne Works plant varied factory lighting and reported that worker productivity rose no matter which way the lighting changed, even when it got dimmer, supposedly because workers respond to being observed rather than to the change itself. The story became so influential that “the Hawthorne effect” entered the language as a general term for people changing behaviour simply because they know they're being watched. For decades the original data were assumed lost. When economists Steven Levitt and John List finally located and formally analysed the records in 2011, the dramatic pattern largely evaporated: once day-of-week and pay-period effects were accounted for, the productivity swings tracked the calendar far better than they tracked the lighting.

What it teaches: an idea can become permanently embedded in a field's vocabulary and taught for generations before anyone analyses the data it was built on. “Everyone knows this effect is real” and “the data shows this effect” are different claims, and for the study the whole concept is named after, they turned out to say quite different things.

Original studies documented in Roethlisberger, F. J., & Dickson, W. J. (1939). Management and the Worker. Harvard University Press. Reanalysis: Levitt, S. D., & List, J. A. (2011). “Was There Really a Hawthorne Effect at the Hawthorne Plant?” American Economic Journal: Applied Economics, 3(1), 224–238
Overgeneralised

Ten Minutes of Mozart, One State Policy

Does listening to Mozart make a baby smarter, or did the claim just outrun the study?

In 1993, 36 college students briefly scored higher on one specific spatial-reasoning task after listening to ten minutes of a Mozart piano sonata, compared with sitting in silence or hearing relaxation instructions. The boost, roughly 8 to 9 IQ-equivalent points on that one task, faded within 10 to 15 minutes, and the original researchers never claimed it applied to general intelligence, children, or long-term development. Media coverage and popular books extended it into “classical music makes babies smarter” anyway, and in 1998 the governor of Georgia budgeted funds to mail every newborn in the state a classical music CD. A later meta-analysis of many replication attempts found the true effect, if any, is small and specific to certain spatial tasks, nothing like the general intelligence boost the popular version claimed.

What it teaches: the size and scope of a claim can grow enormously between the study and the policy built on it without anyone along the way deliberately lying. A short-lived effect on one narrow task, in adults, quietly became a permanent effect on general intelligence, in infants, and a government spent real money on the gap.

Rauscher, F. H., Shaw, G. L., & Ky, K. N. (1993). “Music and Spatial Task Performance.” Nature, 365, 611. Meta-analysis: Pietschnig, J., Voracek, M., & Formann, A. K. (2010). “Mozart Effect-Shmozart Effect.” Intelligence, 38(3), 314–323
Overgeneralised

A Magazine Essay Became a National Policing Strategy

How much data was behind the idea that one broken window invites more crime?

In a 1982 Atlantic essay, James Q. Wilson and George Kelling argued that visible disorder, a broken window left unrepaired, signals nobody is watching and invites more serious crime. The essay contained no original data or study; it was a theoretical argument illustrated with anecdotes. It nonetheless became the intellectual basis for “zero-tolerance” and “order-maintenance” policing adopted at city scale, most famously in 1990s New York, and is cited in criminology as one of the most influential ideas in the field's history. Decades later, a National Research Council review and later meta-analyses found little support for the specific claim that disorder itself causes serious crime; the most rigorous analysis found a modest crime-reduction effect from disorder policing, but concentrated in cooperative, community-based approaches rather than the aggressive enforcement the essay's name became attached to.

What it teaches: an idea doesn't need data behind it to become policy at national scale if it's intuitive and confidently argued. The gap between a compelling essay and a tested causal claim was never closed before the theory was already reshaping how millions of people were policed.

Wilson, J. Q., & Kelling, G. L. (1982). “Broken Windows.” The Atlantic. Review: Braga, A. A., Welsh, B. C., & Schnell, C. (2015). “Can Policing Disorder Reduce Crime?” Journal of Research in Crime and Delinquency, 52(4), 567–588

Part three: testing without a theory of why

The first ten cases are all honestly run studies undone by a statistics problem. This last set is different, and it's the one most product and marketing teams actually live in day to day: a test that never had a real behavioural hypothesis behind it in the first place, just intuition about what might work, and a metric that moved without anyone checking whether that metric actually predicts the outcome the business cares about.

The first half of this is a strategy problem. Better statistics can't fix it. A team testing button colours, headline wording, or where a badge sits on a page, without a theory of the actual psychological mechanism at work, can generate a short list of plausible tweaks and then runs out: there's no underlying model producing the next idea, only whatever's left to try. A real behavioural mechanism, the kind cited throughout the rest of this site, does the opposite. One theory (loss aversion, social proof, the goal-gradient effect) predicts a whole family of testable changes across a product, not just the one screen someone happened to be looking at, which is what actually lets a team make a strategic bet instead of nudging one small stage of a funnel over and over.

The second half compounds the first: even a real, statistically solid lift in an engagement metric, more clicks, more opens, more time on a page, is often assumed to cascade automatically into the outcome that actually matters (revenue, retention, real behaviour change) without that connection ever being tested. Sometimes it does. Sometimes the two move in opposite directions.

The pattern behind both
No behavioural mechanism, just a guess
Tweak something, a metric moves
Assume it cascades to the real outcome
Never validated. Sometimes it doesn't.

A metric is not a mechanism. Moving one, without a theory connecting it to the other, is a guess wearing a graph.

No theory

Ranking by Clicks Rewarded the Wrong Posts

What happens when a platform optimises for clicks instead of the value a click is supposed to represent?

Facebook's News Feed ranking historically weighted how often people clicked a link. By 2013 and 2014, publishers had adapted with headlines engineered purely to trigger a click, “click-baiting,” without describing what the story actually contained. Clicks on this content were genuinely high. Facebook's own data showed something else moving the opposite way: people who clicked a click-bait headline spent less time reading afterwards and were far less likely to like, comment, or share once they returned, exactly the signals the click was supposed to predict.

What it teaches: Facebook had to explicitly rebuild its ranking to measure bounce-back rate and post-click engagement, not just the click itself, because the click alone had stopped meaning what it was assumed to mean. The fix wasn't a bigger sample or a better p-value. It was admitting the metric being optimised was never validated against the outcome it stood in for.

Facebook Newsroom (2014). “News Feed FYI: Click-baiting.”
No theory

Even Netflix Had to Build a Model to Trust a Short-Term Number

How much statistical machinery does it take before a 14-day metric can stand in for a real outcome 63 days later?

Waiting months for every test's true long-term effect is expensive, so Netflix built a “surrogate index,” a statistical model combining several short-term signals into one number meant to predict a longer-term outcome, and checked it directly against 1,098 test arms from 200 real A/B tests. Even this purpose-built, heavily validated model only agreed with the true 63-day result about 95 percent of the time, and correctly identified which tests were actually worth launching only 65 to 79 percent of the time.

What it teaches: if it takes a dedicated model, tested against 200 real experiments, to get to 95 percent confidence that a short-term number predicts a long-term one, assuming an engagement metric will “obviously cascade” down the funnel, with no model and no validation at all, is not a safe default anywhere.

Zhang, V., Zhao, M., Dimakopoulou, M., Le, A., & Kallus, N. (2024). “Evaluating the Surrogate Index as a Decision-Making Tool Using 200 A/B Tests at Netflix.” arXiv:2311.11922. Framework: Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press
Experiment Teardown
News Feed ranking, 2013–2014
Same posts, ranked by a different signal
Ranked by raw clicks
Vague, curiosity-gap headlines rank higher
High click-through rate
Fast bounce back, little liking or sharing afterwards
Ranked by bounce rate + engagement
Descriptive headlines rank higher
Somewhat lower click-through rate
More time reading, more likes and shares afterwards

The metric that looked better was the one predicting less real value. Raw click-through rewarded exactly the headlines least likely to satisfy the person who clicked them.

Six questions that would have caught most of these

None of the thirteen cases above needed fraud to go wrong. Asking these six questions, before treating a finding or a test result as settled, would have flagged nearly all of them years earlier.

  1. Was the sample large enough to detect the claimed effect?A study with 20 to 40 participants per group is rarely powered to reliably detect anything but a very large effect. Ask what the study would have needed, not just what it had.
  2. Was the analysis locked in before the results came in?If a researcher could have tried several outcomes, splits, or models and reported the one that worked, a significant p-value proves much less than it appears to.
  3. How wide is the confidence interval, not just whether it excludes zero?A significant result with a wide interval is still a highly uncertain one. Two “significant” findings can carry very different amounts of real information.
  4. Does the sample resemble the population the claim is being applied to?Ninety Stanford preschoolers, 24 self-selected volunteers, and one factory in the 1920s are each a real result about a specific group, not a free pass to generalise to everyone.
  5. Is there a real behavioural hypothesis behind this test, or just a guess at what might work?A mechanism generates the next ten tests on its own. A guess that happened to work once generates nothing else to try.
  6. Has the metric being moved actually been shown to predict the outcome that matters?Clicks, opens, and time-on-page are only useful if someone has checked that they cascade to revenue, retention, or whatever the real goal is. Otherwise a rising metric is a fact about the metric, not about the business.