Principles

83 principles, one mechanism at a time

Each one is a real cognitive bias or heuristic, broken down to the study that first documented it. Browse by category, or jump straight to any principle by name.

All Principles

83 principles, searchable and filterable

Search by name or keyword, or filter by the underlying psychological mechanism. Tap any row to read the full breakdown.

01

Anchoring

Why does the first number you see decide what a fair price feels like?

The first number you see becomes the reference point every later number gets judged against, even when it's arbitrary.

A RIGGED WHEEL LANDING ON 65 MADE PEOPLE GUESS FAR HIGHER WHEEL LANDS ON 10 25% MEDIAN GUESS WHEEL LANDS ON 65 45% MEDIAN GUESS Everyone knew the wheel was random. The number it landed on still anchored their guess.
Likely mechanismA first number anchors judgment, and adjusting away from it stops too early

The psychology. This is the "anchoring-and-adjustment" heuristic: instead of estimating from scratch, people start from the anchor and adjust away from it, but the adjustment reliably stops too early, as soon as the number feels plausible rather than when it's actually correct.

The full write-up: study, numbers, and caveats
The psychology

This is the "anchoring-and-adjustment" heuristic: instead of estimating from scratch, people start from the anchor and adjust away from it, but the adjustment reliably stops too early, as soon as the number feels plausible rather than when it's actually correct.

Where it causes errors

Overpaying because a "was" price inflates the reference point; accepting a worse negotiation outcome because the other side's opening number reframed what "reasonable" looks like; being measurably swayed by numbers you know are meaningless, like a rigged wheel.

Where it can help

Opening a negotiation with a well-calibrated first offer sets a favourable anchor for both sides. Giving someone a realistic reference figure before they decide helps them avoid over- or under-shooting from having no anchor at all.

Tversky, A., & Kahneman, D. (1974). "Judgment under Uncertainty: Heuristics and Biases." Science, 185(4157), 1124–1131
The foundational paper: anchoring is one of three heuristics it identifies, alongside representativeness and availability.
StrengthParticipants watched a wheel of fortune rigged to land on 10 or 65, then estimated the percentage of African countries in the UN. Because the wheel was visibly random, the design cleanly isolates the anchor from any actual information.
WeaknessSmall, non-representative samples typical of 1970s lab psychology, and a single numeric-estimation task, generalising to real decisions like pricing required later replication.
Key findings
Verbatim“the median estimates of the percentage of African countries in the United Nations were 25 and 45 for groups that received 10 and 65, respectively, as starting points. Payoffs for accuracy did not reduce the anchoring effect.”
ParaphrasedThe paper frames anchoring as a general-purpose heuristic underlying everyday probability judgment, not a quirk specific to one task.
See also: Behavioural Economics vs. Behavioural Science vs. Psychology vs. UX, on why this particular finding kept its original name across psychology, economics, and product design, when most don't.
Run this one live: Live Sessions: Anchoring, the same paradigm and the same two anchors (10 and 65), run with a real room instead of a 1974 lab.
02

Sludge (Obstruction)

Why does cancelling take five clicks when signing up took one?

Friction added deliberately against your interest, like a cancellation flow that interrupts you with a retention offer instead of letting you leave.

CANCELLING TOOK 4 PAGES, SIGNING UP TOOK ONE CLICK SIGNING UP 1 CLICK DONE CANCELLING 6 CLICKS ACROSS 4 PAGES Same company, same account. Only the exit was made this much harder.
Likely mechanismThe effort of a hard step is felt right now; the benefit it unlocks feels distant and abstract

The psychology. Exploits present bias and effort aversion: the cost of navigating the friction is paid right now and viscerally, while the benefit of following through, money saved, an unwanted service ended, is delayed and abstract. People disproportionately abandon the effortful path even when they still want the outcome.

The full write-up: study, numbers, and caveats
The psychology

Exploits present bias and effort aversion: the cost of navigating the friction is paid right now and viscerally, while the benefit of following through, money saved, an unwanted service ended, is delayed and abstract. People disproportionately abandon the effortful path even when they still want the outcome.

Where it causes errors

People keep paying for services they've genuinely decided to leave, purely because leaving costs more perceived effort than staying, design friction overrides real intent, not a change of mind.

Where it can help

The inverse: removing friction from actions that help people, like auto-enrolling employees into retirement savings, or one-click opt-outs, is "sludge removal," and is the explicit basis of the FTC's Click-to-Cancel rule requiring cancellation to be no harder than sign-up.

Mathur, A., et al. (2019). "Dark Patterns at Scale: Findings from a Crawl of 11K Shopping Websites." Proceedings of the ACM on Human-Computer Interaction (CSCW)
The largest systematic study of dark patterns to date, sludge is one of several categories it documents.
StrengthAutomated crawl of 53,180 product pages across 11,286 shopping sites, producing a reusable taxonomy still cited in FTC enforcement actions today. Scale far beyond any single case study.
WeaknessAutomated detection struggles with patterns that only surface deep in a multi-step flow, like a cancellation buried behind a required phone call, so true prevalence is likely understated.
Key findings
Verbatim (abstract)“Analyzing ~53K product pages from ~11K shopping websites, we discover 1,818 dark pattern instances, together representing 15 types and 7 broader categories.”
Verbatim“We identified 22 third-party entities that provide shopping websites with the ability to create and implement dark patterns on their sites.”
Verbatim“Obstruction makes it easy for users to sign up for recurring subscriptions and memberships, but it makes it hard for them to subsequently cancel the subscriptions.”
ParaphrasedThe paper documents this specific pattern, which it calls Hard to Cancel, on 31 of the ~11,000 shopping websites it crawled.
Also worth citing: FTC v. Amazon (2025): $2.5B settlement over a Prime cancellation flow requiring 4 pages, 6 clicks, and 15 options.
See also: Good Friction, the same mechanism, pointed the other way: friction added against your interest here, in your interest there. Which one you're looking at depends entirely on who it protects.
03

Zero Price Effect

Why does free feel like a different universe from one cent?

Free isn't just a low price. It's a different category of decision. Demand jumps disproportionately the moment a price hits $0 .

MAKING ONE CHOCOLATE FREE TRIPLED HOW OFTEN IT WAS PICKED 1¢ VS. 15¢ CHOCOLATE 14% PICKED THE 1¢ ONE FREE VS. 14¢ CHOCOLATE 42% PICKED THE FREE ONE The 14¢ price gap never changed. Making one side free did.
Likely mechanismZero treated as a categorically different price, not just a low one

The psychology. Standard theory treats price as continuous. A 1-cent difference should feel roughly the same wherever it falls. People instead treat "free" as categorically different: with $0 there's no possibility of feeling like you made a bad trade, so the transaction reads as pure gain rather than an exchange.

The full write-up: study, numbers, and caveats
The psychology

Standard theory treats price as continuous. A 1-cent difference should feel roughly the same wherever it falls. People instead treat "free" as categorically different: with $0 there's no possibility of feeling like you made a bad trade, so the transaction reads as pure gain rather than an exchange.

Where it causes errors

Choosing a free but inferior option over a cheap superior one, even when the price gap between them is identical, in cents, to a gap elsewhere that barely moves anyone's preference.

Where it can help

Structuring a genuinely valuable free tier or sample lets people experience real value at zero risk, provided the free offer isn't bait for a later obstruction (see Sludge) or a degraded product.

Shampanier, K., Mazar, N., & Ariely, D. (2007). "Zero as a Special Price: The True Value of Free Products." Marketing Science, 26(6), 742–757
The paper that gives this principle its name: the ordinary vs. luxury chocolate study.
Experiment Teardown
398 participants
Only the Hershey's price moved, to zero
Cost condition
Hershey's, 15¢ Lindt
Picked one, or neither
Free condition
Hershey's, 14¢ Lindt
Picked one, or neither
Result Cost Free Swing
Hershey's 14% 42% +28
Lindt 36% 19% −17

Only the zero moved. The 14¢ gap between the two chocolates was identical in both conditions, yet making the Hershey's free flipped which one people picked.

Internal validityConditions were alternated in fixed time blocks rather than randomized per person, a limitation the paper states itself. The reported percentages come from a real chocolate booth in MIT's student center (n=398); a second field replication at an MIT cafeteria, at different price points, used n=232.
External validitySmall, low-stakes candy only, the exact swing size hasn't been shown to hold at higher prices.
Key findings
ParaphrasedChoosing between a 1¢ Hershey's Kiss and a 15¢ Lindt truffle, participants picked the Lindt over the Hershey's by more than two to one (36% vs. 14%). Dropping both prices by exactly one cent, making the Hershey's free, reversed the pattern: the now-free Hershey's was chosen more than twice as often as the Lindt (42% vs. 19%), even though the 14-cent gap between them never changed.
Verbatim (abstract)“After documenting this basic effect, we propose and test several psychological antecedents of the effect, including social norms, mapping difficulty, and affect. Affect emerges as the most likely account for the effect.”
Also worth citing: Pantry Story, a bakery and café at 336 Parramatta Road, replaces the usual request for a jug of water with a self-serve Purezza tap built into the counter, sparkling and still, poured free into a paper cup with no barista to ask. It's a real, uncontrolled example rather than a study, but the same zero-price pull shows up: a water station that costs nothing and requires no request gets used far more freely than the identical water would if it carried even a token price or a wait for staff.
See also: Zero Price Paradox, free doesn't always win. When people suspect a hidden catch (shipping, data, lock-in), zero pricing can backfire instead.
See also: Reciprocity, related but distinct: this is demand jumping at a $0 price point, regardless of who's giving it or how personal it feels; that's the felt obligation triggered specifically by an unsolicited personal gift. A free sample handed to you by name can trigger both at once.
04

Decoy Effect (Asymmetric Dominance)

Why does adding a worse option make you want a completely different one more?

Add a third, clearly worse option and people switch to the option it makes look better, even though nothing about that option changed.

A CLEARLY WORSE OPTION MAKES PEOPLE PICK WHAT BEATS IT AB JUST OPTIONS A AND B SPLIT PREFERENCE ABC A WORSE OPTION C ADDED MORE PICK OPTION B A and B never changed. The decoy just made B’s case easier to see.
Likely mechanismAn easy win against a decoy colours a harder comparison it never actually settled

The psychology. Choices are made by relative comparison, not absolute judgment. A decoy gives an easy, confident basis for one comparison ("B beats C on everything"), and that confidence spills over into the harder comparison (A vs. B) the decoy never actually resolved.

The full write-up: study, numbers, and caveats
The psychology

Choices are made by relative comparison, not absolute judgment. A decoy gives an easy, confident basis for one comparison ("B beats C on everything"), and that confidence spills over into the harder comparison (A vs. B) the decoy never actually resolved.

Where it causes errors

Paying for an upgrade or bundle you wouldn't have chosen from a straight two-option menu, purely because a deliberately uncompetitive third option was added to manufacture comparative confidence.

Where it can help

Mostly protective: recognising when a "worse" third option exists purely to move you off your actual preference is itself the useful skill. Genuine decision-support tools should remove decoys, not add them.

Huber, J., Payne, J. W., & Puto, C. (1982). "Adding Asymmetrically Dominated Alternatives: Violations of Regularity and the Similarity Hypothesis." Journal of Consumer Research, 9(1), 90–98
The original paper, still taught as a foundational violation of rational-choice theory.
StrengthA direct test of "regularity": a core axiom stating that adding an option should never increase another option's market share. The experiments broke that axiom cleanly.
WeaknessUsed hypothetical choice tasks rather than real purchases with real money, so the exact size of the effect needed real-money replication to confirm.
Key findings
Verbatim (abstract)"An asymmetrically dominated alternative is dominated by one item in the set but not by another. Adding such an alternative to a choice set can increase the probability of choosing the item that dominates it."
ParaphrasedAcross 153 students choosing among six product categories (cars, beers, restaurants, lotteries, film, and TV sets), adding the decoy raised the target's average share by 9.2 percentage points between subjects, and from 53% to 56% within the same subjects when the decoy was removed two weeks later.
Also worth citing: Dan Ariely's Economist-subscription classroom demonstration (from Predictably Irrational, 2008) is a popular illustrative case, not a peer-reviewed study.
05

Operational Transparency

Why does watching your food get made make it taste better?

Watching the work happen increases how much people value it, effort becomes visible, and visible effort feels worth more.

WATCHING THE SEARCH HAPPEN MADE THE RESULT FEEL WORTH MORE INSTANT RESULT SHOWN VALUED LESS SAME RESULT, PROCESS SHOWN +26% SALES (DOMINO’S TRACKER) The identical result, either way. Watching the work happen changed what it felt worth.
Likely mechanismVisible effort is read as value, and watching someone work hard invites returning the favour

The psychology. People use a mental shortcut of "effort = value", the labour illusion. Watching work happen also triggers reciprocity: if a provider visibly works hard for me, I feel inclined to value that effort in return, the way identical service can feel worse when it seems effortless.

The full write-up: study, numbers, and caveats
The psychology

People use a mental shortcut of "effort = value", the labour illusion. Watching work happen also triggers reciprocity: if a provider visibly works hard for me, I feel inclined to value that effort in return, the way identical service can feel worse when it seems effortless.

Where it causes errors

Overvaluing an outcome purely because visible effort was performed, including theatrical effort, like a search deliberately slowed down, independent of whether the effort improved anything real.

Where it can help

Genuine transparency corrects the opposite problem too: excellent, fast work being undervalued because the skill behind it is invisible. Showing the real process (not fake theatre) restores an accurate sense of what went into it.

Buell, R. W., & Norton, M. I. (2011). "The Labor Illusion: How Operational Transparency Increases Perceived Value." Management Science, 57(9), 1564–1579
Harvard Business School: the paper that names this specific mechanism.
StrengthFive experiments manipulating whether users saw a visible "searching" process versus an instant result, while holding the actual outcome constant, isolating transparency itself as the cause.
WeaknessSimulated web services in a lab/online setting; whether the effect holds for higher-stakes professional services, where visibly fake effort could backfire once noticed, is left open.
Key findings
ParaphrasedGiven a straight choice between an instant result and one that took longer, participants chose the slower option 62% of the time at a 30-second wait and 63% of the time at 60 seconds when that wait visibly showed work being done, versus just 42% and 23% when the identical wait showed a plain progress bar.
Verbatim (abstract)“When websites engage in operational transparency by signaling that they are exerting effort, people can actually prefer websites with longer waits to those that return instantaneous results.”
Verbatim“Operational transparency is positively associated with perceptions of effort (β = 0.23; p < 0.01), which in turn is positively associated with reciprocity (β = 0.58; p < 0.01), which has a positive association with perceived value (β = 0.68; p < 0.01).”
Also worth citing: Domino's Pizza Tracker (~26% associated sales lift), Kayak's deliberately delayed search results, TurboTax's "checking 350 deductions" messaging.
See also: Tradeoff Transparency, this principle shows that work happened; that one shows what you gave up to get it. Related instinct, different job.
06

Not Enough Choice

Why does taking an option away make people want it more?

Removing an option doesn't just narrow a decision. It can make people want the eliminated choice more, and trust the chooser less.

CHILDREN WANT A CANDY BAR MORE ONCE SOMEONE TAKES IT AWAY BAR SIMPLY UNAVAILABLE RATED LIKE ANY OTHER AN ADULT TAKES IT AWAY RATED MORE APPEALING Same missing bar, either way. Only how it disappeared changed how much kids wanted it.
Likely mechanismLosing an option makes people want it more, simply because it was taken away

The psychology. Psychological reactance: when a freedom you had, including the freedom to choose an option, is threatened or removed, you become motivated to restore it, sometimes wanting the eliminated option more specifically because it was taken away.

The full write-up: study, numbers, and caveats
The psychology

Psychological reactance: when a freedom you had, including the freedom to choose an option, is threatened or removed, you become motivated to restore it, sometimes wanting the eliminated option more specifically because it was taken away.

Where it causes errors

A retailer discontinuing a low-margin but popular option to "simplify" the range can trigger complaints disproportionate to how often that option was actually chosen. The option's mere existence mattered, independent of its sales.

Where it can help

Genuinely necessary restriction benefits from explaining the reasoning. Reactance is triggered by the sense of an arbitrary or unexplained restriction, transparency about why an option was removed defuses it far better than silently narrowing choice. Naming the freedom directly works too, even when nothing is being restricted at all: a real street experiment found that adding a single line, “but you are free to accept or to refuse,” to an otherwise identical request for bus fare raised compliance from about 10% to about 47.5%. Explicitly restoring the freedom a plain request implicitly threatens defuses the reactance before it starts, without narrowing anyone's real choice.

Hammock, T., & Brehm, J. W. (1966). "The Attractiveness of Choice Alternatives when Freedom to Choose Is Eliminated by a Social Agent." Journal of Personality, 34, 546–554
The founding experiment behind psychological reactance theory.
StrengthIsolated a social agent actively removing a choice (versus an option simply never being available) as the key variable, showing reactance is triggered by the act of restriction, not just absence.
WeaknessUsed children and candy bars, a long-replicated paradigm, but confirming the same mechanism scales to adult purchasing and subscription decisions took decades of later work.
Key findings
ParaphrasedChildren rated a candy bar as more attractive after being told an adult had eliminated it as an option, compared to an equally unavailable bar that wasn't actively taken away by someone.
ParaphrasedThe result generalised into psychological reactance theory, now used broadly to explain resistance to advice, warnings, and restricted choices well beyond the original experiment.
Also worth citing: Guéguen, N., & Pascual, A. (2000). “Evocation of Freedom and Compliance: The ‘But You Are Free’ Technique.” Current Research in Social Psychology, 5, 264–270. A real street experiment: asking passersby for coins for a bus fare, a plain request got about 10% compliance, the identical request followed by “but you are free to accept or to refuse” got about 47.5%, naming the freedom explicitly defused the reactance a bare request would otherwise trigger. A later meta-analysis (Carpenter, 2013) found a real but smaller pooled effect across many replications than this original study.
See also: Choice Overload, too little and too much choice can both backfire, for opposite reasons.
07

Choice Overload

Why do more options sometimes mean nobody picks anything at all?

More options can look generous, but past a point they don't help you choose. They make choosing itself the hard part, and many people opt out entirely.

MORE JAM OPTIONS DRAW A CROWD, FEWER PEOPLE BUY 6 JAMS ON THE TABLE USUAL PURCHASE RATE 24 JAMS ON THE TABLE 1/10th THE PURCHASE RATE More jams brought more people to look. Far fewer of them actually bought one.
Likely mechanismComparison cost outweighing the value of more options

The psychology. Comparing many options raises cognitive load and the fear of choosing wrong. Past a certain set size, the mental cost of comparing outweighs the value of having more options, and people are more likely to defer the decision entirely rather than pick imperfectly.

The full write-up: study, numbers, and caveats
The psychology

Comparing many options raises cognitive load and the fear of choosing wrong. Past a certain set size, the mental cost of comparing outweighs the value of having more options, and people are more likely to defer the decision entirely rather than pick imperfectly.

Where it causes errors

A product page offering dozens of variants can convert worse than one offering six, even though the wider range was meant to serve more customers, abundance reads as effort saved for the seller, not the shopper.

Where it can help

Deliberately curating a smaller, well-chosen set, or grouping a large range into simple categories first, increases both satisfaction and completion. "Fewer, better options" is a legitimate service, not a limitation.

Iyengar, S. S., & Lepper, M. R. (2000). "When Choice is Demotivating: Can One Desire Too Much of a Good Thing?" Journal of Personality and Social Psychology, 79(6), 995–1006
The famous "jam study."
Experiment Teardown
502 shoppers passed the booth
Only the number of jams on display changed
6-jam display
6 jams on the table
Stopped, or walked on
24-jam display
24 jams on the table
Stopped, or walked on
Result 6 jams 24 jams Swing
Stopped to look 40% 60% +20
Purchased (of stoppers) 30% 3% −27

More choice drew a crowd, not a sale. The 24-jam display attracted more browsers, but of everyone who stopped, ten times as many bought from the small display as the large one.

Internal validityA real grocery store, real purchases, not stated preference. The two displays rotated hourly across two Saturdays, with the display order counterbalanced by day.
External validityOne product category (jam) at one upscale store; later meta-analyses found the effect genuine but smaller and more context-dependent than this original study alone suggests.
Key findings
Verbatim“Nearly 30% (31) of the consumers in the limited-choice condition subsequently purchased a jar of Wilkin & Sons jam; in contrast, only 3% (4) of the consumers in the extensive-choice condition did so.”
ParaphrasedFollow-up meta-analyses since have found the "too much choice reduces satisfaction or completion" effect is real but variable across contexts, strongest when options are hard to compare or the chooser lacks expertise.
Also worth citing: A 2025 randomised trial of 402 primary care physicians found the opposite pattern in a different context, showing two or three appropriate treatment alternatives (instead of just one) in the electronic health record increased how often physicians chose a more personalised care plan, from 44% to 62%. A genuine boundary condition, not a contradiction: it depends on how comparable the options are and who's choosing, not a universal rule that fewer options always wins. Altinger, G., et al. (2025). "Multiple Suggested Care Alternatives and Decision-Making of Primary Care Physicians: A Randomised Clinical Trial." JAMA Network Open
Also worth reading: The Biases Draining Your Super Could Also Fill It, on how every unintended duplicate super account multiplies the number of fund and insurance decisions a saver would need to make to actually opt out.
08

Zero Price Paradox

Why would a $0 price sell worse than a 99-cent one?

Sometimes free doesn't win. A price of exactly $0 can convert worse than a small positive price, because it invites people to scrutinise the catch instead of feeling the pull of "free."

A HIDDEN SHIPPING FEE MAKES "FREE" LOOK SUSPICIOUS FREE +$ ship ITEM PRICED "FREE" SEEMS LIKE A CATCH, FEWER PEOPLE WANT IT $1 +$ ship SAME ITEM PRICED $1 NOTHING TO QUESTION, MORE PEOPLE WANT IT Same shipping fee, either way. Only the headline price changed what it signalled.
Likely mechanismScrutiny of the “catch” overriding the zero-price emotional pull

The psychology. Zero pricing triggers two competing responses at once: a positive emotional pull (the Zero Price Effect) and heightened scrutiny of what the catch must be. When the incidental costs of "free" are salient, shipping, time, data, a longer commitment, the scrutiny response can override the emotional one, and demand actually drops compared to a small paid price.

The full write-up: study, numbers, and caveats
The psychology

Zero pricing triggers two competing responses at once: a positive emotional pull (the Zero Price Effect) and heightened scrutiny of what the catch must be. When the incidental costs of "free" are salient, shipping, time, data, a longer commitment, the scrutiny response can override the emotional one, and demand actually drops compared to a small paid price.

Where it causes errors

Assuming "free" always converts best, then being confused when a $0 offer underperforms a $1 one, and misdirecting resources chasing an emotional effect that isn't the one actually in play.

Where it can help

Deliberately choosing between framing something as free versus a small nominal price, depending on whether the "catch" is genuinely invisible (a frictionless trial) or something people will reasonably suspect (data collection, an upsell funnel), matching the framing to the honesty of the offer.

Fan, X., Cai, F. C., & Bodenhausen, G. V. (2022). "The Boomerang Effect of Zero Pricing: When and Why a Zero Price Is Less Effective Than a Low Price for Enhancing Consumer Demand." Journal of the Academy of Marketing Science, 50(3), 521–537
Directly tests the boundary condition of the original Zero Price Effect.
StrengthTests when the classic zero-price effect (Shampanier et al., 2007) does not hold, rather than just re-confirming it, a genuine boundary-condition contribution, rarer and more useful than another confirmation study.
WeaknessFour of the five studies frame the incidental cost as effort, a commute, a lengthy survey, writing a resume. Real financial stakes appear only in the field data (a K-12 tutoring market with prices up to about $25 an hour) and genuine physical risk appears only in one study, a hepatitis C vaccine offer. Whether the same reversal holds for a cost framed purely in money, a shipping fee or a deposit, hasn't been tested here.
Key findings
Verbatim (abstract)"Zero pricing triggers both positive affect and cognitive scrutiny of incidental costs; when incidental costs are high, the scrutiny pathway overrides the affective pathway and decreases demand."
ParaphrasedThe reversal specifically shows up when the "free" offer carries non-obvious costs, making this a boundary condition of the original effect, not a contradiction of it.
See also: Zero Price Effect, the two aren't opposites. Zero pricing pulls demand up when there's no catch to suspect; it can push demand down when people sense there must be one. The paradox is knowing which situation you're in.
What the paper actually tested: The main field study analysed 3,502 short-term K-12 tutoring classes in China, comparing online classes (low incidental cost, no commute) against offline classes (high incidental cost, a real commute). Demand for online classes rose the closer the price got to zero. Demand for offline classes followed an inverted U, peaking at a small positive price and dropping at exactly $0, the boomerang this principle is named for. A follow-up study used a hepatitis C vaccine offer to show the same reversal with a risk-based incidental cost instead of a time-based one.
09

Tradeoff Transparency

Why does showing what you'd give up change what you choose?

Showing people exactly what they're giving up to get something, not just what they're getting, leads to better decisions than presenting benefits alone.

SEEING A CARD’S DRAWBACKS FIRST GREW SPEND, CUT CANCELLATIONS SHOWN ONLY THE PERKS BASELINE SPEND AND CANCEL RATE SHOWN THE DRAWBACKS TOO +9.9% MONTHLY SPEND, −20.5% CANCELS Take-up barely moved either way. Seeing the drawbacks first grew spend and cut cancellations.
Likely mechanismNaming the tradeoff forces a real comparison instead of a quick gut reaction

The psychology. Decisions default to being framed around benefits, with costs and tradeoffs often left implicit. Making the tradeoff explicit, "you get this, but you give up that," engages more deliberate comparison, rather than the quick, affect-based judgment a benefits-only pitch invites.

The full write-up: study, numbers, and caveats
The psychology

Decisions default to being framed around benefits, with costs and tradeoffs often left implicit. Making the tradeoff explicit, "you get this, but you give up that," engages more deliberate comparison, rather than the quick, affect-based judgment a benefits-only pitch invites.

Where it causes errors

A plan comparison that only shows what's included lets people commit without registering what they've traded away, until it bites later, a missing feature, a longer lock-in, a cost that only shows up on the first real bill.

Where it can help

This is the constructive end of the whole subject: naming a product's real drawbacks alongside its strengths is a direct route to better decisions, not just a defence against manipulation, and, as the study below found, it doesn't have to cost the business anything to do it.

Buell, R. W., & Choi, M. (2025). “Improving Customer Compatibility with Tradeoff Transparency.” Management Science, 71(2), 1335–1355
A field experiment run inside Commonwealth Bank of Australia's real credit card acquisition funnel, not a lab survey.
Experiment Teardown
393,036 prospective credit card customers
Only whether drawbacks were shown as prominently as benefits differed
Standard marketing
Card pages led with strengths; drawbacks minimised or absent
Tradeoff transparency
Each card's drawbacks shown with equal billing to its strengths
Result Hidden Shown Swing
Monthly spend +9.9%
Cancel rate −20.5%

Naming the downsides didn't cost CBA customers. It earned better ones. Take-up was statistically unchanged, but customers who saw each card's drawbacks went on to spend more, cancel less, and make late payments less often.

Internal validityA large randomised field experiment (n=393,036, roughly split evenly across conditions) measured against real spending, cancellation, and payment records, not self-reported satisfaction.
External validityRun inside one nationwide bank's credit card funnel; whether the same pattern holds for other financial products, or in markets with different disclosure norms, wasn't tested here.
Key findings
Verbatim (abstract)“Although we find tradeoff transparency to have an insignificant effect on acquisition rates, customers who were shown each offering's tradeoffs selected different products than those who were not.”
ParaphrasedA supplementary analysis found early, mixed evidence that this shift moved customers toward cards suited to their financial profile: customers with fewer existing banking products became more likely to pick the lower-rate card under transparency, though not every one of these secondary patterns reached significance.
Verbatim“Customers who were randomly-selected to experience transparency into each offering's tradeoffs, and who moved forward in opening an account, used their cards more intensively, spending 9.9% more per month, and were 10.8% less likely to make late payments on a monthly basis. Furthermore, customers who experienced transparency were 20.5% less likely to cancel their credit cards during the first nine months of their relationships.”
Also worth citing: Buell, R. W., “Commonwealth Bank of Australia: Unbanklike Experimentation,” Harvard Business School Case 619-018, the teaching case documenting how CBA's Behavioural Economics team designed and ran this “Good and the Bad” experiment inside the bank.
See also: Operational Transparency, related instinct, different job: that one shows work happened, this one shows what you gave up to get it.
Tested as an experiment: Does tradeoff transparency work between products, not just within one?, taking CBA's own finding one level up, from comparing cards to comparing ways to pay.
10

Behavioural Labels

Why does calling someone “a voter” make them more likely to vote?

Label someone with the identity behind a behaviour, "a voter," not "someone who votes", and they act to stay consistent with it.

CALLING SOMEONE "A VOTER" GETS THEM TO THE POLLS MORE ASKED ABOUT "VOTING" TURNOUT AS USUAL CALLED "A VOTER" TURNOUT WENT UP Same election, same people asked to vote. Naming the identity, not the action, moved turnout.
Likely mechanismBeing given an identity nudges people to keep acting consistently with it

The psychology. Self-perception theory: people infer their own traits partly by observing their own behaviour, and once an identity is named for them, they work to keep future behaviour consistent with it. A noun ("a voter") invokes a stable identity; a verb ("voting") only describes a one-off action. Identities pull harder on what happens next than actions do.

The full write-up: study, numbers, and caveats
The psychology

Self-perception theory: people infer their own traits partly by observing their own behaviour, and once an identity is named for them, they work to keep future behaviour consistent with it. A noun ("a voter") invokes a stable identity; a verb ("voting") only describes a one-off action. Identities pull harder on what happens next than actions do.

Where it causes errors

The same lever runs in reverse: labelling someone with a negative identity, "you're lazy," "you're bad with money", can lock in the very behaviour you meant to change, since people conform to identities imposed on them just as readily as ones they choose for themselves.

Where it can help

Deliberately using a positive identity label, "you're someone who shows up," "you're a voter," "you're a saver", reliably outperforms straightforward exhortation to act. It's one of the most replicated, low-cost, genuinely ethical levers in the field: it works by affirming an identity someone already holds or aspires to, not by concealing anything.

Bryan, C. J., Walton, G. M., Rogers, T., & Dweck, C. S. (2011). "Motivating Voter Turnout by Invoking the Self." PNAS, 108(31), 12653–12656
Field experiments across two real state elections, not just survey responses.
StrengthMeasured actual turnout against official state voting records, not self-report, a far stronger design than most attitude studies, across two real elections rather than one lab session.
WeaknessThe turnout gains here are unusually large for this literature: 10.9 to 13.7 percentage points across the two real elections, which the authors themselves call among the largest effects ever recorded on objectively measured voter turnout. An effect that size is worth treating carefully. It's the kind of result that doesn't always hold up at population scale or in later replications.
Key findings
Verbatim (abstract)“[T]he personal-identity phrasing significantly increased interest in registering to vote (experiment 1) and, in two statewide elections in the United States, voter turnout as assessed by official state records (experiments 2 and 3).”
Verbatim (abstract)“These results provide evidence that people are continually managing their self-concepts, seeking to assume or affirm valued personal identities.”
Also worth citing: Miller, Brickman & Bolen (1975): labelling children "neat and tidy" cut littering far more effectively than repeatedly telling them to be neat and tidy, with the effect still holding two weeks later. An older, foundational demonstration of the same mechanism.
See also: Framing Effect, easy to confuse with this one, but different: that's about how the same fact is worded ("75% lean" vs. "25% fat"); this is about naming a person's identity ("a voter").
11

Framing Effect

Why does “90% survive” sound so much better than “10% die”, for the same surgery?

The same fact, framed two different ways, changes how people feel about it, even though the underlying information is identical.

SHOPPERS RATE THE SAME BEEF BETTER LABELLED "75% LEAN" 25% LABELLED "25% FAT" RATED LESS FAVOURABLY 75% LABELLED "75% LEAN" RATED MORE FAVOURABLY Identical beef, both times. Only the label’s frame changed how good it seemed.
Likely mechanismThe same fact feels different depending on whether it's framed as a gain or a loss

The psychology. This is attribute framing: people don't process factual information neutrally. Whether a fact is stated in positive terms ("75% lean") or the mathematically identical negative terms ("25% fat") changes the immediate emotional read, even though nothing about the product changed.

The full write-up: study, numbers, and caveats
The psychology

This is attribute framing: people don't process factual information neutrally. Whether a fact is stated in positive terms ("75% lean") or the mathematically identical negative terms ("25% fat") changes the immediate emotional read, even though nothing about the product changed.

Where it causes errors

Choosing a product based on a positively-framed label, low-fat, "90% success rate", without registering the equivalent negative framing (still has fat, still fails one in ten). The decision shifts based on labelling alone, not on any real change in the thing being labelled.

Where it can help

Well-designed labels that show both framings, or use a single standardised comparable metric, nutrition panels, energy ratings, help people see past the framing to the actual fact. Honest labelling is one of the most effective, low-cost nudges towards better decisions that exists.

Levin, I. P., & Gaeth, G. J. (1988). "How Consumers Are Affected by the Framing of Attribute Information Before and After Consuming the Product." Journal of Consumer Research, 15(3), 374–378
The classic "75% lean / 25% fat" framing study.
StrengthTests the same framing effect both before and after direct product experience (tasting the meat), most framing studies only measure the pre-experience judgment, so this one shows whether real information corrects the bias.
WeaknessSingle product category and a simple binary framing (lean/fat), more complex labels, like multi-attribute nutrition panels or financial disclosures, may not behave identically.
Key findings
Verbatim (abstract)“The consumers’ evaluations were more favorable toward the beef labeled ‘75% lean’ than that labeled ‘25% fat.’”
ParaphrasedThe framing effect was strongest before consumers tasted the product and measurably weaker afterwards, direct experience partially, but not fully, corrected the labelling bias.
See also: Behavioural Labels, related but different: this is about how the same fact is worded; that's about naming a person's identity.
Decoded on The Science Behind: Why do you have to be “into double denim” to win a better savings rate?, where a savings account's ordinary consolation rate gets relabelled a “win.”
12

Illusion of Explanatory Depth

Why can't you actually explain how a zipper works, even though you're sure you understand it?

You feel certain you understand how something works, a zipper, a bill, a policy, right up until you're asked to actually explain the mechanism, step by step.

EXPLAINING A DEVICE OUT LOUD DROPS CONFIDENCE FAST RATES OWN UNDERSTANDING HIGH BEFORE EXPLAINING TRIES TO EXPLAIN HOW IT WORKS CONFIDENCE DROPS SHARPLY The device never changed. Trying to explain it out loud was what exposed the gap.
Likely mechanismFamiliarity gets mistaken for actually understanding how something works

The psychology. This is a specific failure of metacognition, not general overconfidence: people confuse familiarity, having seen or used something many times, with having a working causal model of it. The mental model feels complete because it's never actually run end-to-end, only skimmed. The act of generating a real explanation is what exposes the gap; simply asking “do you understand this?” doesn't.

The full write-up: study, numbers, and caveats
The psychology

This is a specific failure of metacognition, not general overconfidence: people confuse familiarity, having seen or used something many times, with having a working causal model of it. The mental model feels complete because it's never actually run end-to-end, only skimmed. The act of generating a real explanation is what exposes the gap; simply asking “do you understand this?” doesn't.

Where it causes errors

People form strong opinions on things, a tax policy, a medical treatment, a technical claim, using a mental model they've never actually tested by explaining it aloud. Confidence and real understanding are only weakly correlated until someone is forced to walk through the mechanism, at which point both the confidence and, sometimes, the underlying opinion tend to soften.

Where it can help

The fix is built into the mechanism itself: asking someone (or yourself) to explain a claim, plan, or product mechanistically, step by step, before committing is one of the few debiasing techniques that reliably works on contact. It doesn't require distrust or scepticism, just the act of trying to explain.

Rozenblit, L., & Keil, F. (2002). “The Misunderstood Limits of Folk Science: An Illusion of Explanatory Depth.” Cognitive Science, 26(5), 521–562
Twelve studies establishing the core effect across knowledge domains.
StrengthA within-subject before/after design across twelve separate studies and several knowledge domains (mechanical devices, natural phenomena, procedures, narratives), measuring the actual change in self-rated understanding caused by attempting an explanation, not just a single static confidence score.
WeaknessUnderstanding is self-rated on a numeric scale before and after explaining, a subjective measure of a subjective experience. Some of the measured drop could reflect participants adjusting their answer to match what they expect after being put on the spot, rather than a pure recalibration of belief.
Key findings
Verbatim“Nearly all participants showed drops in estimates of what they knew when confronted with having to provide a real explanation, answer a diagnostic question, and compare their understanding to an expert description.”
Verbatim“[W]e found no decrease in knowledge ratings over time with procedures or narratives, and a significantly smaller drop with facts.”
Also worth citing: Fernbach, P. M., Rogers, T., Fox, C. R., & Sloman, S. A. (2013), “Political Extremism Is Supported by an Illusion of Understanding,” Psychological Science, 24, 939–946, found that asking people to explain a policy mechanistically moderated their stated position. Honest caveat: a 2021 preregistered replication attempt (published in Cognition) failed to closely reproduce that political-moderation extension specifically, the core illusion-of-explanatory-depth effect itself remains well-replicated, but this particular applied extension should be treated as unsettled.
13

Social Proof

Why do you trust a restaurant more because it has a line outside?

When people are unsure what to do, they copy what everyone else appears to be doing, treating popularity itself as evidence of quality.

SHOWING DOWNLOAD COUNTS MAKES HITS LESS PREDICTABLE DOWNLOAD COUNT HIDDEN SUCCESS TRACKS SONG QUALITY DOWNLOAD COUNT SHOWN ONE EARLY HIT SNOWBALLS INTO A HIT The best songs rarely finished last. But who saw early counts decided almost everything else.
Likely mechanismOther people's visible choices are read as evidence about what's actually good

The psychology. In ambiguous situations, people use others' visible behaviour as information, a mental shortcut that's often correct (if 500 people chose this, it's probably not terrible) but that breaks down the moment the crowd's choice was itself shaped by something other than quality, like which option simply got noticed first.

The full write-up: study, numbers, and caveats
The psychology

In ambiguous situations, people use others' visible behaviour as information, a mental shortcut that's often correct (if 500 people chose this, it's probably not terrible) but that breaks down the moment the crowd's choice was itself shaped by something other than quality, like which option simply got noticed first.

Where it causes errors

A product, app, or song can become genuinely more popular purely because it was shown as popular early on, not because it was actually better, and the effect compounds: more visible popularity draws more followers, regardless of the underlying quality gap to the alternatives.

Where it can help

Genuine social proof, showing real, verified numbers (verified reviews, real completion rates), helps people make faster, reasonably good decisions under real uncertainty. The line to watch is whether the number shown is real and representative, or manufactured to engineer a herd.

Spotted in the wild

Two examples from the same real checkout flow, a domain registrar upselling an add-on, then asking for payment on the very next screen. Both lean on the identical mechanism this principle names: let a crowd's size stand in for a quality judgment the buyer hasn't made themselves.

A GoDaddy domain checkout upsell for "Full Domain Protection" at $13.95 per year, with a badge reading CHOSEN BY OVER 225,000 CUSTOMERS EACH MONTH highlighted

GoDaddy, domain checkout upsell. “CHOSEN BY OVER 225,000 CUSTOMERS EACH MONTH” sits directly under a $13.95/yr add-on marked “RECOMMENDED,” right where a buyer is deciding whether to add it. Captured 2026-08-26.

The same GoDaddy checkout's order summary screen, showing a 4.4 out of 5 Trustpilot rating from 139,923 reviews, highlighted

Same checkout, order summary screen. A 4.4-star Trustpilot rating and “139,923 reviews” count appear directly above the “Complete Purchase” button, the last thing shown before payment. Captured 2026-08-26.

Both numbers are GoDaddy's own claims, not independently verified figures, the same honesty check this site applies to every citation. Showing them here demonstrates the tactic in current, active use, not that it worked on these particular screens. The actual evidence that the mechanism is real is the cited study below.

Salganik, M. J., Dodds, P. S., & Watts, D. J. (2006). “Experimental Study of Inequality and Unpredictability in an Artificial Cultural Market.” Science, 311(5762), 854–856
The “MusicLab” experiment.
StrengthA genuine live field experiment with 14,341 real participants downloading real, previously unknown songs, some given no information about others' choices, others shown download counts, letting the researchers isolate the causal effect of visible social proof from actual song quality.
WeaknessThe songs were from unsigned bands with no prior track record, and downloads (not payment or lasting engagement) were the measured outcome, a lower-stakes, lower-commitment behaviour than many of the real purchase or voting decisions the effect gets applied to.
Key findings
Verbatim (abstract)“Increasing the strength of social influence increased both inequality and unpredictability of success.”
Verbatim (abstract)“Success was also only partly determined by quality: The best songs rarely did poorly, and the worst rarely did well, but any other result was possible.”
See also: Social Norm, related but distinct: this is about copying others' visible choices under uncertainty; that's about which specific message (what's common vs. what's approved) actually changes behaviour, and when.
Decoded on The Science Behind: How does Up Bank's behavioural design pay off for both the bank and its customers?, where a real-time feed of other people winning a promotion runs alongside a visibly draining prize pool.
14

Social Norm (Descriptive vs. Injunctive)

Why can telling people “most don't do this” backfire and make more of them do it?

Two different messages both count as “the norm”, what most people actually do (descriptive), and what most people approve or disapprove of (injunctive), and mixing them up can backfire.

A LITTERED PARKING STRUCTURE MADE PEOPLE LITTER MORE CLEAN PARKING STRUCTURE 11% LITTERED THE HANDBILL ALREADY-LITTERED STRUCTURE 41% LITTERED THE HANDBILL Same handbill, same parking structure. Seeing existing litter nearly quadrupled the littering rate.
Likely mechanismA norm only changes behaviour if it's front of mind at the moment of choosing

The psychology. Norm messaging only works when the norm is “in focus” at the moment of the decision, simply being true isn't enough. And the two types of norm pull differently: descriptive norms (“most people do X”) tell you what's typical; injunctive norms (“most people approve of X”) tell you what's sanctioned. Pointing at the wrong one, or accidentally highlighting that undesirable behaviour is common, can increase the very behaviour you meant to reduce.

The full write-up: study, numbers, and caveats
The psychology

Norm messaging only works when the norm is “in focus” at the moment of the decision, simply being true isn't enough. And the two types of norm pull differently: descriptive norms (“most people do X”) tell you what's typical; injunctive norms (“most people approve of X”) tell you what's sanctioned. Pointing at the wrong one, or accidentally highlighting that undesirable behaviour is common, can increase the very behaviour you meant to reduce.

Where it causes errors

A campaign that says “too many people still litter here” is factually honest but behaviourally counterproductive. It broadcasts a descriptive norm (littering is common) that undercuts the injunctive message (littering is disapproved of), and can increase littering among people who weren't planning to litter at all.

Where it can help

Correctly identifying and stating the actual descriptive norm, “the majority of guests in this room reuse their towels”, reliably outperforms appeals based purely on abstract values like environmental protection, without requiring any deception: the norm has to be real.

Cialdini, R. B., Reno, R. R., & Kallgren, C. A. (1990). “A Focus Theory of Normative Conduct: Recycling the Concept of Norms to Reduce Littering in Public Places.” Journal of Personality and Social Psychology, 58(6), 1015–1026
Study 1: the parking-structure littering experiment.
Experiment Teardown
139 drivers
Only the visible state of the parking structure changed
Clean environment
No litter visible on arrival
Found a handbill on the windshield
Littered environment
Litter already scattered around
Found the same handbill
Result Clean Littered Swing
Littered the handbill 11% 41% +30

The environment did the persuading. Seeing existing litter nearly quadrupled the littering rate, same handbill, same parking structure, just a different visible norm.

Internal validityA real, physically altered environment, not a survey, with real littering behaviour covertly observed, not self-reported.
External validityOne behaviour (littering), tested across several different physical settings within the same paper: two hospital parking garages, a walkway, dormitory mailboxes, and a library car park. The paper's own focus theory predicts the mechanism should generalise to other behaviours, which needed separate confirmation elsewhere.
Key findings
ParaphrasedAcross five experiments, a relevant norm (against littering) only shaped behaviour when something in the environment made that norm psychologically salient at the moment of the decision, the same norm, unfocused, had little effect.
ParaphrasedMaking people aware of a norm that runs counter to what's wanted (pointing out that a space is already littered) could increase the undesirable behaviour, by making the descriptive norm salient instead of the injunctive one.
Also worth citing: Goldstein, N. J., Cialdini, R. B., & Griskevicius, V. (2008), “A Room with a Viewpoint: Using Social Norms to Motivate Environmental Conservation in Hotels,” Journal of Consumer Research, 35(3), 472–482, two real hotel field experiments found a sign stating the majority of guests in that specific room reuse their towels outperformed a standard environmental-appeal sign.
See also: Social Proof, related but distinct: that's about copying others' visible choices under uncertainty; this is about which specific norm message actually changes behaviour, and when.
15

Compromise Effect

Why does the middle option win, just for being in the middle?

Add a middle option to a lineup of two, and people disproportionately pick the middle one, not because it's objectively best, but because it's the easiest choice to defend.

ADDING A THIRD, PRICIER CAMERA GREW THE MIDDLE ONE’S SHARE AB ONE OF TWO CAMERAS 50% CHOSE IT ABC NOW THE MIDDLE OF THREE 57% CHOSE IT Same camera, same price, same features. Only its position in the lineup changed.
Likely mechanismExtremeness aversion, avoiding the risk of looking wasteful or cheap

The psychology. Extreme options carry a hidden cost: choosing the priciest option risks looking wasteful, choosing the cheapest risks looking cheap. A middle option lets someone avoid both criticisms at once, especially when they expect to have to justify the choice to someone else, the decision reasons itself, rather than requiring weighing the actual attributes.

The full write-up: study, numbers, and caveats
The psychology

Extreme options carry a hidden cost: choosing the priciest option risks looking wasteful, choosing the cheapest risks looking cheap. A middle option lets someone avoid both criticisms at once, especially when they expect to have to justify the choice to someone else, the decision reasons itself, rather than requiring weighing the actual attributes.

Where it causes errors

A three-tier pricing menu reliably pushes disproportionate share towards the middle tier regardless of whether it's genuinely the best value, the same product framed as the extreme of only two options sells differently than when it's repositioned as the safe middle of three.

Where it can help

Recognising this is mostly a defensive skill: before choosing “the middle one,” ask whether you'd still pick it if it were relabelled the cheapest or the most expensive option in a differently bounded set, if the answer changes, the compromise position, not the actual value, made the decision.

Simonson, I. (1989). “Choice Based on Reasons: The Case of Attraction and Compromise Effects.” Journal of Consumer Research, 16(2), 158–174
The paper that names and formally tests the compromise effect.
StrengthDirectly tested whether the compromise effect strengthens under “need for justification”, participants told they'd have to explain their choice to someone else showed a larger compromise effect than those who didn't, directly supporting the reason-based account rather than just documenting that the effect exists.
WeaknessUses hypothetical, forced-choice lab scenarios rather than real purchases with real money on the line, the size of the effect in real, high-stakes purchases needed further field replication.
Key findings
Verbatim“[T]he market shares of alternatives in the TV, apartment, calculator, mouthwash, and calculator battery categories were, on average, 17.5 percent larger when they were compromise brands than when they were not.”
ParaphrasedThe attraction effect was measurably stronger among participants who expected to justify their decision to someone else. An earlier pilot study found the same pattern for the compromise effect, but that difference wasn't statistically significant in the main study, so the compromise result is less settled than the attraction one.
See also: Decoy Effect and Similarity Effect, collectively known as “context effects”: how a choice set is built changes which option wins, independent of anyone's actual preferences.
16

Similarity Effect

Why does adding a near-identical option make the best one look worse?

Add a second option that's very similar to an existing one, and the two split attention and share between them, leaving a more distinct alternative relatively better off, even though nothing about its own quality changed.

ADDING A NEAR-TWIN TO OPTION A GROWS OPTION B’S SHARE TWO OPTIONS, A AND B B HOLDS HALF THE MARKET A GAINS A NEAR-TWIN, B UNCHANGED B’S SHARE OF THE MARKET GROWS A and its twin split the same slice. B never had to change to win more of it.
Likely mechanismTwo near-identical options split the same pool of support instead of adding to it

The psychology. People compare options partly by how easy they are to tell apart. Two near-identical options compete directly with each other for the same slice of consideration, effectively dividing, not adding to, their combined share, while a more distinctive option doesn't have to fight that internal battle and keeps its share intact.

The full write-up: study, numbers, and caveats
The psychology

People compare options partly by how easy they are to tell apart. Two near-identical options compete directly with each other for the same slice of consideration, effectively dividing, not adding to, their combined share, while a more distinctive option doesn't have to fight that internal battle and keeps its share intact.

Where it causes errors

A retailer or platform that adds a new option nearly identical to an existing bestseller can inadvertently cannibalise that bestseller's share and hand a relative advantage to a totally different, less similar product, the near-duplicate doesn't steal share evenly, it mostly steals from its closest twin.

Where it can help

Deliberately differentiating similar options, rather than letting them cluster, helps people actually compare on the attributes that matter, instead of getting stuck between two choices that are hard to tell apart for reasons that have nothing to do with quality.

Tversky, A. (1972). “Elimination by Aspects: A Theory of Choice.” Psychological Review, 79(4), 281–299
The foundational paper identifying similarity, attraction, and compromise as distinct context effects.
StrengthFormally derives and experimentally tests the similarity effect as a specific, falsifiable violation of “regularity” and “independence of irrelevant alternatives”, foundational axioms of rational choice theory, using genuine probabilistic choice experiments, not just informal demonstrations.
WeaknessAs with much choice-theory work from this era, the experimental choice sets are abstract and hypothetical rather than real products in a real market, later applied work (Huber & Puto, 1983) tested the effect with real product categories to confirm it held outside the lab.
Key findings
ParaphrasedThe model, and the accompanying experiments,showed that adding a close substitute for an existing option reduces that option's relative share by more than simple probability models predict, because the two similar options compete directly with each other for the same underlying preference.
ParaphrasedThe paper formally derives and tests the similarity effect itself, using classic thought experiments (choosing between records, travel agencies, trips to Paris or Rome) to show that adding a near-duplicate option violates basic axioms of rational choice like regularity and the constant-ratio rule. It doesn't name “attraction” or “compromise” effects: those specific effects and terms came later, from Huber, Payne, and Puto (1982) and Simonson (1989), building on the elimination-by-aspects framework this paper introduces.
See also: Decoy Effect and Compromise Effect, together, the three classic “context effects”: similarity splits share between near-twins, attraction (decoy) boosts a dominant option via a deliberately worse one, and compromise favours the safe middle. Same underlying lesson: how a menu is built shapes the outcome as much as what's actually on it.
17

Goal Gradient Effect

Why do you try harder for a loyalty stamp you're 8 steps into, not 2?

Effort and motivation increase the closer you get to a goal, which is why a loyalty card with two stamps already filled in gets used faster than a blank one, even for the exact same reward.

CUSTOMERS WHO PERCEIVE THEY’RE CLOSER BUY THE SAME 10 FASTER 0 OF 10 STAMPED 10 PURCHASES TO GO STEADY PACE, EVENLY SPACED 2 OF 12 STAMPED STILL 10 PURCHASES TO GO PACE PICKS UP NEAR THE END Both cards need exactly 10 more purchases. Only the one that felt closer sped people up.
Likely mechanismFeeling closer to a finish line speeds people up, even when the real distance hasn't changed

The psychology. Originally observed in animals running a maze, they move faster the closer they get to the food, the same gradient shows up in humans pursuing any goal with a visible finish line. Perceived proximity to completion, not just objective distance, drives the acceleration: this is why pre-filling a loyalty card with “free” stamps that weren't actually earned still speeds people up, even though the real number of purchases needed hasn't changed.

The full write-up: study, numbers, and caveats
The psychology

Originally observed in animals running a maze, they move faster the closer they get to the food, the same gradient shows up in humans pursuing any goal with a visible finish line. Perceived proximity to completion, not just objective distance, drives the acceleration: this is why pre-filling a loyalty card with “free” stamps that weren't actually earned still speeds people up, even though the real number of purchases needed hasn't changed.

Where it causes errors

A 10-stamp loyalty card gives someone the same effective distance to go as an 8-purchase card that starts with 2 stamps pre-filled, but the second version reliably accelerates purchasing, exploiting a felt sense of progress rather than the real requirement.

Where it can help

Genuine progress bars, showing someone how close they actually are to finishing a form, a course, or a savings goal, use the identical mechanism to help people follow through on things they already want to do, with nothing fabricated about the progress shown.

Kivetz, R., Urminsky, O., & Zheng, Y. (2006). “The Goal-Gradient Hypothesis Resurrected: Purchase Acceleration, Illusionary Goal Progress, and Customer Retention.” Journal of Marketing Research, 43(1), 39–58
The real coffee-card field study.
StrengthCombined a real café field experiment (real customers, a real 10-coffee loyalty card, real purchase records over time) with a separate online field study, converging evidence from two different real-world reward programmes, not just one lab task.
WeaknessThe main purchase-acceleration pattern comes from real customer card records rather than a randomised experiment, since people chose for themselves whether to join and use the loyalty programme. A separate test of the 10-stamp versus 12-stamp cards was randomly assigned, but with only 108 customers at a single café.
Key findings
ParaphrasedCafé customers bought coffee at an increasing rate the closer they got to their free 10th coffee. In a separate randomised test, 108 customers given a 12-stamp card with 2 stamps already filled in finished their 10 required purchases in 12.7 days on average, 20% faster than the 15.6 days it took customers given a blank 10-stamp card, for the identical real distance to the reward.
ParaphrasedAfter earning a reward, participants' effort dropped and then re-accelerated as they approached the next reward, the gradient resets and repeats with each new goal, rather than motivation simply declining over time.
See also: Medium Maximisation, a related but different reward-programme mechanism: this one is about accelerating effort near a finish line; that one is about chasing the reward token itself past the point where it still tracks real value.
The finish line doesn't have to be external: Law of Round Numbers is the same acceleration effect aimed at a number someone sets for themselves rather than a card a business hands them.
Experiment built on this: Does showing progress through a home loan appointment booking increase completion?, testing the exact mechanism proposed above on a real multi-step lending form.
Decoded on The Science Behind: Why does GoalSaver's 4.75% bonus feel like money you'd be losing, not money you haven't earned yet?, where CommBank's Goal Tracker breaks a savings target into weekly milestones using this exact acceleration effect.
Also decoded on The Science Behind: How does Up Bank's behavioural design pay off for both the bank and its customers?, where every Saver renders its own visible progress bar against its real target, right in the account list.
Special report built on this: Breaking a Form Into Steps Changes How It Feels. It Doesn't Always Change How It Performs., where this exact mechanism explains the Progress Flow form pattern, and a real head-to-head test finds it can backfire on a phone.
18

Medium Maximisation

Why do people chase airline miles worth less than the cash they spent to get them?

Give people a token, points, miles, stamps, standing in for a reward, and they'll sometimes work to maximise the token itself, even in choices where doing so makes no material sense.

ADDING POINTS MORE THAN DOUBLES WHO PICKS THE HARDER TASK NO POINTS SHOWN 25% PICK THE HARDER TASK PTS POINTS ADDED TO THE TASK 53% PICK THE HARDER TASK Their taste for the ice cream barely moved. The points did the deciding.
Likely mechanismChasing more points crowds out what the points are actually worth

The psychology. A medium (points, miles) is supposed to be a transparent stand-in for the value it can be traded for. But once effort is being tracked in the medium's units, people can lose the thread back to real value, chasing “more points” the way they'd chase more money, even in situations where the points-maximising option is worth objectively less once redeemed.

The full write-up: study, numbers, and caveats
The psychology

A medium (points, miles) is supposed to be a transparent stand-in for the value it can be traded for. But once effort is being tracked in the medium's units, people can lose the thread back to real value, chasing “more points” the way they'd chase more money, even in situations where the points-maximising option is worth objectively less once redeemed.

Where it causes errors

Choosing the option that earns more loyalty points over one that earns fewer points but is worth objectively more once you factor in redemption value. The token becomes the goal, not the proxy it was meant to be, especially when the point-earning option makes an otherwise inferior choice feel “certain” or advantageous.

Where it can help

There isn't a genuinely benign version of exploiting this in others. The honest use is entirely defensive: converting any points-based choice back into its real cash-equivalent value before deciding, rather than comparing the number of points directly.

Hsee, C. K., Yu, F., Zhang, J., & Zhang, Y. (2003). “Medium Maximization.” Journal of Consumer Research, 30(1), 1–14
The paper that names the effect: Study 1 is the ice cream/task experiment.
Experiment Teardown
96 students recruited
Only a points system was added
Control condition
6m6 minvanilla
7m7 minpistachio
Chose a task
Medium condition
6m6 min60 ptsvanilla
7m7 min100 ptspistachio
Chose a task
Result Control Medium Swing
Chose long task ~25% ~53% +28
Preferred pistachio ~27% ~29% +2

Choice moved, taste didn't. In the control condition, task choice and flavour preference tracked closely together. Add points, and choice of the long task roughly doubled, while real preference for pistachio barely shifted at all.

Internal validityIdentical recruitment and an identical follow-up preference question in both conditions, though the paper doesn't state how people were assigned to each group.
External validityOne campus, one reward. The authors re-ran the same logic in a bank CD scenario (Study 2) and found the same pattern.
Key findings
ParaphrasedIn Study 1, participants offered a points-based reward chose the longer, more point-heavy task at a significantly higher rate than participants offered the identical flavour tradeoff directly (χ² = 7.43, p < .01), even though a separate question showed their actual flavour preference didn't differ between the two groups.
ParaphrasedThe effect traces to specific illusions the medium creates, for example, a medium can make an option feel more certain or more proportionally rewarding than it actually is, rather than simple confusion about how the redemption maths works.
See also: Goal Gradient Effect, related reward-programme mechanism, different failure mode: that one is about accelerating effort as a goal nears; this one is about losing track of real value while chasing the token itself.
See also: Symbolic Rewards, related but different: this is a token that stands in for a real reward and gets over-optimised for its own sake; that one is a token with no exchange value at all, that still motivates through recognition alone.
19

Noise

Why do the same judge, the same case, on different days, land on different verdicts?

Ask the same expert to judge the same thing twice, on different days, and you'll often get two different answers, not because they're biased in a consistent direction, but because judgment itself is noisy.

THE SAME JUDGE SCORES THE SAME WINE GOLD, THEN BRONZE TASTING 1, SAME BOTTLE GOLD SCORED TOP MARKS TASTING 2, SAME BOTTLE BRONZE SCORED LOW MARKS Same bottle, same judge, poured twice. Only 1 in 10 judges scored it the same way twice.
Likely mechanismUnwanted, inconsistent scatter in judgments, not a bias pointing the same way every time

The psychology. Bias and noise are different failure modes: bias is a consistent, predictable skew in one direction (a scale that always reads 2kg heavy); noise is scatter, the same judge, judging the same case, lands in different places depending on factors that shouldn't matter, like mood or simply the day. Because it doesn't point in a consistent direction, noise is far harder to notice than bias, and can be just as costly.

The full write-up: study, numbers, and caveats
The psychology

Bias and noise are different failure modes: bias is a consistent, predictable skew in one direction (a scale that always reads 2kg heavy); noise is scatter, the same judge, judging the same case, lands in different places depending on factors that shouldn't matter, like mood or simply the day. Because it doesn't point in a consistent direction, noise is far harder to notice than bias, and can be just as costly.

Where it causes errors

Two identical insurance claims, loan applications, or court sentences can get meaningfully different outcomes purely because of who reviewed them and on what day, an inconsistency that's invisible in any single case, but shows up starkly the moment the same case is quietly shown to the same expert twice.

Where it can help

Structuring judgment, breaking a decision into independently scored components, aggregating multiple independent judges' scores rather than relying on one, using checklists and decision rules, measurably reduces noise without requiring anyone to be less expert or less caring about the decision.

Hodgson, R. T. (2008). “An Examination of Judge Reliability at a Major U.S. Wine Competition.” Journal of Wine Economics, 3(2), 105–113
The wine-judges reliability study.
StrengthAn unusually clean real-world reliability test: each expert judge unknowingly re-tasted the identical wine, poured from the same bottle, multiple times within the same flight, across four years of a real, major competition, rather than a constructed lab scenario.
WeaknessWine scoring is a genuinely difficult sensory-judgment task with inherently high natural variance (palate fatigue, order effects), some caution is warranted in assuming the same magnitude of inconsistency applies identically to very different kinds of expert judgment, like medical diagnosis or legal sentencing, even though the general mechanism is the same.
Key findings
Verbatim (abstract)“Between 65 and 70 judges were tested each year. About 10 percent of the judges were able to replicate their score within a single medal group.”
Verbatim (abstract)“Another 10 percent, on occasion, scored the same wine Bronze to Gold.”
Also worth citing: Kahneman, D., Sibony, O., & Sunstein, C. R. (2021), Noise: A Flaw in Human Judgment (Little, Brown Spark), the book that popularised “noise” as a concept distinct from bias, and that specifically cites Hodgson's wine-judge findings as a real-world illustration. It's a trade book rather than a peer-reviewed study, so treat it as a synthesis and framework, not itself the primary evidence.
20

Price Transparency

Why does hiding the total in fees change what people actually buy?

Showing the full, all-in price upfront, instead of a lower headline number with fees added later, changes not just what people buy, but how much and how carefully.

SHOWING THE FULL PRICE UPFRONT MAKES SHOPPERS BUY LESS $40 +fees LOW PRICE, FEES ADDED LATER MORE FULL-PRICE ITEMS PICKED $52 all-in FULL PRICE SHOWN UPFRONT FEWER, CHEAPER ITEMS PICKED Trading down to cheaper items alone caused over a quarter of the whole revenue drop.
Likely mechanismLimited attention anchoring on the headline number

The psychology. Splitting a price into a low headline figure plus separate add-on fees exploits limited attention: people anchor on and remember the first, smaller number, and partially discount or forget the fees layered on after. Making the full price salient from the start removes that gap, forcing the real cost to be weighed against the real decision from the outset rather than after commitment has already begun.

The full write-up: study, numbers, and caveats
The psychology

Splitting a price into a low headline figure plus separate add-on fees exploits limited attention: people anchor on and remember the first, smaller number, and partially discount or forget the fees layered on after. Making the full price salient from the start removes that gap, forcing the real cost to be weighed against the real decision from the outset rather than after commitment has already begun.

Where it causes errors

Drip pricing, a cheap-looking base price with mandatory fees revealed at checkout, reliably increases what people actually pay compared to the same total price shown upfront, precisely because it exploits the anchoring/attention gap rather than offering any real value; it's a major reason many jurisdictions now require all-in pricing by law.

Where it can help

Showing the complete, real price from the very first screen doesn't just protect consumers. The underlying research also found it changes purchase behaviour towards products people actually value more, since decisions get made against real cost rather than a distorted one.

Blake, T., Moshary, S., Sweeney, K., & Tadelis, S. (2021). “Price Salience and Product Choice.” Marketing Science, 40(4), 619–636
A real field experiment on StubHub's live marketplace.
StrengthA genuine large-scale field experiment run on StubHub's actual live marketplace, with real buyers and real money, comparing the platform's usual drip-priced checkout against an all-in, full-price-upfront version, with real transaction and click-stream data, not a survey or lab simulation.
WeaknessThe experiment is specific to one marketplace (ticket resale) with an unusually large gap between the salient sticker price and the real all-in cost; how large the effect is in categories where fees are proportionally smaller wasn't directly tested in this study.
Key findings
Verbatim (abstract)“Making the full purchase price salient to consumers reduces both the quality and quantity of goods purchased.”
Verbatim (abstract)“The effect of salience on quality accounts for at least 28% of the overall revenue decline.”
Also worth citing: Morwitz, V. G., Greenleaf, E. A., & Johnson, E. J. (1998), “Divide and Prosper: Consumers' Reactions to Partitioned Prices,” Journal of Marketing Research, 35(4), 453–463, the earlier classic study establishing the underlying mechanism: splitting a price into a base plus a separate mandatory surcharge lowers what people recall as the total cost and increases demand, compared to one combined price.
21

Salience

Why does the one bold number on a bill matter more than the total?

The option, number, or detail that visually or emotionally stands out gets weighted far more heavily in a decision than its actual importance justifies, simply because it's the thing you noticed.

PEOPLE PICK THE EYE-CATCHING OPTION EVEN AT THE SAME ODDS 50%50% BOTH OPTIONS PLAIN, SAME ODDS CHOSEN EQUALLY OFTEN 50%50% ONE OPTION MADE TO STAND OUT PICKED FAR MORE OFTEN Same options, same odds, both times. What stood out visually decided it.
Likely mechanismWhatever visually stands out gets weighed more heavily than what actually matters more

The psychology. Attention is a scarce resource, and choice requires comparing multiple attributes across multiple options at once, a task people can't do exhaustively. Salience theory formalises what fills that gap: attributes and outcomes that stand out (because they're extreme, unusual, or visually prominent relative to the immediate comparison set) get disproportionate decision weight, while equally or more important but less eye-catching details get systematically underweighted.

The full write-up: study, numbers, and caveats
The psychology

Attention is a scarce resource, and choice requires comparing multiple attributes across multiple options at once, a task people can't do exhaustively. Salience theory formalises what fills that gap: attributes and outcomes that stand out (because they're extreme, unusual, or visually prominent relative to the immediate comparison set) get disproportionate decision weight, while equally or more important but less eye-catching details get systematically underweighted.

Where it causes errors

A single dramatic, low-probability outcome, a jackpot, a worst-case warning, a shockingly discounted price, can dominate a decision even when it's the least likely or least relevant number in the set, purely because it's the one detail that visually or emotionally jumps out from the rest.

Where it can help

Deliberately making the genuinely important number salient, the true total cost, the actual risk that matters most, rather than the most visually dramatic one, uses the same mechanism to help people weight decisions by what's actually important rather than by what simply catches the eye first.

Bordalo, P., Gennaioli, N., & Shleifer, A. (2012). “Salience Theory of Choice Under Risk.” Quarterly Journal of Economics, 127(3), 1243–1285
The formal, unifying theory of salience in choice.
StrengthBuilds a single formal model that generates quantitative, testable predictions for several previously separate, well-known anomalies in choice under risk (like the Allais paradox and specific patterns of risk-seeking behaviour), a genuine unifying contribution rather than another one-off demonstration.
WeaknessThe paper is primarily theoretical, built and tested against existing experimental data on choice under risk rather than a single original field experiment of its own, its predictive power outside gambling-style risk choices, in more everyday consumer decisions, rests on later applied work rather than this paper directly.
Key findings
Verbatim (abstract)“This leads the decision maker to a context-dependent representation of lotteries in which true probabilities are replaced by decision weights distorted in favor of salient payoffs. By endogenizing decision weights as a function of payoffs, our model provides a novel and unified account of many empirical phenomena, including frequent risk-seeking behavior, invariance failures such as the Allais paradox, and preference reversals.”
Verbatim“a lottery's salient payoffs are those which differ most strongly from the payoffs of alternative lotteries, and the decision maker's mind focuses on salient payoffs when making a choice.”
See also: Compromise Effect and Similarity Effect, all three are context-dependent choice theories; salience is the broadest, general-attention mechanism, while the other two describe specific, well-documented patterns salience-style theories are sometimes used to help explain.
Tested as an experiment: Does naming a savings pocket keep the money there?, a picturable label winning by contrast against the generic ones around it, tested as a live product change.
A specific case worth its own entry: Law of Round Numbers, a round number is a specific, reliable way to make one point on a scale more salient than its neighbours, strong enough to change behaviour on its own.
22

Left Digit Bias

Why does $2.99 feel so much cheaper than $3.00?

$2.99 gets processed as “two dollars something,” not $2.99, people anchor on the leftmost digit of a number and round off everything after it, even when the true difference to the next whole number is a single cent.

BUYERS PAY MORE FOR A CAR AT 39,999 MILES THAN 40,000 39,999 mi ODOMETER JUST UNDER 40,000 HIGHER SALE PRICE 40,000 mi ONE MILE MORE, DIGIT ROLLS OVER PRICE DROPS SHARPLY Just one mile changed on the odometer. The leftmost digit rolling over did the rest.
Likely mechanismThe first digit of a number does most of the work in how big it feels

The psychology. The brain encodes multi-digit numbers by their leftmost, most significant digit first, and this initial encoding disproportionately shapes the resulting magnitude judgment. The remaining digits get processed, but with far less weight. $2.99 and $3.00 differ by one cent in reality, but by a whole leading digit in how the price is first encoded, which is why the perceived gap feels much larger than one cent.

The full write-up: study, numbers, and caveats
The psychology

The brain encodes multi-digit numbers by their leftmost, most significant digit first, and this initial encoding disproportionately shapes the resulting magnitude judgment. The remaining digits get processed, but with far less weight. $2.99 and $3.00 differ by one cent in reality, but by a whole leading digit in how the price is first encoded, which is why the perceived gap feels much larger than one cent.

Where it causes errors

Millions of real transactions show the same pattern: a used car's price drops by a disproportionate amount the moment its odometer crosses a round 10,000-mile threshold, even though one extra mile changed nothing about the car, the leftmost digit rolling over changed how buyers processed the number.

Where it can help

Recognising the mechanism defuses it: deliberately reading a price's actual value, not just its leading digit, before comparing options, and for anyone pricing honestly, choosing round, un-gamed numbers rather than exploiting the .99 effect is the straightforward ethical alternative.

Spotted in the wild

McDonald's own app prices a real current offer one cent under a full dollar, the exact pricing shape the cited study measures on car odometers.

The My McDonald's app showing a pickup-only offer: Buy a Large Sundae, get a Hashbrown for $0.99, with the $0.99 highlighted

McDonald's, My McDonald's app. A “Buy a Large Sundae, get a Hashbrown for $0.99” offer, one cent under a full dollar. Captured 2026-08-28.

This is McDonald's own current pricing, not evidence that the missing cent changed what anyone actually bought. The evidence for that is the cited study below.

Lacetera, N., Pope, D. G., & Sydnor, J. R. (2012). “Heuristic Thinking and Limited Attention in the Car Market.” American Economic Review, 102(5), 2206–2236
22 million real used-car auction transactions.
StrengthAnalysed over 22 million real wholesale used-car auction transactions, genuine market data with real dollars, not a lab simulation, letting the researchers detect a discontinuous price drop precisely at each 10,000-mile odometer threshold.
WeaknessThe setting is wholesale used-car auctions specifically, populated partly by professional dealers alongside end consumers; the paper notes the effect looks driven by final customers rather than professional agents, but fully disentangling the two groups from aggregate auction data has real limits.
Key findings
Verbatim (abstract)“Analyzing over 22 million wholesale used-car transactions, we find substantial evidence of this left-digit bias; there are large and discontinuous drops in sale prices at 10,000-mile thresholds in odometer mileage, along with smaller drops at 1,000-mile thresholds.”
Verbatim (abstract)“We also investigate whether this heuristic behavior is primarily attributable to the final used-car customers or the used-car salesmen who buy cars in the wholesale market. The evidence is most consistent with partial inattention by final customers.”
Also worth citing: Thomas, M., & Morwitz, V. (2005), “Penny Wise and Pound Foolish: The Left-Digit Effect in Price Cognition,” Journal of Consumer Research, 32(1), 54–64, the classic lab demonstration behind familiar $X.99 pricing, showing the effect is strongest when the compared prices are otherwise close together, and that it isn't limited to prices alone.
See also: Anchoring, related but distinct: anchoring is about a whole number shaping later judgments; this is specifically about which digit within a number does the heaviest lifting.
Tested as an experiment: Does a home loan rate of 5.99% get remembered as a 5% loan, not almost 6%?, the same threshold effect, moved from a car's odometer onto a home loan's headline rate.
23

Ordering Effects

Why does the same list of options get chosen differently just by shuffling the order?

The same set of options, presented in a different order, gets chosen differently, items placed first or last draw disproportionate attention and votes, independent of their actual merit.

THE SAME CANDIDATE GAINS VOTES JUST BY BEING LISTED FIRST #4 LISTED LOWER DOWN THE BALLOT REFERENCE VOTE SHARE #1 SAME CANDIDATE, LISTED FIRST +2.5 PTS VOTE SHARE Same candidate, same platform, same voters. Only the position on the ballot moved.
Likely mechanismWhat comes first and last gets remembered; what's in the middle gets lost

The psychology. Two distinct memory effects compound here: primacy (early items get deeper, more deliberate processing before attention fades) and recency (the last item is still fresh in working memory at the moment of choosing). Whatever sits in the middle of a list gets neither advantage, evaluated more briefly and remembered less well, regardless of its actual quality.

The full write-up: study, numbers, and caveats
The psychology

Two distinct memory effects compound here: primacy (early items get deeper, more deliberate processing before attention fades) and recency (the last item is still fresh in working memory at the moment of choosing). Whatever sits in the middle of a list gets neither advantage, evaluated more briefly and remembered less well, regardless of its actual quality.

Where it causes errors

Ballot position alone has been shown to shift real vote share in real elections, the same candidate can gain or lose several percentage points purely from being listed first versus later, in races where voters have little other information to go on.

Where it can help

Randomising or rotating the order options are presented in, across ballots, survey questions, search results, or menu items, neutralises the position advantage and forces the comparison to run on the actual content, not where it happened to sit.

Miller, J. M., & Krosnick, J. A. (1998). “The Impact of Candidate Name Order on Election Outcomes.” Public Opinion Quarterly, 62(3), 291–330
Real 1992 Ohio ballot data across 118 races.
Experiment Teardown
118 down-ballot races
Only whether a candidate was listed first changed
Listed later
Same candidate, printed lower on the ballot
Listed first
Same candidate, printed top of the ballot
Result Later First Swing
Avg. vote share vs. reference 0 (ref) +2.5 pts +2.5

Position alone moved real votes. In the roughly half of races where the effect showed up, simply being listed first was worth an average +2.5 percentage points, same candidate, same platform, different print order.

Internal validityBallot position was assigned to precincts by a fixed rotation scheme, not true random assignment. The authors tested this directly and found no meaningful difference in turnout or demographics across the precinct groups, supporting the rotation as a functionally random natural experiment on real votes, not a survey.
External validityEffect concentrated in low-information down-ballot races; not expected, and not claimed, to hold at the same size in high-salience races with strong independent voter cues.
Key findings
Verbatim (abstract)“Reliable name-order effects appeared in 48 percent of 118 races, nearly always advantaging candidates listed first, by an average of 2.5 percent.”
Verbatim (abstract)“These effects were stronger in races when party affiliations were not listed, when races had been minimally publicized, and when no incumbent was involved.”
Also worth citing: Murdock, B. B. (1962), “The Serial Position Effect of Free Recall,” Journal of Verbal Learning and Verbal Behavior, 64, 482–488: the classic lab demonstration of the underlying memory mechanism: people best recall the first and last items in any list, whatever the content.
See also: Order Effect, related but different: this is about which item gets selected from a set by its position in it; that one is about how information order changes the judgement of a single target.
24

Chunking

Why is 5551234567 so much harder to remember than 555-123-4567?

Break a long string of information into small groups, a phone number, a password, a to-do list, and it becomes dramatically easier to hold in mind, without adding a single new fact.

GROUPING A LONG NUMBER HELPS PEOPLE REMEMBER IT 5551234567 ONE LONG STRING, TEN DIGITS HARD TO HOLD IN MIND 555-123-4567 SAME TEN DIGITS, THREE CHUNKS EASY TO HOLD IN MIND Same ten digits, no new information added. The mind just had three groups to hold, not ten.
Likely mechanismThe mind can only hold a handful of grouped pieces of information at once

The psychology. Working memory holds a genuinely small number of independent units at once, not a fixed number of individual digits or facts, but a fixed number of chunks, however much information is packed into each one. Grouping 9472163 into 947-216-3 doesn't reduce the amount of information, but it does reduce the number of separate units the mind has to juggle, from seven down to three.

The full write-up: study, numbers, and caveats
The psychology

Working memory holds a genuinely small number of independent units at once, not a fixed number of individual digits or facts, but a fixed number of chunks, however much information is packed into each one. Grouping 9472163 into 947-216-3 doesn't reduce the amount of information, but it does reduce the number of separate units the mind has to juggle, from seven down to three.

Where it causes errors

A long, unbroken block of instructions, a 16-digit account number with no groupings, or a form asking for a dozen independent inputs at once all overload working memory in a way that has nothing to do with how complex the underlying information actually is, people make more errors and abandon more often, purely because of formatting, not content.

Where it can help

Deliberately chunking information you're presenting, grouping digits, breaking a long form into short steps, structuring instructions as a short numbered list, respects the real limits of working memory and measurably reduces errors and drop-off, without hiding or removing a single piece of information.

Miller, G. A. (1956). “The Magical Number Seven, Plus or Minus Two: Some Limits on Our Capacity for Processing Information.” Psychological Review, 63(2), 81–97
The foundational paper on working-memory capacity and chunking.
StrengthTriangulated the same underlying capacity limit across three genuinely different task types, absolute judgment of single stimuli, memory span for sequences, and rapid counting, converging on a similar number of processable units from independent methods, rather than relying on just one paradigm.
WeaknessThe paper is a synthesis and reinterpretation of many earlier studies conducted by other researchers, using varied methods and materials, rather than a single new controlled experiment of Miller's own, its power comes from the pattern across old studies, which also means it inherits whatever limitations each of those original studies had.
Key findings
Verbatim“Everybody knows that there is a finite span of immediate memory and that for a lot of different kinds of test materials this span is about seven items in length. I have just shown you that there is a span of absolute judgment that can distinguish about seven categories and that there is a span of attention that will encompass about six objects at a glance.”
Verbatim“Since the memory span is a fixed number of chunks, we can increase the number of bits of information that it contains simply by building larger and larger chunks, each chunk containing more information than before.”
Decoded on The Science Behind: How does Up Bank's behavioural design pay off for both the bank and its customers?, where a long borrowing-power form and a six-figure deposit target are each split into several smaller, labelled steps.
Special report built on this: Breaking a Form Into Steps Changes How It Feels. It Doesn't Always Change How It Performs., where this exact mechanism is one of two explaining the Progress Flow form pattern, never tested on a form directly.
25

Translating Information

Why does “$2 a day” feel cheaper than the exact same $730 a year?

The same total cost feels completely different depending on the unit it's expressed in, a $600 yearly cost and “less than $2 a day” are the same number, restated.

CALLING $730 A YEAR "$2 A DAY" MAKES IT FEEL LIKE POCKET CHANGE $730/yr SHOWN AS ONE YEARLY TOTAL COMPARED TO A VACATION $2/day SAME COST, SHOWN PER DAY COMPARED TO A COFFEE Exactly the same $730, restated. Only what it got compared to changed.
Likely mechanismA cost gets compared to whatever similar expense comes to mind first

The psychology. When evaluating a cost, people don't independently calculate whether it's reasonable, they compare it to whatever similar expense comes to mind first. Reframing a cost in small, frequent units (a day, a coffee) prompts comparison against small, frequent expenses; the same cost reframed as one large aggregate figure prompts comparison against big, infrequent ones, and the exact same number can pass or fail that comparison depending purely on which frame was offered.

The full write-up: study, numbers, and caveats
The psychology

When evaluating a cost, people don't independently calculate whether it's reasonable, they compare it to whatever similar expense comes to mind first. Reframing a cost in small, frequent units (a day, a coffee) prompts comparison against small, frequent expenses; the same cost reframed as one large aggregate figure prompts comparison against big, infrequent ones, and the exact same number can pass or fail that comparison depending purely on which frame was offered.

Where it causes errors

A $600 annual subscription reframed as “less than $2 a day” can feel trivially affordable, while the identical cost stated as one $600 charge can feel significant enough to decline. The underlying financial reality hasn't moved at all, only which comparison it got measured against.

Where it can help

The same reframing works honestly in reverse: translating a small recurring cost into its real annual total (“that's $600 a year”) helps people catch subscriptions and fees that feel individually trivial but add up to something genuinely significant, useful specifically because it undoes the distortion rather than exploiting it.

Gourville, J. T. (1998). “Pennies-a-Day: The Effect of Temporal Reframing on Transaction Evaluation.” Journal of Consumer Research, 24(4), 395–408
The paper that names and formally tests the “pennies-a-day” effect.
StrengthTested the reframing effect across a genuinely wide range of real product and donation categories in multiple studies, and directly proposed and tested the underlying two-step mechanism (which comparison prices get retrieved, then how the transaction is evaluated against them) rather than only documenting that the effect exists.
WeaknessStudies are lab-based hypothetical evaluations of stated willingness to consider a purchase or donation, not tracked real purchase behaviour over time, whether the reframing changes actual long-run payment behaviour, not just an initial evaluation, needed further field confirmation.
Key findings
ParaphrasedFraming an identical cost as a small, recurring “pennies-a-day” amount rather than one large aggregate charge produced more favourable evaluations of the transaction at small daily amounts, but the effect reversed at larger daily amounts, a pattern found across a range of product and charitable-donation categories.
Verbatim (abstract)“The PAD framing of a target transaction is shown to systematically foster the retrieval and consideration of small ongoing expenses as the standard of comparison, whereas an aggregate framing of that same transaction is shown to foster the retrieval and consideration of large infrequent expenses. This difference in retrieval is shown to significantly influence subsequent transaction evaluation and compliance.”
See also: Price Transparency, related but different: that's about whether the full cost is shown at all; this is about which unit an already-visible cost gets expressed in.
26

Emergency Reserves (“Extra Lives” for Goals)

Why does giving yourself permission to fail once make you less likely to quit?

Build in a small, pre-approved allowance for slipping up, a “skip day,” a spare life, and people persist at a goal for longer than if the same goal demanded perfection.

A PRE-APPROVED SKIP DAY KEEPS PEOPLE ON A STREAK LONGER ONE SLIP BREAKS THE STREAK GIVES UP SOONER OK, COVERED ONE SLIP IS ALREADY COVERED PERSISTS LONGER Objectively equivalent goals, one slip either way. Only the framing of that slip changed.
Likely mechanismOne missed day can feel like the whole goal is already broken (the “what-the-hell effect”)

The psychology. A rigid, all-or-nothing goal has no room for a single slip: one missed day, one broken streak, and the goal itself can feel already failed, triggering the “what-the-hell effect”, a wave of guilt and rule-abandonment that turns one small lapse into giving up on the whole goal. A pre-approved reserve changes what a slip means: it's no longer a broken rule, just an anticipated and already-covered event, which keeps the goal, and the sense of still being on track, intact.

The full write-up: study, numbers, and caveats
The psychology

A rigid, all-or-nothing goal has no room for a single slip: one missed day, one broken streak, and the goal itself can feel already failed, triggering the “what-the-hell effect”, a wave of guilt and rule-abandonment that turns one small lapse into giving up on the whole goal. A pre-approved reserve changes what a slip means: it's no longer a broken rule, just an anticipated and already-covered event, which keeps the goal, and the sense of still being on track, intact.

Where it causes errors

A savings plan, diet, or habit tracker with zero tolerance for a missed day routinely produces a specific failure pattern: one bad day triggers a much larger collapse than the missed day itself justified, because the person now perceives the entire goal as already broken, not just delayed by one day.

Where it can help

This is close to a pure positive-use principle: explicitly building a small number of pre-approved “passes” into a goal, a skip day, a spare life, a grace period, measurably increases how long people persist at that goal compared to an objectively identical goal with no slack, without requiring the goal itself to become any easier.

Sharif, M. A., & Shu, S. B. (2017). “The Benefits of Emergency Reserves: Greater Preference and Persistence for Goals that Have Slack with a Cost.” Journal of Marketing Research, 54(3), 495–509
Six studies on goal slack, one tracking real day-by-day persistence over a week.
StrengthCombined six separate studies (mostly online scenario studies, plus one incentive-compatible task tracking real behaviour over seven days), directly comparing goals with a built-in emergency reserve against objectively identical goals without one, isolating the reserve itself, rather than any difference in the underlying goal, as the cause of the difference in persistence.
WeaknessThe reserves tested were deliberately given a real cost to use, not a free pass, as part of the design, meant to preserve commitment, a genuinely free, costless skip might behave differently (and potentially undermine the goal instead of protecting it), a boundary condition the paper itself flags rather than treating reserves as unconditionally beneficial.
Key findings
ParaphrasedPeople persisted longer at goals with a small emergency reserve built in than at objectively equivalent goals with no reserve, and also reported preferring goals structured with a reserve over otherwise identical goals without one.
ParaphrasedThe persistence effect ran through resistance to using the reserve itself: people with a reserve pushed harder specifically to avoid having to dip into it, even when using it carried no real downside, and that resistance, not a feeling of being back on track, is what the studies identified as driving the extra effort.
Also worth citing: This is the constructive flip side of the “what-the-hell effect” / abstinence violation effect: Marlatt, G. A., & Gordon, J. R. (1985), Relapse Prevention: Maintenance Strategies in the Addictive Behavior Change, and Cochran, W., & Tesser, A. (1996), in Striving and Feeling (Eds. Martin & Tesser), the well-documented failure mode where one lapse in an all-or-nothing goal triggers total abandonment rather than a simple, recoverable setback.
See also: Goal Gradient Effect, related goal-pursuit mechanism: that's about accelerating effort as a goal nears; this is about surviving the moments effort lapses instead.
27

Pain of Paying

Why does the same coffee taste better on a subscription than paid for by the cup?

The more directly and viscerally you feel the act of paying, the less you enjoy, and the less you consume, even when the price hasn't changed at all.

PAYING BEFORE YOU DRINK MAKES THE COST FEEL LIGHTER PAY FIRST, THEN DRINK FREELY FEELS LIGHT, BARELY NOTICED DRINK FIRST, PAY EACH TIME FELT WITH EVERY SIP Same coffee, same total cost, same timing of use. Only when the paying itself was felt changed.
Likely mechanismPaying hurts in proportion to how visible the act of paying is, not the amount

The psychology. Paying itself carries a small negative jolt, separate from the price paid, and that jolt scales with how salient, immediate, and attention-grabbing the payment method is. Cash, counted out note by note, hurts more than a card tap, which hurts more than a pre-authorised subscription charge you never see. The pain isn't proportional to the amount; it's proportional to how vividly the act of paying registers.

The full write-up: study, numbers, and caveats
The psychology

Paying itself carries a small negative jolt, separate from the price paid, and that jolt scales with how salient, immediate, and attention-grabbing the payment method is. Cash, counted out note by note, hurts more than a card tap, which hurts more than a pre-authorised subscription charge you never see. The pain isn't proportional to the amount; it's proportional to how vividly the act of paying registers.

Where it causes errors

Metered, pay-per-use pricing (paying by the bite, the minute, the click) can suppress consumption of something you'd actually enjoy more of, purely because each unit reminds you you're paying, a flat, prepaid, or bundled price removes that friction and can measurably increase how much of the same thing you use.

Where it can help

Businesses can use this honestly by removing needless payment friction from things that are genuinely good for the customer, a wellness app charging once a year instead of nagging weekly, rather than only using it to numb people into overspending on subscriptions they'd cancel if the pain of paying were more visible.

Prelec, D., & Loewenstein, G. (1998). “The Red and the Black: Mental Accounting of Savings and Debt.” Marketing Science, 17(1), 4–28
The paper that formalised and named the pain of paying.
StrengthBuilt a full “double-entry” mental accounting model rather than a single demonstration, explaining a wide range of previously separate findings (prepayment preference, flat-rate bias, credit-card overspending) as consequences of one mechanism.
WeaknessAs a theoretical/synthesis paper, its evidence is drawn from existing consumer behaviour patterns and earlier experiments rather than one new controlled study of its own. The theory's individual predictions have needed separate empirical testing since.
Key findings
Verbatim (abstract)“Such pay-before sequences confer hedonic benefits because consumption can be enjoyed without thinking about the need to pay for it in the future.”
Verbatim (abstract)“our model predicts that consumers will find it less painful to pay for, and hence will prefer, flat-rate pricing schemes such as unlimited Internet access at a fixed monthly price, even if it involves paying more for the same usage.”
See also: Mental Accounting, the broader model this principle is one application of: how payment method itself changes which mental account an expense gets charged to. Also see Payment Transparency, the more specific mechanism (rehearsal and immediacy) behind why some payment methods hurt more than others.
28

Peak-End Rule (& Duration Neglect)

Why do you remember a vacation by its worst afternoon, not its ten good days?

You remember an experience almost entirely by its worst (or best) moment and how it ended, not by how long it lasted or its average intensity.

ENDING A PROCEDURE GENTLY MAKES PATIENTS RECALL LESS PAIN PAIN: 2.5 ENDS ABRUPTLY AT PEAK PAIN RATED 4.9 OVERALL UNPLEASANTNESS PAIN: 1.7 ENDS GENTLY, PAIN TAPERS OFF RATED 4.4 OVERALL UNPLEASANTNESS The gently-ending patients felt pain for longer overall. A milder ending was still remembered as easier.
Likely mechanismMemory keeps the most intense moment and the ending, and drops almost everything else

The psychology. Memory doesn't store an experience as a continuous record; it compresses it into a small number of representative snapshots, and the most intense moment (the peak) plus the final moments (the end) dominate that compressed summary. Duration barely registers at all, an experience twice as long, with the same peak and end, is remembered almost identically.

The full write-up: study, numbers, and caveats
The psychology

Memory doesn't store an experience as a continuous record; it compresses it into a small number of representative snapshots, and the most intense moment (the peak) plus the final moments (the end) dominate that compressed summary. Duration barely registers at all, an experience twice as long, with the same peak and end, is remembered almost identically.

Where it causes errors

A genuinely good service can be remembered as worse than it was because of one badly-timed moment near the end, a rushed goodbye, a clumsy final step, and a long, mostly excellent experience gets no credit in memory for all the time it was going well, simply because duration isn't what memory tracks.

Where it can help

Deliberately engineering a strong, positive final moment, a genuine thank-you, a small unexpected touch right at the end, reliably improves how an entire experience is remembered, even without changing anything that happened earlier. One of the most replicated, low-cost levers in service design.

Redelmeier, D. A., Katz, J., & Kahneman, D. (2003). “Memories of Colonoscopy: A Randomized Trial.” Pain, 104(1–2), 187–194
The follow-up randomised trial: deliberately extends the procedure to test whether a milder ending changes memory.
Experiment Teardown
682 colonoscopy patients
Only how the procedure ended differed
Standard procedure
Ends as soon as the scope is withdrawn
Extended procedure
Scope held still an extra interval: mild, not zero, discomfort
Result Standard Extended Swing
Final pain 2.5 1.7 −0.8
Overall rating 4.9 4.4 −0.5

A longer, more painful procedure, remembered as better. Adding a milder tail end lowered both the final-moment pain and the overall retrospective score, even though total pain and duration both increased.

Internal validityA randomised trial (n=682, split standard vs. extended) with real-time pain recorded during the procedure, not only recalled afterwards.
External validityTested on colonoscopy specifically; the original 1996 study found the same peak/end pattern in a second, different procedure (lithotripsy), suggesting the mechanism isn't specific to one procedure.
Key findings
ParaphrasedIn this 2003 trial itself, patients' retrospective ratings of how bad the overall procedure was correlated with peak pain and final-minute pain, but the duration of the procedure barely predicted the rating at all (r = 0.10 and 0.09 for the two retrospective measures), the same pattern the original 1996 study first identified.
ParaphrasedPatients in the extended-procedure group also returned for a repeat colonoscopy at a higher rate over a median 5.3-year follow-up (odds ratio 1.41) than patients in the standard-procedure group, the better memory measurably changed later behaviour, not just the retrospective rating.
Also worth citing: Redelmeier, D. A., & Kahneman, D. (1996). “Patients’ Memories of Painful Medical Treatments: Real-Time and Retrospective Evaluations of Two Minimally Invasive Procedures.” Pain, 66(1), 3–8, the original paper (n=154 colonoscopy, n=133 lithotripsy) that first established the peak/end correlation and duration neglect, setting up the randomised follow-up trial above.
29

Illusory Truth Effect

Why does hearing a lie twice make it start to feel true?

Hearing a claim repeated makes it feel truer, even if you don't consciously remember the repetition, and even if you correctly identified it as false the first time.

HEARING A CLAIM THREE TIMES MAKES PEOPLE RATE IT TRUER SESSION 1, FIRST TIME HEARD RATED ONLY A LITTLE TRUE SESSION 3, THIRD TIME HEARD RATED THE TRUEST Same statement, no new evidence, across all three sessions. Only repeated exposure changed how true it felt.
Likely mechanismRepetition alone makes a claim feel more true, even without any new evidence

The psychology. Familiarity is used as a proxy for accuracy: statements that feel easier to process, because they've been encountered before, get misread as more likely to be true, even when the actual source of that ease was mere repetition, not evidence. This runs largely below conscious awareness; knowing the mechanism doesn't fully protect you from it.

The full write-up: study, numbers, and caveats
The psychology

Familiarity is used as a proxy for accuracy: statements that feel easier to process, because they've been encountered before, get misread as more likely to be true, even when the actual source of that ease was mere repetition, not evidence. This runs largely below conscious awareness; knowing the mechanism doesn't fully protect you from it.

Where it causes errors

Misinformation that gets repeated, even repeated specifically to debunk it, can become more believable purely through repeated exposure, independent of whether anyone consciously accepts it as true. Simple repeated claims in advertising or political messaging exploit the same mechanism.

Where it can help

The same mechanism works for the true and useful things worth someone remembering, a genuine safety instruction, a real key fact. Repeating an accurate message across multiple honest touchpoints increases how confidently and correctly people recall it.

Hasher, L., Goldstein, D., & Toppino, T. (1977). “Frequency and the Conference of Referential Validity.” Journal of Verbal Learning and Verbal Behavior, 16(1), 107–112
The founding demonstration of the effect.
StrengthA genuinely longitudinal design, the same participants rated the same statements on three separate occasions two weeks apart, directly showing belief in a statement's truth rose specifically with repeated exposure to that statement over real time, not just within one sitting.
WeaknessStatements used were plausible-sounding trivia rather than claims with real personal or political stakes, whether the effect is weaker or stronger for claims people are already motivated to believe or disbelieve needed later research to establish.
Key findings
Verbatim“As can be seen from Table 2, the average rating assigned to repeated statements increased across successive sessions, while the rating assigned to nonrepeated statements diminished slightly.”
Verbatim“The increase in validity ratings with repetition was equivalent for true and for false statements, despite the fact that subjects succeeded in discriminating between them.”
Also worth citing: Fazio, L. K., Brashier, N. M., Payne, B. K., & Marsh, E. J. (2015), “Knowledge Does Not Protect Against Illusory Truth,” Journal of Experimental Psychology: General, found the effect persists even for statements people could correctly identify as false based on their own prior knowledge, showing repetition can override what someone already knows.
30

Illusion of Control

Why do people blow on dice or pick their own lottery numbers?

People act as though they can influence outcomes that are actually random or already fixed, and that belief can persist indefinitely if nobody ever corrects it.

PEOPLE ASK MORE TO SELL A LOTTERY TICKET THEY CHOSE 07 14 22 THEY CHOSE THESE NUMBERS WON’T SELL IT FOR LESS 03 19 41 HANDED THIS TICKET INSTEAD SELLS IT EASILY Both tickets carry exactly the same odds. Choosing the numbers just felt like skill.
Likely mechanismSkill cues misapplied to chance outcomes

The psychology. Ordinary skill situations reliably increase real control, so the mind uses skill-like cues, choice, competition, involvement, familiarity, as a shortcut for inferring control even in situations that are actually pure chance. Add any of those cues to a random process and confidence in influencing the outcome rises, with no real change in the odds.

The full write-up: study, numbers, and caveats
The psychology

Ordinary skill situations reliably increase real control, so the mind uses skill-like cues, choice, competition, involvement, familiarity, as a shortcut for inferring control even in situations that are actually pure chance. Add any of those cues to a random process and confidence in influencing the outcome rises, with no real change in the odds.

Where it causes errors

People persist in believing an action they take has an effect, a button, a ritual, a specific strategy, long after it stopped doing anything, or never did, simply because nothing ever forces them to test the belief directly.

Where it can help

There's little ethical upside to deliberately manufacturing this illusion in others. The constructive use runs the other way: noticing where you might be pressing a metaphorical dead crosswalk button, checking a phone repeatedly, refreshing a page, a superstitious pre-decision ritual, and testing whether the action actually changes anything.

Langer, E. J. (1975). “The Illusion of Control.” Journal of Personality and Social Psychology, 32(2), 311–328
The foundational series of studies.
StrengthA series of controlled experiments systematically varied specific skill-like cues (choice of ticket number, competition against a weak vs. strong opponent, familiarity with the materials) one at a time within otherwise pure-chance games, isolating exactly which cues drive the inflated confidence.
WeaknessFive of the six studies ran in real workplaces or a public racetrack rather than a lab, and stakes stayed fairly low (a few dollars per ticket, with one exception involving a genuine $2,000 scholarship); whether the same magnitude of illusion holds for larger, life-changing financial or health decisions still needed separate confirmation.
Key findings
Verbatim“The mean amount of money required for the subject to sell his ticket was $8.67 in the choice condition and only $1.96 in the no-choice condition.”
Verbatim“This illusion may be induced by introducing competition, choice, stimulus or response familiarity, or passive or active involvement into a chance situation. When these factors are present, people are more confident and are more likely to take risks.”
Also worth citing: New York City deactivated roughly three-quarters of its pedestrian crosswalk buttons starting in the 2000s without telling the public, many kept being pressed for years, an illusion of control persisting purely because it was never corrected (reported by the New York Times, 2004; discussed in Stuart Vyse's The Uses of Delusion, 2022).
31

Hindsight Bias

Why does every crisis look obvious in hindsight, right after it happens?

Once you know how something turned out, it feels like you basically knew it all along, even on things that were genuinely uncertain beforehand.

LEARNING THE OUTCOME MAKES PEOPLE FEEL THEY "KNEW IT ALL ALONG" BEFOREHAND, OUTCOME UNKNOWN SAID "COULD GO EITHER WAY" AFTER BEING TOLD THE OUTCOME SAID "I KNEW IT ALL ALONG" Same event, same evidence, both times. Only knowing the outcome was new.
Likely mechanismMemory of prior uncertainty reconstructed after the fact

The psychology. New outcome information doesn't just get added to memory alongside what you believed before, it actively reshapes the memory of what you believed before, making past uncertainty feel smaller in hindsight than it actually was. You're not lying about having predicted it; your memory of your own prior uncertainty has been genuinely, invisibly edited.

The full write-up: study, numbers, and caveats
The psychology

New outcome information doesn't just get added to memory alongside what you believed before, it actively reshapes the memory of what you believed before, making past uncertainty feel smaller in hindsight than it actually was. You're not lying about having predicted it; your memory of your own prior uncertainty has been genuinely, invisibly edited.

Where it causes errors

A decision that was genuinely reasonable given what was known at the time gets unfairly judged as careless or obvious once the outcome is known, “how did nobody see this coming”, which discourages honest post-mortems and can punish good decision-making that simply had a bad outcome.

Where it can help

Writing down predictions, confidence levels, and reasoning before an outcome is known, and reviewing that written record afterwards, rather than relying on memory, is the direct, low-cost countermeasure, and is standard practice in well-run forecasting and decision-review processes.

Fischhoff, B. (1975). “Hindsight ≠ Foresight: The Effect of Outcome Knowledge on Judgment Under Uncertainty.” Journal of Experimental Psychology: Human Perception and Performance, 1(3), 288–299
The founding paper, still the standard citation fifty years on.
StrengthDirectly compared what participants said they would have predicted beforehand against what they actually predicted, using real historical events described in advance, a genuine before/after design rather than only measuring after-the-fact confidence.
WeaknessRelied on participants' self-reports of their own prior beliefs, recreated under experimental instruction, rather than a true prospective record of what each person actually believed before the study began, a limitation later, more rigorous prospective designs were built to address.
Key findings
Verbatim“In each of the 24 cases, reporting an outcome increased its perceived likelihood of occurrence.”
Verbatim“Experiment 2 has shown that subjects are either unaware of outcome knowledge having an effect on their perceptions or, if aware, they are unable to ignore or rescind that effect.”
32

Survivorship Bias

Why did WWII engineers almost armour the wrong part of the returning bombers?

Studying only the survivors of a selection process, the funds that didn't collapse, the planes that made it home, the founders who succeeded, gives a systematically distorted picture, because the failures that would correct it are invisible.

ENGINEERS ALMOST ARMOURED THE HOLES, NOT THE ENGINE RETURNING PLANES: HOLES IN WINGS SEEMS LIKE: ARMOUR THE WINGS ENGINE-HIT PLANES NEVER RETURNED REALLY: ARMOUR THE ENGINE The engine showed no holes in the data, because those hits never made it home.
Likely mechanismNon-random sample created by invisible failures

The psychology. There's no active illusion here so much as a data-collection blind spot: when the cases that fail also disappear from view, whatever you can still observe is, by definition, a non-random and misleading sample. The mind treats a vivid, available set of examples as representative, without checking whether an entire category of counter-examples was ever able to be seen at all.

The full write-up: study, numbers, and caveats
The psychology

There's no active illusion here so much as a data-collection blind spot: when the cases that fail also disappear from view, whatever you can still observe is, by definition, a non-random and misleading sample. The mind treats a vivid, available set of examples as representative, without checking whether an entire category of counter-examples was ever able to be seen at all.

Where it causes errors

Studying only successful companies, funds, or people to extract “lessons for success” routinely produces confident-sounding advice that's actually just a description of what survivors happened to have in common, the same traits are often just as common among the failures nobody studied, because they didn't survive to be studied.

Where it can help

Deliberately seeking out the missing failure data, the funds that closed, the products that were discontinued, the applicants who were rejected, before drawing a conclusion is the direct countermeasure, and is exactly the move that made Wald's WWII analysis correct where the intuitive version was backwards.

Mangel, M., & Samaniego, F. J. (1984). “Abraham Wald’s Work on Aircraft Survivability.” Journal of the American Statistical Association, 79(386), 259–267
The paper that reconstructs and formalises Wald's original wartime reasoning.
StrengthReconstructs Wald's actual statistical method from his original wartime reports, rather than relying on the popularised, simplified retelling, showing the real mathematical reasoning behind the now-famous “armour the parts with no bullet holes” conclusion.
WeaknessThis is a historical and methodological reconstruction of a real WWII analysis, not a new controlled study. Its evidentiary weight rests on Wald's original military data and reasoning, which weren't designed as a public, replicable dataset.
Key findings
ParaphrasedWald recognised that damage data collected only from planes that returned from combat systematically excluded the planes that didn't return, and that the correct conclusion was to reinforce the areas showing the least damage on survivors, since planes hit in the heavily-damaged areas were the ones not coming back.
Verbatim“Wald's methods were used in World War II and by the Navy and Air Force during the wars in Korea and Vietnam.”
33

Take-the-Best Heuristic (Less-Is-More)

Why can one good clue beat a spreadsheet full of data?

A simple rule that uses just one good piece of information can out-predict a complicated model trying to weigh everything. More inputs aren't always more accurate.

A SIMPLE ONE-CUE RULE PREDICTS NEW CASES BETTER COMPLEX MODEL, WEIGHS EVERY INPUT STUMBLES ON NEW DATA ONE BEST CUE, IGNORES THE REST GENERALIZES BETTER Less overfitting, not less accuracy. A simple rule is less fooled by noise.
Likely mechanismA simple rule can predict better than a complex one, because it isn't fooled by noise in the data

The psychology. Complex models fit their training data well but can overfit, picking up noise alongside the real signal, while a simple “take the best single cue and stop” rule, by using less information, is less exposed to that noise. Under genuine uncertainty, the simpler rule can generalise better precisely because it asks less of the data.

The full write-up: study, numbers, and caveats
The psychology

Complex models fit their training data well but can overfit, picking up noise alongside the real signal, while a simple “take the best single cue and stop” rule, by using less information, is less exposed to that noise. Under genuine uncertainty, the simpler rule can generalise better precisely because it asks less of the data.

Where it causes errors

Assuming a more complex decision process will always beat a simple one leads people to distrust genuinely good simple heuristics, and to build overengineered decision processes that perform worse, not better, than a well-chosen simple rule, especially in noisy, real-world conditions.

Where it can help

In domains where information is genuinely limited or noisy, deliberately identifying the single most predictive cue and building a decision rule around it, rather than trying to weigh a dozen partially-reliable inputs, can be both simpler to use and more accurate, not a compromise between the two.

Gigerenzer, G., & Goldstein, D. G. (1996). “Reasoning the Fast and Frugal Way: Models of Bounded Rationality.” Psychological Review, 103(4), 650–669
The foundational paper on fast-and-frugal heuristics.
StrengthDirectly pitted the simple Take-the-Best algorithm against five other statistical models, including full multiple regression, across roughly 858 million simulated head-to-head comparisons drawn from one real-world dataset, the actual populations of the 83 largest cities in Germany, not only arguing for simplicity in the abstract.
WeaknessThe test used a single, unusually well-behaved domain (city populations, which the nine cues predict very well) and simulated rather than real decision-makers. Multiple regression was actually given every advantage, the true population figures and individually optimised weights for each simulated case, which makes its tie with Take-the-Best more notable but also means this is one demonstration rather than a test across many kinds of real-world environments.
Key findings
Verbatim (abstract)“The Take The Best algorithm matched or outperformed all competitors in inferential speed and accuracy.”
ParaphrasedThe advantage traced to the recognition principle. Models that folded every cue, including recognition, into one weighted sum sometimes ranked an unfamiliar city above a familiar one, simply because the familiar city carried more negative cue values on record. Take-the-Best and simple tallying let recognition act as its own single cue and avoided that mistake. That is also why every algorithm's accuracy peaked partway through learning, when a person recognized some cities but not all eighty-three, exactly the paper's less-is-more effect.
34

Choice Bracketing (Diversification Bias)

Why do you choose differently picking snacks for the whole week at once vs. one at a time?

Choosing once for a whole set of future decisions produces different choices than making each decision fresh, one at a time, same options, different outcome, depending purely on how the choices are grouped.

CHOOSING SNACKS ALL AT ONCE PRODUCES MORE VARIETY CHOSEN ONE AT A TIME SAME FAVOURITE, EVERY TIME CHOSEN ALL AT ONCE DELIBERATELY VARIED Same occasions, same appetite, same options. Only how bundled the choices felt changed.
Likely mechanismChoosing several things at once favours variety; choosing them one at a time repeats the favourite

The psychology. When several choices are considered together (bracketed broadly), the mind treats them as ingredients in a varied whole and actively seeks contrast between them, pushing towards more variety. When the same choices are made one at a time (bracketed narrowly), each decision gets evaluated on its own immediate merits, usually the same favourite, repeated.

The full write-up: study, numbers, and caveats
The psychology

When several choices are considered together (bracketed broadly), the mind treats them as ingredients in a varied whole and actively seeks contrast between them, pushing towards more variety. When the same choices are made one at a time (bracketed narrowly), each decision gets evaluated on its own immediate merits, usually the same favourite, repeated.

Where it causes errors

Choosing a week's worth of snacks or meals in advance reliably produces more variety than most people actually want in the moment they eat each one, the future self bracketed broadly wants variety in the abstract; the present self making one choice at a time usually just wants the thing it likes.

Where it can help

Deliberately choosing which bracket fits a decision, committing in advance to a single simple choice for a recurring decision that doesn't benefit from variety, versus leaving room for in-the-moment choice where variety and mood genuinely matter, routing around the mismatch instead of defaulting into it by accident.

Read, D., & Loewenstein, G. (1995). “Diversification Bias: Explaining the Discrepancy in Variety Seeking Between Combined and Separated Choices.” Journal of Experimental Psychology: Applied, 1(1), 34–49
The paper that names diversification bias and identifies choice bracketing as one of its causes.
StrengthSystematically tested and ruled out several competing explanations, including satiation and simple utility-maximising logic, before isolating choice bracketing and time-contraction as the actual mechanisms: a genuine process of elimination, not just a single demonstration.
WeaknessExperiments centre on low-stakes, easily-repeated consumption choices like snacks and soft drinks; whether the same bracketing effect operates identically for larger, less frequent decisions needed separate confirmation.
Key findings
Verbatim (abstract)“If people make combined choices of quantities of goods for future consumption, they choose more variety than if they make separate choices immediately preceding consumption.”
Verbatim“All 13 children in the combined choice condition chose two different candy bars, compared with only 48% (12 of 25) in the separate choice condition.”
35

The Precision Effect

Why does $21,947 sound more like a real number than $22,000?

A precise number ($21,947) and a round one ($22,000) trigger different reactions in a negotiation, and precision doesn't always help: it can make a novice seem more credible, but make an expert seem less competent.

A PRECISE OFFER IMPRESSES AMATEURS, NOT EXPERTS $21,947 OFFERED TO AN AMATEUR READS AS CREDIBLE $21,947 OFFERED TO AN EXPERT READS AS INEXPERIENCED The exact same number, offered either way. Only who read it changed what it signalled.
Likely mechanismPrecision read as a competence signal, until it isn't

The psychology. A precise number implies the person behind it did real, careful work to arrive at it, a useful signal when the audience assumes competence is genuine. But among experts who know the domain, an unusually precise number can instead read as artificial confidence, or the tell of someone padding a number to look rigorous rather than actually being so, flipping the same signal from a credibility boost into a credibility cost.

The full write-up: study, numbers, and caveats
The psychology

A precise number implies the person behind it did real, careful work to arrive at it, a useful signal when the audience assumes competence is genuine. But among experts who know the domain, an unusually precise number can instead read as artificial confidence, or the tell of someone padding a number to look rigorous rather than actually being so, flipping the same signal from a credibility boost into a credibility cost.

Where it causes errors

Assuming more precision always signals more credibility can badly misfire when negotiating with someone who actually knows the domain: an experienced negotiator on the other side of the table may read your carefully-precise opening figure as overcompensation, not rigor.

Where it can help

Matching the precision of a number to the actual sophistication of the audience, more precise for people who'll take it as a genuine effort signal, rounder (or precise but explicitly justified) for experts who might otherwise read excess precision as posturing, uses the same finding honestly rather than by accident.

Loschelder, D. D., Friese, M., Schaerer, M., & Galinsky, A. D. (2016). “The Too-Much-Precision Effect: When and Why Precise Anchors Backfire With Experts.” Psychological Science, 27(12), 1573–1587
Five experiments, 1,320 participants including real professional negotiators.
StrengthRan five separate experiments spanning multiple real professional domains (real estate, jewellery, car sales, human resources) with a combined 1,320 participants including genuine domain experts, not just students, giving the finding real breadth across contexts.
WeaknessExperiments 1 and 2 compared different professional groups (real-estate agents and jewellers against people with no negotiation experience) rather than randomly assigning expertise, so part of the original gap could in principle reflect other differences besides negotiation experience. The researchers anticipated this and ran a fourth experiment contrasting car salespeople against mechanics, who share the same product knowledge but differ specifically in negotiation experience, and found the same pattern, which narrows but doesn't fully close the concern for the real-estate and jewelry findings.
Key findings
Verbatim“For amateurs, WTP and counteroffers increased linearly with precision of the first offer. For experts, WTP and counteroffers increased with precision only up to a point, after which greater precision backfired.”
Verbatim (abstract)“Anchor precision backfired because experts saw too much precision as reflecting a lack of competence.”
See also: Left Digit Bias, related but different: that's about which digit in a number dominates how it's read; this is about what a number's overall precision signals about the person offering it.
36

Default Effect

Why does a country's organ donation rate jump from 4% to 99% just by changing one checkbox?

Whatever option is pre-selected for you gets chosen at a dramatically higher rate than the exact same option would if you had to actively pick it, even when picking would take seconds and nothing else stands in the way.

PEOPLE STAY ENROLLED IN A DEFAULT, EVEN A POOR FIT A SAME PLAN, MUST ACTIVELY PICK RARELY CHOSEN A SAME PLAN, SET AS DEFAULT STAYS ENROLLED Which plan was best never changed. Only whether it was pre-selected did.
Likely mechanismStaying with the default feels safer, easier, and like the recommended choice

The psychology. Staying with a default requires no decision at all, so it inherits several advantages at once: it reads as an implicit recommendation, it avoids the small effort of choosing, and it avoids the risk of regretting an active choice that turns out worse, all of which push towards passivity even when someone has every practical ability to switch.

The full write-up: study, numbers, and caveats
The psychology

Staying with a default requires no decision at all, so it inherits several advantages at once: it reads as an implicit recommendation, it avoids the small effort of choosing, and it avoids the risk of regretting an active choice that turns out worse, all of which push towards passivity even when someone has every practical ability to switch.

Where it causes errors

People stay enrolled in a randomly-assigned default option even when it's a genuinely poor fit for them and switching costs nothing but a small amount of effort and attention, the default's power over behaviour holds up even in situations deliberately designed to remove every other reasonable explanation for staying.

Where it can help

This is one of the clearest cases where the same mechanism has an unambiguous constructive use: setting the default to whatever option best serves most people (retirement-saving auto-enrolment, organ-donor registration) reliably improves outcomes at essentially zero cost to anyone's freedom to choose differently.

Brot-Goldberg, Z. C., Layton, T. J., Vabson, B., & Wang, A. Y. (2023). “The Behavioral Foundations of Default Effects: Theory and Evidence from Medicare Part D.” American Economic Review, 113(10), 2718–2758
A genuine natural experiment using randomly-assigned real defaults.
StrengthExploits a real natural experiment in which new low-income Medicare beneficiaries are randomly assigned a default prescription drug plan by the government: true random assignment of the default itself, in a real, high-stakes setting, not a lab simulation of one.
WeaknessThe population studied is specifically low-income Medicare beneficiaries in a subsidised programme, a group that may face particular attention or capacity constraints; how strongly the same passivity generalises to higher-resource populations facing other kinds of defaults needed separate confirmation.
Key findings
Verbatim (abstract)“We estimate that when a beneficiary’s default is exogenously changed from one year to the next, 96% of beneficiaries follow that default.”
Verbatim“These results provide strong evidence that consumers follow defaults in this market, no matter what that default may be, and largely reject the hypothesis that the high levels of choice persistence in this program are explained by consumers actively choosing to remain in their plans in order to avoid real switching costs.”
See also: Smart Defaults, the deliberate version of the same lever. This principle is about the power of whatever's pre-selected, passive by default; that one is about choosing when a default takes effect and what it's measured against, so it works with someone's biases instead of against them.
See also: Every Default Decides Who Pays for Doing Nothing, a full report placing this principle on a wider spectrum: forced choice, opt-in, opt-out, timed defaults, and personalised defaults, each with its own real study and its own honest catch.
37

Mental Accounting

Why does the same $20 feel different depending on whether you found it or earned it?

Money gets mentally filed into separate accounts by source or purpose (a windfall, a grocery budget, a ticket already paid for) and a dollar in one account gets treated as worth less than the identical dollar in another.

LOSING A TICKET STOPS A REBUY MORE THAN LOSING CASH LOST: TICKET LOST THE $10 TICKET 46% WOULD STILL BUY ANOTHER LOST: $10 CASH LOST $10 CASH INSTEAD 88% WOULD STILL BUY A TICKET Same $10 loss, either way. Only which mental account it came from differed.
Likely mechanismMoney gets mentally sorted into separate buckets, even though it all spends the same

The psychology. Standard economic theory treats money as fungible: a dollar is a dollar, regardless of where it came from or what it's earmarked for. People don't actually budget that way: money gets sorted into mental categories by source (a bonus vs. a paycheck) or intended use (rent vs. entertainment), and those category boundaries change how willingly the money gets spent, even when moving it between categories would leave someone objectively better off.

The full write-up: study, numbers, and caveats
The psychology

Standard economic theory treats money as fungible: a dollar is a dollar, regardless of where it came from or what it's earmarked for. People don't actually budget that way: money gets sorted into mental categories by source (a bonus vs. a paycheck) or intended use (rent vs. entertainment), and those category boundaries change how willingly the money gets spent, even when moving it between categories would leave someone objectively better off.

Where it causes errors

The clearest everyday version: carrying high-interest credit card debt while keeping money sitting in a low-interest savings account, because “savings” and “debt repayment” live in separate mental accounts that don't get netted against each other, even though doing the maths would make the trade obviously worth it.

Where it can help

The same partitioning, used deliberately, is a legitimate savings tool: earmarking money for a real goal (a house deposit, an emergency fund) in a separate, named account makes it psychologically harder to raid for everyday spending. That's a self-imposed commitment device, not something done to someone else, and a specific version of the wider lever covered in Precommitment Devices.

Kahneman, D., & Tversky, A. (1984). “Choices, Values, and Frames.” American Psychologist, 39(4), 341–350
Two independent samples, surveyed with one-sentence changes to which mental account a $10 loss was charged to.
Experiment Teardown
383 respondents, two independent samples
Only which mental account the $10 loss was charged to differed
Lost the ticket
A $10 ticket, already bought, is lost before entering
Lost the cash
$10 in cash is lost before any ticket is bought
Result Ticket Cash Swing
Would still buy 46% 88% +42

The identical $10 loss, filed under a different mental account, nearly doubled people's willingness to spend. Losing the ticket felt like overspending on “the theatre”; losing the equivalent cash didn't touch that account at all.

Internal validityThe headline 46%/88% comparison came from two large independent samples (n=200 and n=183) with random, between-subjects assignment to one framing or the other. A separate check, presenting both versions to the same respondents, found the ticket answer shifted after seeing the cash version first, but not the reverse: evidence the two scenarios were being read as connected once people could compare them.
External validityA hypothetical vignette, not real money on the line: a genuine limitation of this specific study, though the same asymmetry has since been observed in studies using real, not imagined, money.
Key findings
ParaphrasedRespondents told they'd lost the $10 theatre ticket itself were much less willing to pay $10 for a replacement than respondents told they'd lost an equivalent $10 bill on the way to the theatre, even though both groups faced an identical net cost.
Verbatim“Buying a second ticket increases the cost of seeing the play to a level that many respondents apparently find unacceptable. In contrast, the loss of the cash is not posted to the account of the play, and it affects the purchase of a ticket only by making the individual feel slightly less affluent.”
Also worth citing: Thaler, R. (1985). “Mental Accounting and Consumer Choice.” Marketing Science, 4(3), 199–214, the paper that formalised mental accounting into a full theoretical model, already cited on this site as the foundation for Pain of Paying.
See also: Pain of Paying, a direct application of the same mental-accounting model to how payment method itself changes spending, and Payment Transparency, the separate mechanism (rehearsal and immediacy) that governs whether a payment gets tracked into any mental account at all.
Tested as an experiment: Does naming a savings pocket keep the money there?, a live-product test of this mechanism, built on the high school field session.
Decoded on The Science Behind: Why does GoalSaver's 4.75% bonus feel like money you'd be losing, not money you haven't earned yet?, where a named, goal-labelled GoalSaver balance is the real, everyday version of a separate mental account.
Also decoded on The Science Behind: How does Up Bank's behavioural design pay off for both the bank and its customers?, where Up's Maybuys saves towards a photo of the actual product a customer wants, not just a labelled balance.
38

Endowment Effect

Why do you suddenly want more money for your mug the moment it's yours?

Merely owning something, even briefly, makes you value it more than you would if you didn't own it, so people demand more to give it up than they'd ever pay to acquire the identical thing.

OWNING A MUG MAKES SELLERS ASK TWICE WHAT BUYERS OFFER $5.25 OWNS THE MUG (SELLER) $5.25 WON’T SELL FOR LESS $2.25 DOESN’T OWN IT (BUYER) $2.25 WON’T PAY MORE The identical mug, no other difference. Owning it roughly doubled its price.
Likely mechanismOwning something makes giving it up feel like a loss, which looms larger than an equivalent gain

The psychology. Loss aversion applies to ownership itself: giving up something you already possess registers as a loss, and losses are felt more sharply than equivalent gains. That pulls the minimum price someone would accept to sell an item upward, above what a neutral, non-owning buyer would pay for the identical thing, a gap created entirely by which side of the transaction someone happens to be standing on.

The full write-up: study, numbers, and caveats
The psychology

Loss aversion applies to ownership itself: giving up something you already possess registers as a loss, and losses are felt more sharply than equivalent gains. That pulls the minimum price someone would accept to sell an item upward, above what a neutral, non-owning buyer would pay for the identical thing, a gap created entirely by which side of the transaction someone happens to be standing on.

Where it causes errors

Free trials and generous return windows work partly because of this. Once a product has sat in your home for two weeks, sending it back starts to feel like a loss. That inflates its perceived value well past what you'd have paid for it cold, and can quietly push people into keeping, and paying for, things they wouldn't have chosen from scratch.

Where it can help

A genuinely no-obligation trial period lets people discover real fit before committing, which is informationally useful rather than manipulative on its own. And naming the bias personally is a clean debiasing habit: asking “would I buy this today, at this price, if I didn't already own it?” cuts through the ownership pull when deciding whether to sell, renew, or keep something.

Spotted in the wild

NRMA's member magazine, Open Road, is exactly the kind of object this principle predicts people would rather hold onto than scroll past.

Cover of NRMA's Open Road magazine, Spring 2026 issue, showing the cover story Australia in Bloom: Four Spectacular Wildflower Road Trips and a Best Value Car Awards sidebar.

NRMA, Open Road magazine, Spring 2026 issue: cover story “Australia in Bloom: Four Spectacular Wildflower Road Trips,” plus a “Best Value Car Awards” sidebar. Photographed August 2026.

NRMA keeps printing and mailing this instead of moving to digital-only, consistent with a physical copy carrying the ownership pull this principle describes. That's a real company's ongoing choice, not a controlled test of the endowment effect itself. Full case: Printing costs more than digital. The paper copy still wins.

Kahneman, D., Knetsch, J. L., & Thaler, R. H. (1990). “Experimental Tests of the Endowment Effect and the Coase Theorem.” Journal of Political Economy, 98(6), 1325–1348
The classic Cornell coffee-mug experiment, replicated across several trials within the same paper.
StrengthUsed real goods and real money, with random assignment to seller, buyer, or chooser roles, and repeated the core comparison across several independent trials with different objects (mugs, pens, chocolate), not a hypothetical vignette, and not a one-off result.
WeaknessThe exact reservation-price gap varied somewhat across the paper's several replications rather than landing on one clean number each time, worth being upfront that the robust, oft-replicated finding is the direction and rough size (sellers demanding roughly twice what buyers offer), not one precise dollar figure.
Key findings
Verbatim“The median selling prices in the mug and pen markets were more than twice the median buying prices.”
Verbatim“The close similarity of results for buyers and choosers indicates that there was relatively little reluctance to pay for the mug.”
See also: Loss Aversion, the general asymmetry this is a specific case of: giving something up is felt more sharply than gaining the identical thing, ownership or not. Also Mental Accounting and IKEA Effect, the same family, applied to how money and self-made objects get valued.
39

IKEA Effect

Why do you love the wobbly shelf you built more than a better one you bought?

Assembling or partly creating something yourself makes you value the finished result more than an identical item you didn't build: labour itself generates attachment, independent of how good the result actually is.

PEOPLE RATE A SHELF THEY BUILT HIGHER THAN A PRE-BUILT ONE PRE-ASSEMBLED, DIDN’T BUILD IT 2.50 RATING OUT OF 7 SELF-ASSEMBLED, BUILT IT THEMSELVES 3.81 RATING OUT OF 7 The identical piece of furniture. Only who built it differed.
Likely mechanismPutting in the effort makes the finished result feel more valuable than it objectively is

The psychology. Effort invested becomes evidence of value through a kind of self-signalling: successfully completing a task creates a sense of authorship and competence, and that feeling gets folded into how the finished object itself is perceived, not just pride in the process, but a genuinely inflated read of the product's quality.

The full write-up: study, numbers, and caveats
The psychology

Effort invested becomes evidence of value through a kind of self-signalling: successfully completing a task creates a sense of authorship and competence, and that feeling gets folded into how the finished object itself is perceived, not just pride in the process, but a genuinely inflated read of the product's quality.

Where it causes errors

A business can offload real assembly or setup work onto the customer under the banner of “customisation,” at full price, and the customer ends up defending, and overvaluing, the result more than they rationally would, purely because they did the labour, not because the product is actually better.

Where it can help

Letting customers genuinely customise or assemble part of a product (a meal kit, a build-your-own bundle) can create authentic satisfaction and a real sense of fit, not just manufactured attachment, as long as the labour is optional and clearly doesn't substitute for a fair price.

Norton, M. I., Mochon, D., & Ariely, D. (2012). “The IKEA Effect: When Labor Leads to Love.” Journal of Consumer Psychology, 22(3), 453–460
Four studies, IKEA boxes, origami, and Lego sets, testing where the effect holds and where it breaks.
StrengthTested the effect across four different build tasks to check it wasn't specific to one kind of object, and included a crucial boundary condition: the elevated valuation disappeared when a participant's creation was later destroyed or left incomplete, isolating successful completion, not effort alone, as the actual driver.
WeaknessValuations were self-reported liking and willingness-to-pay for low-stakes household objects in a lab or online setting; whether the same size effect holds for expensive or high-stakes purchases wasn't directly tested here.
Key findings
Verbatim“Builders reporting greater liking (M = 3.81, SD = 1.56) than non-builders (M = 2.50, SD = 1.03), t(50) = 3.58, p < .001.”
Verbatim (abstract)“Our account suggests that labor leads to increased valuation only when labor results in successful completion of tasks; thus when participants built and then destroyed their creations, or failed to complete them, the IKEA effect dissipated.”
Also worth citing: Ling, I.-L., Liu, Y.-F., Lin, C.-W., & Shieh, C.-H. (2020). “Exploring IKEA Effect in Self-Expressive Mass Customization: Underlying Mechanism and Boundary Conditions.” Journal of Consumer Marketing, 37(4), 365–374. Extends the original assembly-based finding to design-choice customisation tools specifically: offering more real choice in a customisation toolkit gave shoppers more room for self-expression, which in turn raised how much they valued the finished product.
See also: Endowment Effect, related but distinct: that's valuation inflated by mere ownership; this is valuation inflated specifically by the labour of creating something.
40

Ostrich Effect

Why do people stop checking their bank balance exactly when they should check it most?

People selectively avoid checking information that might be bad news (an account balance, a bill, a test result) even though not looking doesn't change the underlying reality, only their awareness of it.

PEOPLE CHECK THEIR BALANCE LESS WHEN THE MARKET IS DOWN MARKET IS UP CHECKS BALANCE OFTEN MARKET IS DOWN LOGINS DROP OFF Same account, same information available. Only whether the news was likely good changed.
Likely mechanismInformation avoidance to dodge anticipated negative emotion

The psychology. Anticipated negative emotion from confirming bad news is unpleasant enough that avoiding the information itself becomes a coping strategy, even when the information is freely available and genuinely actionable. It's a short-term emotional relief traded for worse decision quality, since not knowing doesn't stop the underlying problem from getting worse.

The full write-up: study, numbers, and caveats
The psychology

Anticipated negative emotion from confirming bad news is unpleasant enough that avoiding the information itself becomes a coping strategy, even when the information is freely available and genuinely actionable. It's a short-term emotional relief traded for worse decision quality, since not knowing doesn't stop the underlying problem from getting worse.

Where it causes errors

Investors who stop checking their portfolio during a downturn miss the window to rebalance or cut losses; someone avoiding a subscription's usage-versus-cost breakdown keeps paying for something no longer worth it, purely because checking would confirm what they already suspect.

Where it can help

Products that proactively surface bad news in a calm, low-friction way, rather than requiring the user to go looking for it, route around the ostrich effect entirely. A bank sending a gentle “your balance is lower than usual” nudge does the checking for the person who would otherwise avoid it.

Karlsson, N., Loewenstein, G., & Seppi, D. (2009). “The Ostrich Effect: Selective Attention to Information.” Journal of Risk and Uncertainty, 38(2), 95–115
Real brokerage and retirement-account login records from Scandinavian and American investors.
StrengthUsed actual account-login records rather than self-reported checking behaviour, across two independent datasets from different markets, tracking what investors actually did, not what they said they'd do.
WeaknessThe paper reports the relationship as a regression coefficient rather than a single before-and-after percentage: in the Vanguard data, a 1 percentage point rise in the prior four-day return was associated with 18,000 to 23,000 extra daily logins, about 5 to 6% of the average daily total. That's a real, precise effect size, but it doesn't translate into one clean “X% more likely to check” headline figure.
Key findings
Verbatim (abstract)“In both datasets, investors monitor their portfolios more frequently in rising markets than when markets are flat or falling.”
ParaphrasedThe pattern held across two independent datasets, Scandinavian and American investors, suggesting it isn't specific to one brokerage platform or one market.
41

Order Effect

Why does “smart, then difficult” sound different from “difficult, then smart”?

The same information about a person or product, presented in a different order, produces a different overall impression: whatever arrives first anchors how everything after it gets read.

THE SAME SIX TRAITS SEEM HAPPIER WHEN GOOD ONES COME FIRST "INTELLIGENT..." READ FIRST 32% CALLED THEM "HAPPY" "ENVIOUS..." READ FIRST 5% CALLED THEM "HAPPY" Identical six traits, only the order reversed. What arrived first anchored how the rest read.
Likely mechanismEarly information anchors how later information gets read

The psychology. Early information gets used to build an initial hypothesis about a person or thing, and later information is then interpreted through that frame rather than weighed on equal terms. The same fact lands differently depending on whether it confirms or complicates the story already forming.

The full write-up: study, numbers, and caveats
The psychology

Early information gets used to build an initial hypothesis about a person or thing, and later information is then interpreted through that frame rather than weighed on equal terms. The same fact lands differently depending on whether it confirms or complicates the story already forming.

Where it causes errors

A performance review that opens with weaknesses before strengths leaves a more negative overall impression than the identical content in reverse order. Nothing in either version is untrue, but the read is different, which is exactly why the “compliment sandwich” exists as a workaround.

Where it can help

Deliberately leading with the most decision-relevant, honestly-earned positive information, when that's genuinely the right place to start, can prevent an unfairly negative first impression from colouring an evaluation, without hiding or misrepresenting anything that comes after.

Asch, S. E. (1946). “Forming Impressions of Personality.” Journal of Abnormal and Social Psychology, 41(3), 258–290
The founding demonstration that reordering identical facts changes the overall impression they create.
StrengthOne of the earliest controlled demonstrations that identical factual content, simply reordered, changes a global impression, not just which details people remember, but their overall evaluative read of a person.
WeaknessThe original method used short, decontextualised adjective lists read aloud rather than realistic information (a real conversation, a real review); later replication attempts have found the effect size for this specific paradigm is smaller and more context-dependent than the original 1946 results suggested.
Key findings
Verbatim“The impression produced by A is predominantly that of an able person who possesses certain shortcomings which do not, however, overshadow his merits. On the other hand, B impresses the majority as a ‘problem,’ whose abilities are hampered by his serious difficulties.”
ParaphrasedOn the specific trait “happy,” 32% of participants who heard the positive-first ordering later described the person that way, versus only 5% of participants who heard the identical traits in negative-first order.
See also: Ordering Effects, related but different: that's about which item gets selected from a set by its position in it; this is about how information order changes the judgement of a single target.
42

Shooting the Messenger

Why do people blame the doctor who diagnoses them, not the disease?

People who deliver bad news get judged more harshly, seen as less likeable and less trustworthy, even when everyone agrees they had nothing to do with causing the bad news.

PEOPLE RATE A MESSENGER LESS LIKEABLE FOR BAD NEWS + ASSIGNED TO GIVE GOOD NEWS RATED LIKEABLE ASSIGNED TO GIVE BAD NEWS RATED LESS LIKEABLE Neither messenger chose the news they carried. Only one of them got blamed for it.
Likely mechanismBad news creates a bad feeling, and that feeling gets pinned on whoever delivered it

The psychology. Bad news creates an unpleasant emotional state, and that negative feeling gets misattributed onto whoever is socially linked to delivering it, an automatic transfer of blame from the news itself onto its bearer, even when the bearer's innocence is explicit and known to everyone involved.

The full write-up: study, numbers, and caveats
The psychology

Bad news creates an unpleasant emotional state, and that negative feeling gets misattributed onto whoever is socially linked to delivering it, an automatic transfer of blame from the news itself onto its bearer, even when the bearer's innocence is explicit and known to everyone involved.

Where it causes errors

A support agent, doctor, or manager who has to relay unwelcome news (a delay, a diagnosis, a rejection) absorbs a likability penalty they didn't earn, which can push people to avoid or delay delivering necessary bad news altogether, at the cost of the recipient finding out later, and worse.

Where it can help

The same research found the penalty shrinks when a messenger explicitly signals benevolent intent. Honestly framing bad news around genuine care for the recipient, “I'm telling you now because I think you'd want to know,” is a legitimate way to soften an unfair reaction, not a way to spin the news itself.

John, L. K., Blunden, H., & Liu, H. (2019). “Shooting the Messenger.” Journal of Experimental Psychology: General, 148(4), 644–666
Eleven experiments testing whether innocent messengers of bad news get penalised, and why.
StrengthEleven separate experiments, an unusually high replication bar for one paper, systematically ruled out alternative explanations, showing the effect is specific to the actual messenger and distinct from simply disliking the bad news itself.
WeaknessStudy 1's stakes were a real but small $2 bonus, and several of the later studies (including a skin-cancer biopsy scenario) ask participants to imagine themselves in a hypothetical situation rather than living through a real one, a genuine limit on how much the effect size generalises to consequential real-world bad news.
Key findings
Verbatim“The messenger was judged to be significantly less likeable when delivering bad news relative to good news (Mbad = 6.26, SD = 2.34; Mgood = 7.21, SD = 2.08), t(239) = 3.32, p = .001.”
Verbatim“Study 2A demonstrates the specificity of the effect: dislike is directed at innocent messengers of bad news, and not innocent bystanders.”
43

Symbolic Rewards

Why does a badge worth literally nothing make people work harder?

A token of recognition with no monetary or material value (a badge, an award, a public “thank you”) increases effort and performance anyway, purely through the status and social recognition it carries.

A WORTHLESS BADGE STILL MAKES PEOPLE EDIT MORE NO BARNSTAR GIVEN BASELINE EDIT RATE GIVEN A BARNSTAR +60% EDIT RATE Worth nothing, spendable nowhere. The recognition itself moved real behaviour.
Likely mechanismSocial-esteem motivation: the recognition itself is the reward

The psychology. People are motivated by social esteem, not just material payoff. A symbolic reward publicly signals competence and contribution to a community, and people work to earn and maintain that signal even when it can't be spent, traded, or redeemed for anything. The recognition itself is the reward.

The full write-up: study, numbers, and caveats
The psychology

People are motivated by social esteem, not just material payoff. A symbolic reward publicly signals competence and contribution to a community, and people work to earn and maintain that signal even when it can't be spent, traded, or redeemed for anything. The recognition itself is the reward.

Where it causes errors

Symbolic recognition can be substituted for real compensation or resources (a badge, a certificate, an “Employee of the Month” photo), extracting genuine extra effort from people while giving up nothing of material value in return. That's a real risk of exploitation when it's used to avoid paying for labour rather than as a genuine complement to fair treatment.

Where it can help

In genuinely voluntary, non-exploitative contexts (open-source projects, volunteer work, peer recognition inside a team that's already fairly paid), a well-placed symbolic reward is a low-cost, honest way to make existing contributions visible and appreciated, without needing to gamify or manipulate anything. It works because the recognition is real, not manufactured.

Restivo, M., & van de Rijt, A. (2012). “Experimental Study of Informal Rewards in Peer Production.” PLOS ONE, 7(3), e34358
A randomised field experiment on real Wikipedia contributors, not a lab task or survey.
Experiment Teardown
200 top-1% Wikipedia contributors, none previously awarded a barnstar
Only whether a barnstar was awarded differed
No barnstar
No public recognition given for existing high-quality edits
Given a barnstar
Awarded a barnstar, a public, symbolic badge with no monetary value
Result None Given Swing
Edits (90 days) +60%
Awarded again 2% 12% +10

A badge with zero monetary value nearly doubled editing activity. Recipients edited 60% more over the next 90 days, and other community members, entirely unprompted, went on to award them six times as many additional barnstars of their own.

Internal validityA true randomised field experiment (n=200, 100/100 split) measured against real editing behaviour, not self-reported effort or a lab task.
External validityTested on already-top-1%-productive editors on one platform; whether the same size effect holds for average contributors, or on platforms without Wikipedia's particular reputation culture, wasn't tested here.
Key findings
Verbatim (abstract)“Receiving a barnstar increases productivity by 60% and makes contributors six times more likely to receive additional barnstars from other community members, revealing that informal rewards significantly impact individual effort.”
Verbatim“Twelve experimental subjects were subsequently awarded one or more barnstars from other contributors, compared to two subjects in the control group.”
Also worth citing: Kosfeld, M., & Neckermann, S. (2011). “Getting More Work for Nothing? Symbolic Awards and Worker Performance.” American Economic Journal: Microeconomics, 3(3), 86–99, a field experiment with paid workers on a real data-entry task, told in advance that top performers would receive a purely symbolic congratulatory card. Performance rose by about 12% on average, corroborating the same mechanism in a paid workplace rather than a volunteer community.
See also: Medium Maximisation, related but different: that's a token that stands in for a real reward and gets over-optimised for its own sake; this is a token with no exchange value at all, that still motivates through recognition alone.
44

Scarcity

Why does “only 2 left” make you want something you weren't even sure about?

Making something look limited in quantity or time left makes it more wanted and more likely to be chosen right now, even when nothing about the thing itself has changed at all.

A JAR WITH ONLY 2 COOKIES IS RATED MORE DESIRABLE JAR WITH 10 COOKIES RATED LESS DESIRABLE JAR WITH 2 COOKIES RATED MORE DESIRABLE Same cookies, same jar shape, same shop. Only the count left inside changed.
Likely mechanismScarcity read as an inferred demand signal

The psychology. Scarcity is read as a signal, not just a constraint. If something is hard to get, people infer other people must want it too, and a supply that's visibly shrinking reads as stronger proof of rising demand than a supply that was simply always small, even though neither one says anything about the object's actual quality.

The full write-up: study, numbers, and caveats
The psychology

Scarcity is read as a signal, not just a constraint. If something is hard to get, people infer other people must want it too, and a supply that's visibly shrinking reads as stronger proof of rising demand than a supply that was simply always small, even though neither one says anything about the object's actual quality.

Where it causes errors

The signal doesn't require a real shortage to fire. A countdown timer, an “only 2 left” label, or a flash sale works on the same mechanism whether or not supply is actually constrained, which is exactly why fabricated urgency claims have become one of the most commonly cited dark patterns, and have drawn direct regulatory action when the underlying claim wasn't true.

Where it can help

Genuine scarcity (a real registration deadline, a real capacity limit, an actual last-batch product run) is useful information, and surfacing it honestly helps people decide in time instead of endlessly deferring a decision they'd otherwise assume they could make later. The harm was never the information; it's inventing the shortage.

Worchel, S., Lee, J., & Adewole, A. (1975). “Effects of Supply and Demand on Ratings of Object Value.” Journal of Personality and Social Psychology, 32(5), 906–914
A controlled two-experiment lab study, not a real purchase or field setting.
StrengthManipulated the actual physical supply of an identical object, the same cookies, the same jar, only the count inside it changed, isolating scarcity itself as the cause, rather than a claim about scarcity layered on top of some other difference.
WeaknessA single-session lab task using hypothetical desirability ratings, not real purchases with real money on the line, on a sample of female undergraduates in 1975; whether the same size effect holds for actual buying decisions, or across a much broader population today, wasn't tested here.
Key findings
ParaphrasedIdentical cookies drawn from a jar holding only two, versus a jar holding ten, were rated as significantly more desirable and attractive. Participants tasted the cookie in every condition, and taste ratings themselves never differed: only the desirability and cost ratings moved with supply.
Verbatim (abstract)“Cookies were rated as more valuable when their supply changed from abundant to scarce than when they were constantly scarce.”
Also worth citing: Aggarwal, P., Jun, S. Y., & Huh, J. H. (2011). “Scarcity Messages.” Journal of Advertising, 40(3), 19–30, a more recent consumer study finding limited-quantity messages (“only 2 left”) outperform limited-time messages (“sale ends soon”) at raising purchase intent, mediated by a felt sense of competition with other buyers, the same two claims shown side by side on the flash-sale card above.
Seen live: Field Session: the high school money talk, where this ran alongside Anchoring and Social Proof as one of three “spending traps” on the same slide.
See also: Anchoring and Social Proof, the three most commonly paired urgency tactics on real checkout and sale pages, each working through a different mechanism but often deployed together.
Decoded on The Science Behind: How does Up Bank's behavioural design pay off for both the bank and its customers?, where a real, published prize pool counts down live on screen rather than just being described as limited.

A real example: a “flash sale ends in 09:47, only 2 left” product card shown to a room of Year 8/9 students, printed next to its two cousins on the same slide: social proof (“1,200+ bought this month”) and anchoring (“was $160, now $80”).

45

Checklists

Why did a simple paper checklist cut surgical deaths almost in half?

A short list of steps, read aloud, measurably reduces errors in complex, high-stakes tasks, not by teaching anyone anything new, but by protecting skilled people from the predictable ways complexity causes a step to get skipped.

READING A CHECKLIST ALOUD CUT SURGICAL DEATHS NEARLY IN HALF DEATHS: 1.5% BEFORE THE CHECKLIST COMPLICATIONS: 11% DEATHS: 0.8% CHECKLIST READ ALOUD COMPLICATIONS: 7% 8 hospitals, 4 continents, the same surgeries and surgeons.
Likely mechanismWriting steps down catches the slip a busy, tired mind would otherwise miss

The psychology. As a task gets more complex, the volume of things that must be remembered and executed correctly, under time pressure, interruption, and fatigue, exceeds what working memory can reliably hold, even for experts who could recite every individual step from memory. A checklist doesn't add knowledge; it externalises memory, catching the predictable slip rather than the exotic one, and frees up attention for judgment instead of recall.

The full write-up: study, numbers, and caveats
The psychology

As a task gets more complex, the volume of things that must be remembered and executed correctly, under time pressure, interruption, and fatigue, exceeds what working memory can reliably hold, even for experts who could recite every individual step from memory. A checklist doesn't add knowledge; it externalises memory, catching the predictable slip rather than the exotic one, and frees up attention for judgment instead of recall.

Where it causes errors

A checklist completed as a box-ticking formality, initialled without the verbal exchange or genuine pause it's meant to prompt, can create false confidence with none of the real protection. Mandating the paper without the practice is exactly what a large real-world rollout found: no measurable drop in harm when a checklist becomes a compliance step instead of a team conversation.

Where it can help

Used as intended, read aloud, out loud, as the trigger for an actual conversation between the people in the room, not just a form to initial, a checklist is one of the few low-cost, low-tech interventions with a directly measured drop in real-world harm. The idea traces back to a single 1935 index card, written after a fatal Boeing test-flight crash. It now spans aviation cockpits and operating rooms alike.

Haynes, A. B., Weiser, T. G., Berry, W. R., et al. (2009). “A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population.” New England Journal of Medicine, 360(5), 491–499
A prospective before-and-after study across eight real hospitals on four continents, not a lab task.
Experiment Teardown
3,733 patients having major surgery across 8 hospitals in 8 countries
Same hospitals, same teams, only whether the checklist was used differed
Before
Standard care, no structured checklist
After
19-item WHO Surgical Safety Checklist, read aloud at sign-in, time-out, and sign-out
Result Before After Swing
Death rate 1.5% 0.8% −0.7
Complication rate 11% 7% −4

A 19-item card cut the death rate almost in half. Complications fell by more than a third, across a mix of hospitals ranging from a rural Tanzanian district hospital to a major Seattle medical centre, the same simple card, working in wildly different settings.

Internal validityA large, real, prospectively-collected clinical sample (n=3,733 before, n=3,955 after) at the same sites with the same teams, not a simulation or a lab task.
External validityA before-after design, not a randomised trial, so secular trends and a Hawthorne effect from being observed can't be fully ruled out; and a much larger 2014 rollout across 101 Ontario hospitals (109,341 procedures) found no significant change in mortality or complications when the same checklist was mandated without the same implementation support.
Key findings
Verbatim (abstract)“The rate of death was 1.5% before the checklist was introduced and declined to 0.8% afterward (P = 0.003). Inpatient complications occurred in 11.0% of patients at baseline and in 7.0% after introduction of the checklist (P<0.001).”
ParaphrasedThe improvement held across sites with very different baseline resources and complication rates, from a low-income-country district hospital to a high-income academic medical centre, and no single site accounted for the overall effect. The paper itself calls the exact mechanism unclear and “most likely multifactorial,” spanning both systemic changes (like moving antibiotic administration into the operating room) and changes in team behaviour.
Also worth citing: Urbach, D. R., Govindarajan, A., Saskin, R., Wilton, A. S., & Baxter, N. N. (2014). “Introduction of Surgical Safety Checklists in Ontario, Canada.” New England Journal of Medicine, 370(11), 1029–1038, a population-level natural experiment across 101 hospitals (109,341 procedures before, 106,370 after) found mortality (0.71% → 0.65%) and complications (3.86% → 3.82%) barely moved and weren't statistically significant, once the checklist was a mandated form rather than a trained, supported practice.
See also: Illusion of Explanatory Depth, both are about the gap between feeling you've got a complex process covered and actually verifying it step by step; and Noise, where a checklist is one of the standard tools recommended for making inconsistent expert judgment more consistent.
46

False Positives

What if the exciting result from your pilot is just noise wearing a lab coat?

A handful of small, individually reasonable choices in how a result is measured and analysed (when did we stop collecting, which groups did we compare, which outcome did we report) can turn pure chance into something that looks statistically real.

A FEW FLEXIBLE CHOICES TURN 5% FALSE POSITIVES INTO 60%+ 5% ONE FIXED ANALYSIS PLAN FALSE-POSITIVE RATE 60%+ A FEW "DEFENSIBLE" CHOICES ADDED FALSE-POSITIVE RATE Nobody has to cheat for this to happen. Every single choice felt reasonable at the time.
Likely mechanismEnough small, reasonable analysis choices will eventually turn up a “significant” result by chance alone

The psychology. Nobody has to cheat for this to happen. Given enough small, defensible choices (which week to start counting, which segment to look at, which of several metrics to report), chance alone will hand you something that clears the bar for “significant” far more often than the 5% everyone assumes. The result feels earned because every individual decision felt reasonable at the time.

The full write-up: study, numbers, and caveats
The psychology

Nobody has to cheat for this to happen. Given enough small, defensible choices (which week to start counting, which segment to look at, which of several metrics to report), chance alone will hand you something that clears the bar for “significant” far more often than the 5% everyone assumes. The result feels earned because every individual decision felt reasonable at the time.

Where it causes errors

A pilot gets greenlit for a full rollout on the strength of a result that was never really there, and because nobody deliberately faked anything, the team has no idea the win was noise until it fails to reappear at scale.

Where it can help

Pre-registering exactly what you'll measure and how, before you see the data, closes off the undisclosed flexibility that manufactures false positives. The fix isn't more caution, it's committing to the analysis plan in advance.

Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). “False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant.” Psychological Science, 22(11), 1359–1366
A simulation plus two real demonstration experiments, not a single anecdote.
StrengthThe simulation isolates exactly how much a few common analysis choices inflate the false-positive rate, and the demonstration experiments show the same flexibility manufacturing a result everyone would recognise as impossible.
WeaknessThe demonstration studies used small samples (around 20 participants per cell) deliberately, to show how easy manipulation is at that scale, not to claim the specific “findings” themselves are real effects worth replicating.
Key findings
Verbatim“...the bottom row reporting the false-positive rate if the researcher uses all of these degrees of freedom, a practice that would lead to a stunning 61% false-positive rate! A researcher is more likely than not to falsely detect a significant effect by just using these four common researcher degrees of freedom.”
Verbatim“According to their birth dates, people were nearly a year-and-a-half younger after listening to ‘When I’m Sixty-Four’ (adjusted M = 20.1 years) rather than to ‘Kalimba’ (adjusted M = 21.5 years), F(1, 17) = 4.92, p = .040.”
See also: Noise, on inconsistency in judgment even without any of this flexibility being exploited deliberately; and the Voltage Effect field session, where this is one of the “vital signs” checked before scaling a pilot.
47

Representativeness (Who was actually tested)

What if the people your pilot tested on look nothing like everyone else you're about to scale to?

A result only describes the population it was measured on, and the people who end up in most trials, pilots, and studies are a poor stand-in for the people a full rollout will actually reach.

STUDY SUBJECTS ARE 96% WESTERN, JUST 12% OF THE WORLD STUDY SUBJECTS 96% WESTERN, INDUSTRIALISED WORLD POPULATION 12% OF EVERYONE Psychology undergrads were the sole subject pool in two-thirds of US studies.
Likely mechanismSelection bias in who opts into a pilot

The psychology. Recruiting for a pilot is itself a selection process, not a neutral step: whoever is easiest to reach, already engaged, or already opted in ends up over-represented, and that group's behaviour becomes the evidence base for a rollout meant for everyone else too.

The full write-up: study, numbers, and caveats
The psychology

Recruiting for a pilot is itself a selection process, not a neutral step: whoever is easiest to reach, already engaged, or already opted in ends up over-represented, and that group's behaviour becomes the evidence base for a rollout meant for everyone else too.

Where it causes errors

A feature that tests brilliantly with an early-adopter beta group can fail with the mainstream user base whose needs, patience, and technical comfort were never in the sample. The pilot wasn't wrong, it just answered a narrower question than the one being asked of it.

Where it can help

Deliberately over-sampling the least convenient, least engaged, or most sceptical segment during a pilot, rather than the easiest volunteers, gives a far more honest preview of what a full rollout will actually face.

Henrich, J., Heine, S. J., & Norenzayan, A. (2010). “The Weirdest People in the World?” Behavioral and Brain Sciences, 33(2–3), 61–83
A review of the subject pools behind published psychology research, not a single study.
Key findings
Verbatim“This means that 96% of psychological samples come from countries with only 12% of the world's population.”
Verbatim“67% of the American samples (and 80% of the samples from other countries) were composed solely of undergraduates in psychology courses (Arnett 2008).”
See also: Lab vs. Field, on this site's Reading the Research page, for the related gap between how people behave in a study and how they behave for real; and the Voltage Effect field session.
48

Chef or the Ingredients

What if the pilot worked because of who ran it, not what it was?

A pilot staffed with your very best people proves your very best people are good at their jobs, not that the idea itself will work once it's handed to an average team at real scale.

ORDINARY TEACHERS, NOT STAR TEACHERS, STILL PRODUCED THE GAIN USUAL PILOT: BEST PEOPLE RUN IT WORKS, BUT WAS IT THE PROGRAMME? CHECC: ORDINARY LOCAL TEACHERS +0.22 SD GAIN HELD ANYWAY Ordinary staff, not star staff. The gain still held.
Likely mechanismSuccess gets credited to the idea, when it may really be the exceptional team that ran it

The psychology. It's natural to give a new idea its best shot by putting your strongest people on it, but that quietly changes the question being tested. “Does this work?” and “does this work when run by exceptional people?” can have completely different answers, and only one of them is the question a full rollout actually needs answered.

The full write-up: study, numbers, and caveats
The psychology

It's natural to give a new idea its best shot by putting your strongest people on it, but that quietly changes the question being tested. “Does this work?” and “does this work when run by exceptional people?” can have completely different answers, and only one of them is the question a full rollout actually needs answered.

Where it causes errors

A training programme, a sales script, or a support process that shines in a pilot run by your top performers can fall flat the moment it's handed to the average employee. The “ingredients” were never the whole recipe; the “chef” was doing more of the work than anyone credited.

Where it can help

Deliberately staffing a pilot with an average, representative team, not your best, costs you a flashier pilot result, but buys you a real answer to whether the idea itself, not the people running it, is what's working.

List, J. A., Fryer, R. G., Levitt, S. D., & Samek, A.: the Chicago Heights Early Childhood Center (CHECC), documented across several NBER working papers including NBER Working Paper No. 21477
A real early-childhood field experiment, funded by a $10M Griffin Foundation grant, not a lab task.
Key findings
ParaphrasedThe CHECC preschool was deliberately staffed with a representative sample of ordinary local public-school teachers, not hand-picked star teachers, specifically so its result would still hold once the programme left the pilot.
ParaphrasedAt the 9-month follow-up, children in the programme showed a real gain over the control group, roughly +0.22 SD on early academic skills and +0.16 SD on executive functioning, a result attributable to the design, not to unusually gifted staff.
See also: Survivorship Bias, on the related trap of learning only from the successes that happened to make it through; and the Voltage Effect field session, where this exact study anchors the “very best” trap.
49

Spillovers

What if helping the people in your pilot means quietly hurting everyone around them?

A pilot too small to move the surrounding system can hide an effect that only appears once an idea is running at real scale, because scale changes the system the idea operates inside of.

HELPING JOB SEEKERS AT SCALE DISPLACES OTHER JOB SEEKERS LOW SATURATION, FEW TREATED REAL GAIN NET EMPLOYMENT EFFECT HIGH SATURATION, MOST TREATED NEAR ZERO NET EMPLOYMENT EFFECT ~30,000 job seekers, sharing the same jobs whether treated or not.
Likely mechanismScaled up, a programme starts competing directly with the people it didn't include

The psychology. A small pilot is, almost by definition, too small to change the market or system around it. Scale it up, and the people who weren't part of the programme are still sharing the same jobs, the same fares, the same shelf space, so a benefit to the treated group can come directly at the expense of everyone else competing for the same thing.

The full write-up: study, numbers, and caveats
The psychology

A small pilot is, almost by definition, too small to change the market or system around it. Scale it up, and the people who weren't part of the programme are still sharing the same jobs, the same fares, the same shelf space, so a benefit to the treated group can come directly at the expense of everyone else competing for the same thing.

Where it causes errors

A programme that looks like a clear win in a small trial can look close to worthless once it's rolled out to everyone, because the gain it produces for participants is offset, unmeasured, by a loss for the non-participants they were quietly competing against.

Where it can help

Testing at varying levels of saturation, some markets at low intensity, some at high, rather than one uniform small pilot, is one of the few designs that can actually catch a spillover before a full rollout does.

Crépon, B., Duflo, E., Gurgand, M., Rathelot, R., & Zamora, P. (2013). “Do Labor Market Policies Have Displacement Effects? Evidence from a Clustered Randomized Experiment.” The Quarterly Journal of Economics, 128(2), 531–580
A clustered field experiment across roughly 235 French labour markets, not a single-site trial.
StrengthEach labour market was randomly assigned a different treatment saturation, from 0% to 100% of eligible young job seekers receiving job-placement assistance, a design specifically built to catch effects a single-site pilot never could.
WeaknessSpecific to one labour-market intervention in one country; the size of the displacement effect elsewhere depends on how tight the local job market is, something this design can flag but not generalise a single number for.
Key findings
ParaphrasedNearly 30,000 young, college-educated, long-term-unemployed job seekers took part. Treated job seekers found stable work faster than untreated ones in the same market.
Verbatim (abstract)“These gains are transitory, and they appear to have come partly at the expense of eligible workers who did not benefit from the program, particularly in labor markets where they compete mainly with other educated workers, and in weak labor markets. Overall, the program seems to have had very little net benefits.”
See also: When Interventions Backfire, on this site's Reading the Research page; and the Voltage Effect field session.
50

Cost Traps (Diminishing Returns at Scale)

What if the same idea, at real scale, buys you a fraction of what the pilot promised?

An effect that looked strong and cheap in a small, closely-run pilot routinely shrinks once it's handed to a real operating team running it for real, so the true cost of each unit of impact is far higher than the pilot implied.

NUDGES SHRINK 80% ONCE RUN AT REAL GOVERNMENT SCALE +8.7 PTS PUBLISHED IN JOURNALS 33% RELATIVE LIFT +1.4 PTS RUN AT REAL SCALE 8% RELATIVE LIFT 126 government nudge trials, roughly 23 million people.
Likely mechanismFounder-team expertise that doesn't transfer to scale

The psychology. A pilot is usually run by the people who designed it: motivated, closely watching the details, free of the workarounds and shortcuts a busy real-world team adopts under normal pressure. None of that expertise or attention transfers automatically once the same idea is handed off to be run at scale by people who didn't design it.

The full write-up: study, numbers, and caveats
The psychology

A pilot is usually run by the people who designed it: motivated, closely watching the details, free of the workarounds and shortcuts a busy real-world team adopts under normal pressure. None of that expertise or attention transfers automatically once the same idea is handed off to be run at scale by people who didn't design it.

Where it causes errors

A budget built on the pilot's cost-per-outcome routinely blows out at scale, not because anyone did anything wrong, but because the pilot's efficiency was partly a feature of being small, closely supervised, and run by people with a personal stake in it working.

Where it can help

Budgeting a rollout against the effect size a real operating team is likely to achieve, not the effect size the pilot team achieved, is a more honest starting point than assuming the pilot's numbers travel unchanged.

DellaVigna, S., & Linos, E. (2022). “RCTs to Scale: Comprehensive Evidence from Two Nudge Units.” Econometrica, 90(1), 81–116
A meta-analysis of 126 real government-run nudge trials, not a single case study.
Key findings
Verbatim (abstract)“In papers published in academic journals, the average impact of a nudge is very large, an 8.7 percentage point take-up effect, a 33.5% increase over the average control. In the Nudge Unit trials, the average impact is still sizable and highly statistically significant, but smaller at 1.4 percentage points, an 8.1% increase.”
ParaphrasedThat's roughly an 80–85% shrinkage in effect size between the published academic result and the same style of intervention running for real, the same budget buying a fraction of the impact the pilot literature implied.
Also worth citing: Al-Ubaydli, O., List, J. A., & Suskind, D. (2017). “What Can We Learn from Experiments? Understanding the Threats to the Scalability of Experimental Results.” American Economic Review: Papers & Proceedings, 107(5), 282–286, List's own academic treatment of why effects and returns decay at scale.
See also: The Voltage Drop at Scale, on this site's Reading the Research page; and the Voltage Effect field session, where all five of these “vital signs” are checked before a rollout.
51

Reciprocity

Why does a stranger's small, unasked-for favour make you agree to something you'd normally say no to?

An unsolicited gift or favour creates a felt obligation to give something back, often something worth far more than what was received, and even when the gift was never requested.

AN UNASKED-FOR COKE ROUGHLY DOUBLED TICKET SALES 1 NO FAVOUR, EMPTY-HANDED ~1 TICKET BOUGHT ON AVERAGE $ UNPROMPTED COKE GIVEN ~2X AS MANY TICKETS BOUGHT A small, unasked-for gift, from a stranger. It created a debt worth repaying.
Likely mechanismAccepting something creates a felt obligation to give back, often more than it cost

The psychology. Reciprocity norms evolved to make exchange and cooperation possible without a contract enforcing it: a society where favours get returned can trade and cooperate without every exchange being simultaneous or formally guaranteed. That leaves people with a genuine aversive feeling, a psychological debt, the moment they accept something, and the discomfort of carrying that debt is often enough to make people repay it with more than the original gift was worth, even when the gift was small, unsolicited, and given by someone they've just met.

The full write-up: study, numbers, and caveats
The psychology

Reciprocity norms evolved to make exchange and cooperation possible without a contract enforcing it: a society where favours get returned can trade and cooperate without every exchange being simultaneous or formally guaranteed. That leaves people with a genuine aversive feeling, a psychological debt, the moment they accept something, and the discomfort of carrying that debt is often enough to make people repay it with more than the original gift was worth, even when the gift was small, unsolicited, and given by someone they've just met.

Where it causes errors

A free sample, a promotional pen, a “complimentary” consultation, or an unsolicited gift mailed with a donation request all work the same lever: a two-dollar giveaway extracting a decision worth fifty times as much, from someone who never asked for the gift and might not have wanted it. Because the obligation is felt rather than reasoned through, it survives contact with a person's own better judgment: people report feeling they “have to” reciprocate even when they can articulate, out loud, that the exchange isn't a fair one.

Where it can help

The same mechanism, used without engineering an obligation on purpose, is just genuine goodwill: an unconditional gesture with no purchase requirement or ask attached tends to build real trust rather than manufactured debt. The clearest version of this on this site is a costly, specific apology after a real service failure. See Peak-End Rule and the Voltage Effect field session for the Uber research on why that works. The honest line: give first because it's the right thing to do, and let reciprocation happen or not. Don't manufacture a debt sized to what you intend to ask for later.

Regan, D. T. (1971). “Effects of a Favor and Liking on Compliance.” Journal of Experimental Social Psychology, 7, 627–639
A lab experiment with independently randomised favour and liking manipulations, and a real behavioural outcome (money spent) rather than a self-reported one.
Experiment Teardown
81 Stanford freshman men, paired with a confederate posing as a fellow participant, “Joe,” to rate paintings together
Only whether Joe returned from a break with an unasked-for Coke changed
No favour
Joe steps out for a short break and returns empty-handed
Unprompted favour
Joe returns with two Cokes, hands one over unasked, saying he got himself one and grabbed an extra
Result No favour Favour Swing
Subjects who bought more than one raffle ticket 25% 58% ~2×

The favour worked whether or not participants said they liked Joe. A weak, inconsistent liking effect sat alongside a strong, consistent favour effect. The debt from an unasked-for gift pulled people towards compliance even when their stated opinion of the person asking hadn't moved much at all. Reciprocity didn't need liking to do its work; it operated as its own lever.

Internal validityA real behavioural outcome, money spent on raffle tickets, not a self-reported intention, with the favour and liking manipulations independently randomised in the same design.
External validity77 Stanford freshman men in one lab in 1971; whether the same swing holds outside that population, or outside a favour delivered face-to-face by the same person making the later ask, wasn't tested here.
Key findings
Verbatim“the favor more than doubled the proportion of the subjects buying more than a single ticket, raising this proportion from 25% in the two control conditions to 58% in the Favor condition.”
Verbatim (abstract)“the relationship between favors and compliance is mediated, not by liking for the favor-doer, but by normative pressure to reciprocate.”
Verbatim“The lack of any difference in compliance between subjects who received no favor and those who received a favor from someone other than the requester allows us to reject the notion that simply receiving the soft drink might lead to greater compliance.”
Also worth citing: Strohmetz, D. B., Rind, B., Fisher, R., & Lynn, M. (2002). “Sweetening the Till: The Use of Candy to Increase Restaurant Tipping.” Journal of Applied Social Psychology, 32(2), 300–309, real restaurant field experiments, not a lab: a candy given with the check raised average tips from about 15.1% to 17.8%, and giving one candy, then unexpectedly returning with a second as a personal, spontaneous extra, outperformed giving both at once, roughly a 21% tip increase in that condition. The manner of the gift mattered as much as receiving one at all.
See also: Zero Price Effect, related but distinct: that's demand jumping at a $0 price point, regardless of who's giving it or how personal it feels; this is the felt obligation triggered specifically by an unsolicited personal gift, which is why a free sample can trigger both mechanisms in the same moment. See also the dentist field session, where free toothpaste at checkout does exactly that.
Tested as an experiment: Does a costly, specific apology beat a generic one after a service failure?, the costly-signal half of this mechanism, built on the Uber research above.
52

Good Friction (Asymmetric Effort)

Why should signing up for savings take one click, and buying on impulse take five?

The same tool as Sludge, aimed the other way: effort added on purpose to a choice you'd regret, and stripped away from one you'd endorse in a calmer moment.

A ONE-TAP SIGN-UP TOOK ENROLMENT FROM 3% TO 20% FRICTION ADDED TO WITHDRAWALS HARDER TO ACT ON IMPULSE 1 TAP FRICTION REMOVED FROM ENROLLING 20% ENROLLED, UP FROM 3% Same lever, opposite direction. Effort moved to the side that needed it.
Likely mechanismThe same pull towards the easier path right now that powers Sludge, aimed on purpose

The psychology. Sludge and Good Friction run on the identical mechanism: whichever path costs less immediate effort usually wins, almost regardless of which path actually serves the person walking it. Present bias and effort aversion don't check which side of someone's interest they're operating on. Good Friction doesn't fight that bias; it uses it on purpose. A few extra seconds or steps, placed exactly when temptation is hottest, buy slower judgment a chance to catch up before the impulsive choice locks in. Removed from the other direction, the same lever works just as hard the other way: a decision someone already wants to make, but keeps failing to start, collapses from several small decisions into one.

The full write-up: study, numbers, and caveats
The psychology

Sludge and Good Friction run on the identical mechanism: whichever path costs less immediate effort usually wins, almost regardless of which path actually serves the person walking it. Present bias and effort aversion don't check which side of someone's interest they're operating on. Good Friction doesn't fight that bias; it uses it on purpose.

A few extra seconds or steps, placed exactly when temptation is hottest, buy slower judgment a chance to catch up before the impulsive choice locks in. Removed from the other direction, the same lever works just as hard the other way: a decision someone already wants to make, but keeps failing to start, collapses from several small decisions into one.

Where it causes errors

The line between Good Friction and Sludge is who the friction serves, not how it looks from the outside. A “cooling-off” step framed as protecting the customer can just as easily be protecting the business from a refund, and a savings app that makes withdrawing your own emergency money take three days and a phone call has crossed from protective into obstructive, whatever it started out as.

The same confusion runs in reverse, too: stripping friction from the wrong side of a decision, a one-tap purchase confirmation, a buy-now-pay-later checkout engineered to feel like spending nothing at all, removes exactly the pause that would have let someone catch a purchase they'd regret. See the “safe to spend” field session for a real, cited case: making a banking app's balance look more available, not less, made a discretionary purchase 22.2% more likely.

Where it can help

Placed deliberately, the same lever cuts both ways on the same goal. Add a little friction to the choice you don't want to make easy: a short delay and a typed reason before money leaves a “locked” savings pot, the kind of intentional speed bump the “where to add friction back” review above argues for.

Strip it from the choice you do: collapsing a multi-step retirement enrolment into a single pre-filled “yes” raised participation by double digits (see the study below), and Irrational Labs found that asking one question during payroll onboarding, at the exact moment someone was already looking at their paycheck, lifted savings enrolment roughly 5x, from 3% to 20%. Neither example removed anyone's choice, they just moved the effort tax to the side of the decision that actually needed it.

Beshears, J., Choi, J. J., Laibson, D., & Madrian, B. C. (2013). “Simplification and Saving.” Journal of Economic Behavior & Organization, 95, 130–145
The field experiment behind “Quick Enrolment”: collapsing a multi-step 401(k) decision into a single yes/no, run inside two real US companies' live retirement plans.
StrengthA genuine field experiment run inside two employers' actual, live 401(k) systems. The outcome is real enrolment and contribution records, not a stated intention in a survey or a lab choice task.
WeaknessBoth employers already ran a standard plan with an employer match, so the effect is measured on a workforce already somewhat primed to consider saving; how much the same one-step form would move a population with no existing plan, or no match, wasn't tested here.
Key findings
Verbatim (abstract)“Individuals received an opportunity to enroll in a retirement savings plan at a pre-selected contribution rate and asset allocation, allowing them to collapse a multidimensional problem into a binary choice between the status quo and the pre-selected alternative. The intervention increases plan enrollment rates by 10 to 20 percentage points.”
Verbatim (abstract)“We find that a similar intervention can be used to increase contribution rates among employees who are already participating in a savings plan.”
Also worth citing: Kristen Berman / Irrational Labs, “How to Get America's Savings Rate Up”, not a peer-reviewed study, but a real, named applied case: Intuit added a single question, “How much would you like to automatically set aside each pay period?”, to payroll onboarding, at the moment employees were already looking at their paycheck. Self-reported result: savings enrolment roughly 5x, from 3% to 20%, and $1,000+ in average extra annual savings per person.
See also: Sludge, the same mechanism, pointed the other way: friction added in someone's interest here, against it there. The two aren't opposites so much as the same dial, and which one you're looking at depends entirely on who the friction protects.
See also: Smart Defaults, a different lever entirely: not how hard a decision is to make, but when it's made and what it's measured against.
Seen live: the “safe to spend” field session's “where to add friction back” review, and the envelope-saving “out of sight, out of mind” case from the high school talk, hidden savings that took real effort to reach were exactly the ones that survived.
53

Halo Effect

Why does one great impression get extended, unverified, to everything else about the same source?

A single strong, positive (or negative) trait quietly colours judgment of other, unrelated traits of the same person, brand, or product, even ones nobody actually checked.

ONE TRAIT CHANGES HOW PEOPLE SCORE TRAITS THEY NEVER CHECKED SEES ONE STRONG TRAIT FORMS ONE "SEEMS IMPRESSIVE" RATES SKILL AND CHARACTER BOTH SCORED HIGH, NEITHER CHECKED Only one trait was actually observed. The impression spread to traits people never checked.
Likely mechanismOne strong impression quietly stands in for judgments that were never actually made

The psychology. Rating several distinct qualities of the same target independently is cognitively expensive, so people substitute one easy question for several hard ones: instead of separately asking “how competent, how honest, how reliable is this,” they ask “how do I feel about this overall,” then let that single answer stand in for all the others. The first vivid, salient trait usually sets that overall feeling, and everything judged afterwards gets checked against it rather than against its own evidence.

The full write-up: study, numbers, and caveats
The psychology

Rating several distinct qualities of the same target independently is cognitively expensive, so people substitute one easy question for several hard ones: instead of separately asking “how competent, how honest, how reliable is this,” they ask “how do I feel about this overall,” then let that single answer stand in for all the others. The first vivid, salient trait usually sets that overall feeling, and everything judged afterwards gets checked against it rather than against its own evidence, not because people are lazy, but because holding several independent trait-judgments in mind at once is harder than it feels like it should be.

Where it causes errors

Phil Rosenzweig's business-book critique of Fortune's “World's Most Admired Companies” survey is the cleanest corporate example: respondents rate the same company's strategy, leadership, and culture glowingly when its stock price is up, and rate the identical practices as rigid and poorly led when the price falls, without having independently audited the strategy, leadership, or culture either time. Rosenzweig names ABB and IBM as real, documented cases of the swing.

The same error runs quieter in everyday judgment: a beautifully designed product gets assumed to be more reliable and more usable than a plain one, with no evidence either way (design researcher Don Norman calls this specific case the “aesthetic-usability effect”), and a well-dressed, articulate job candidate gets rated as more competent on skills the interview never actually tested.

Where it can help

The honest version of this isn't faking a halo, it's letting a real strength do legitimate work. A genuinely excellent, checkable experience at one visible touchpoint reasonably earns some trust in the parts of the relationship a customer hasn't tested yet; that's not manipulation, it's an imperfect but real signal, provided the trust it buys is actually deserved elsewhere too.

The other honest use is defensive. Organisations that know about the bias can structure evaluations to blunt it: rating one trait across every candidate before moving to the next trait, instead of rating one candidate across every trait before moving to the next candidate. That order alone stops a strong first impression on trait A from quietly deciding the score on unrelated trait B.

Thorndike, E. L. (1920). “A Constant Error in Psychological Ratings.” Journal of Applied Psychology, 4(1), 25–29
The paper that coined the term: real supervisors and commanding officers rating real subordinates, a 1915 sample from two large industrial firms, plus military ratings, on multiple, logically independent traits.
StrengthLarge-scale, consequential, real-world ratings, supervisors scoring actual subordinates on real personnel records, not a lab vignette with invented targets, giving the finding strong ecological validity for how evaluators behave when the rating actually matters.
WeaknessCorrelational, not experimental: Thorndike could show trait-correlations were implausibly high and uniform across pairs of traits that should vary independently, but couldn't rule out that raters had some genuine shared information linking them. A later controlled experiment (Nisbett & Wilson, 1977, below) closed that gap by manipulating only one trait while holding the “unrelated” ones physically identical.
Key findings
ParaphrasedRatings of traits that should be logically or practically independent, physical bearing and intelligence, or technical skill and character, correlated far more highly than plausible, suggesting raters were extending one general impression across specific judgments rather than assessing each trait on its own evidence.
ParaphrasedThe pattern held across both the 1915 industrial sample (supervisors rating employees on traits including intelligence, technical skill, and reliability) and the military ratings, indicating the error wasn't specific to one relationship type or rating context.
Verbatim“Their ratings were apparently affected by a marked tendency to think of the person in general as rather good or rather inferior and to color the judgments of the qualities by this general feeling. This same constant error toward suffusing ratings of special features with a halo belonging to the individual as a whole appeared in the ratings of officers made by their superiors in the army.”
Also worth citing: Nisbett, R. E., & Wilson, T. D. (1977). “The Halo Effect: Evidence for Unconscious Alteration of Judgments.” Journal of Personality and Social Psychology, 35(4), 250–256, the controlled follow-up. 118 undergraduates watched one of two videotaped interviews with the same instructor: warm and friendly in one, cold and distant in the other. His appearance, mannerisms, and accent were physically identical in both, yet students who saw the warm version rated those fixed attributes as appealing, and students who saw the cold version rated the same attributes as irritating, and denied their attribute ratings had been influenced by likability at all.
See also: Social Proof and Framing Effect, related but distinct: those describe trust borrowed from other people's behaviour or from how a fact is worded; this is trust borrowed from one unrelated trait of the same source, no wording or crowd required.
Tested as an experiment: Does a great digital home loan lift home loan trust everywhere else, too?, testing whether the halo from one excellent channel spills into unrelated, unchanged channels of the same bank.
54

Loss Aversion

Why does losing $20 hurt more than finding $20 feels good?

Losses are felt more intensely than equivalent gains, so people go out of their way to avoid a loss they'd happily walk past as a foregone gain.

THE JOB LABELLED "CURRENT" IS THE ONE THEY KEEP QUIETSOCIAL "CURRENT": THE QUIET JOB KEEPS THE QUIET JOB QUIETSOCIAL "CURRENT": THE SOCIAL JOB KEEPS THE SOCIAL JOB The same two jobs, offered either way. Whichever matched their labelled routine is the one each kept.
Likely mechanismLosing something hurts noticeably more than gaining the same thing feels good

The psychology. Loss aversion is the core asymmetry inside prospect theory: the psychological pain of losing something is felt more intensely than the pleasure of gaining the identical thing, so choices are shaped more by what might be lost relative to a reference point than by the objective size of the gain or loss on offer. Change the reference point without changing the actual options, and preferences can flip.

The full write-up: study, numbers, and caveats
The psychology

Loss aversion is the core asymmetry inside prospect theory: the psychological pain of losing something is felt more intensely than the pleasure of gaining the identical thing, so choices are shaped more by what might be lost relative to a reference point than by the objective size of the gain or loss on offer. Change the reference point without changing the actual options, and preferences can flip.

Where it causes errors

Framing a switch as giving something up (losing a current low rate, losing a seat, losing an accrued discount) suppresses switching even when the alternative is objectively better, because leaving gets coded as a loss of the current plan's benefits rather than a gain of the new plan's. Investors hold falling stocks too long for the same reason: selling would "lock in" a loss that, on paper, has already happened.

Where it can help

Framing a change that's genuinely already in someone's interest as protecting them from a future loss ("keep your rate from rising," "don't lose your streak") can honestly help people follow through on a decision they already wanted to make. This only holds up when the loss being described is real; borrowing the frame to manufacture urgency around a change that isn't actually protective is a different, dishonest use of the same lever.

Tversky, A., & Kahneman, D. (1991). “Loss Aversion in Riskless Choice: A Reference-Dependent Model.” Quarterly Journal of Economics, 106(4), 1039–1061
The paper that extended loss aversion beyond risky gambles into ordinary riskless choices: no probabilities involved, just two options compared against different starting points.
StrengthA between-subjects design manipulated only which option was described as matching a person's current position, while holding the two options being compared completely fixed, isolating the effect of the reference point itself from the effect of the options.
WeaknessLike most of Tversky and Kahneman's individual-choice studies, this relies on a hypothetical described scenario rather than a real decision with something real being given up, and the job-choice experiment's headline 70%-versus-33% split comes from a single sample of 106 people.
Key findings
ParaphrasedParticipants imagined leaving a training job for one of two new jobs: Job X, with limited social contact and a 20-minute commute, or Job Y, moderately sociable with a 60-minute commute. The two jobs never changed. What changed was a separate, third reference job used for comparison: an isolated one with a 10-minute commute for one group, a highly sociable one with an 80-minute commute for the other, so the identical trade-off between X and Y was framed as a loss on a different dimension depending only on which reference job the comparison started from.
Verbatim“Job x was chosen by 70 percent of the participants in version 1 and by only 33 percent of the participants in version 2 (N = 106, p < 0.01).”
Also worth citing: Tversky, A., & Kahneman, D. (1992). “Advances in Prospect Theory: Cumulative Representation of Uncertainty.” Journal of Risk and Uncertainty, 5(4), 297–323, a separate paper, using risky gamble-choice experiments with 25 graduate students, that produced the widely-cited loss-aversion coefficient of roughly 2 to 2.5 (losses looming about twice as large as equivalent gains). That figure comes from this different, risk-based paradigm, not the riskless job-choice study above. The two papers are often blurred together, but they're testing the same underlying asymmetry with different methods.
See also: Endowment Effect, a specific case of this same asymmetry: once something is treated as already owned, giving it up gets coded as a loss, not a foregone gain.
See also: Clawback, a deliberately engineered version of the same asymmetry: a reward is explicitly framed as already given so that failing to keep it registers as a loss rather than a missed gain.
See also: Risk Aversion, a related but distinct asymmetry: this principle is about losses hurting more than same-sized gains help, while risk aversion is about the curve of value itself, present even in a choice with no possible loss at all.
Decoded on The Science Behind: Why does GoalSaver's 4.75% bonus feel like money you'd be losing, not money you haven't earned yet?, where a savings account's bonus-interest condition uses this exact asymmetry: the same rate feels different depending on whether it's framed as protected or as still to be earned.
Also decoded on The Science Behind: How does Up Bank's behavioural design pay off for both the bank and its customers?, where a proactive warning frames a bonus rate as already the customer's to lose, sent before the condition actually lapses.
Also decoded on The Science Behind: Why do you have to be “into double denim” to win a better savings rate?, where skipping an optional app download gets framed as a loss instead of a neutral extra step.
Also worth reading: The Biases Draining Your Super Could Also Fill It, on real superannuation members who converted a temporary 2020 market downturn into a permanent loss by switching to cash at the bottom.
55

Signpost Effect

Why does the same fuel-economy number change what car you buy, depending only on how it's written?

Expressing the same underlying fact in different terms (cost, efficiency, or environmental impact) can activate a goal that was otherwise sitting dormant, and steer the choice towards whichever option best serves it.

THE SAME CAR IS CHOSEN MORE OFTEN LABELLED BY ITS GAS RATING 32 MPG LABELLED AS A PLAIN MPG FIGURE NO VALUE ACTIVATED 6/10 GAS RATING LABELLED BY A GREENHOUSE-GAS RATING ENVIRONMENTAL VALUE ACTIVATED Same car, same fuel economy, differently framed. Strongest for people who already valued the environment.
Likely mechanismRestating a fact in new terms can wake a dormant value and steer the choice towards it

The psychology. Most attributes can be truthfully expressed multiple ways: a car's fuel economy as miles-per-gallon, dollars of fuel cost, or a comparative greenhouse-gas rating. The specific translation used isn't neutral: it can act as a "signpost," activating whichever goal that framing happens to be congruent with. A greenhouse-gas rating specifically wakes up pro-environmental values in a way a plain MPG figure doesn't, and once that goal is active, it directs attention and choice towards the option that best satisfies it.

The full write-up: study, numbers, and caveats
The psychology

Most attributes can be truthfully expressed multiple ways: a car's fuel economy as miles-per-gallon, dollars of fuel cost, or a comparative greenhouse-gas rating. The specific translation used isn't neutral: it can act as a "signpost," activating whichever goal that framing happens to be congruent with. A greenhouse-gas rating specifically wakes up pro-environmental values in a way a plain MPG figure doesn't, and once that goal is active, it directs attention and choice towards the option that best satisfies it. The effect is strongest for people who already hold the relevant value; the same underlying facts, in a values-neutral translation, don't reliably wake it up on their own.

Where it causes errors

A retailer or car dealer choosing which framing to lead with can quietly tilt customers towards, or away from, the efficient option without changing a single underlying fact. Someone who would have bought the fuel-efficient car if shown its greenhouse-gas rating might buy the inefficient one if shown only its miles-per-gallon figure, not because their values changed, but because nothing woke them up.

Where it can help

Product and policy design that deliberately choose the value-congruent translation, a real per-year dollar cost instead of an abstract MPG figure, a real tons-of-CO2 number instead of an unlabelled score, genuinely helps people who hold pro-environmental values act on values they already have, rather than leaving them dormant. This is disclosure done well: the same real number, presented in the terms that actually let someone apply their own priorities to it.

Ungemach, C., Camilleri, A. R., Johnson, E. J., Larrick, R. P., & Weber, E. U. (2018). “Translated Attributes as Choice Architecture: Aligning Objectives and Choices Through Decision Signposts.” Management Science, 64(5), 2445–2459
Three car-choice experiments testing whether re-expressing the same fuel-economy attribute in different terms changes which car gets chosen, and why.
StrengthTested three different translations of the identical underlying fuel-economy attribute, annual fuel cost, gallons-per-100-miles, and a comparative greenhouse-gas rating, against each other across three controlled choice experiments, isolating the effect of the framing itself from the effect of the number.
WeaknessAll three experiments (341, 799, and 606 participants respectively) recruited paid respondents through Amazon Mechanical Turk, a convenience sample making a hypothetical choice, not real car buyers with their own money on the line.
Key findings
ParaphrasedEvery attribute can be truthfully expressed multiple ways, for a car's fuel economy, as miles per gallon, dollars of annual fuel cost, or a comparative greenhouse-gas rating, and which translation is used is not neutral: it can activate an otherwise-dormant goal congruent with that specific framing.
ParaphrasedAcross three experiments, participants shown fuel economy expressed in a way congruent with pro-environmental values were more likely to choose the more fuel-efficient car than participants shown the identical underlying fuel economy expressed in a values-neutral way, an effect that was stronger among participants who reported valuing the environment more.
Also worth citing: Larrick, R. P., & Soll, J. B. (2008). “The MPG Illusion.” Science, 320(5883), 1593–1594, a related paper by an overlapping research group, showing miles-per-gallon itself is a nonlinear, easily-misjudged unit; that separate finding helped motivate the US EPA's 2013 addition of a gallons-per-100-miles figure to new car stickers. It's a different specific mechanism than the signpost effect above (a measurement-scale illusion, not a dormant-value activation), but the same broader lesson: how a fuel-economy number is expressed changes real purchase decisions and real policy, not just lab ones.
56

Perceived Fairness (Dual Entitlement)

Why is raising a snow shovel's price the morning after a blizzard "unfair," but raising it because costs went up is fine?

People judge a price change against a reference transaction, not in isolation. Protecting profit from a real cost increase reads as fair; exploiting a shift in demand or bargaining power reads as exploitative, even at the identical new price.

A SNOW SHOVEL PRICED UP THE MORNING AFTER A STORM READS UNFAIR $15 SAME SHOVEL, BEFORE THE STORM NORMAL PRICE $20 SAME SHOVEL, MORNING AFTER 82% CALLED IT UNFAIR Same store, same shovel. The store’s own costs never changed.
Likely mechanismA price rise feels fair when it tracks a real cost increase, not just higher demand

The psychology. People carry an implicit reference transaction: the price, wage, or rent that has applied to them so far, and treat it as an entitlement both sides are assumed to honour. A firm is entitled to protect its existing reference profit, so passing on a real cost increase is fair; it isn't entitled to exploit a shift in its bargaining power to raise price beyond what's needed to protect that same profit. The same number, a price going up, reads as fair or unfair depending on which of these two stories it appears to belong to, not on the number itself.

The full write-up: study, numbers, and caveats
The psychology

People carry an implicit reference transaction: the price, wage, or rent that has applied to them so far, and treat it as an entitlement both sides are assumed to honour. A firm is entitled to protect its existing reference profit, so passing on a real cost increase is fair; it isn't entitled to exploit a shift in its bargaining power to raise price beyond what's needed to protect that same profit. The same number, a price going up, reads as fair or unfair depending on which of these two stories it appears to belong to, not on the number itself.

Where it causes errors

A company that raises prices citing "increased demand" alone, with no visible cost story, reliably reads as exploitative even to customers who would accept the identical increase framed as covering rising costs, meaning a factually accurate explanation can land worse than a less complete one, purely because of which reference transaction it invokes.

Where it can help

Genuinely disclosing the real cost driver behind a price change, a bank citing a real rise in its own funding costs, a delivery platform citing a real rise in fuel costs, lets customers judge the change against the fair reference point instead of assuming the worst, without requiring anything misleading. This only works, and should only be used, when the cited cost increase is real.

Kahneman, D., Knetsch, J. L., & Thaler, R. (1986). “Fairness as a Constraint on Profit Seeking: Entitlements in the Market.” American Economic Review, 76(4), 728–741
A telephone survey of Toronto and Vancouver residents conducted between May 1984 and July 1985, asking people to judge realistic pricing, wage, and rent vignettes as fair or unfair.
StrengthUsed many parallel vignettes that varied only the story behind an identical price or wage change, a real cost increase versus a demand spike, an existing customer versus a new one, isolating exactly which detail of a price change drives the fairness judgment, not just whether people dislike price increases in general.
WeaknessFairness judgments were collected as hypothetical telephone-survey responses to described scenarios, not decisions in a real market with real money on either side. How strongly the same asymmetry shows up in real purchase or loyalty behaviour, rather than a stated opinion, is a separate question the design doesn't answer.
Key findings
ParaphrasedIn the paper's best-known vignette, a hardware store that had been selling snow shovels for $15 raised the price to $20 the morning after a large snowstorm; roughly 82% of respondents rated the increase as "Unfair" or "Very Unfair," even though nothing about the store's own costs had changed.
Verbatim“Parallel results were obtained in questions concerning residential tenancy. As in the case of wages, many respondents apply different rules to a new tenant and to a tenant renewing a lease. A rent increase that is judged fair for a new lease may be unfair for a renewal.”
57

Law of Round Numbers

Why would someone retake an entire test just to turn a 990 into a 1,000?

A round number acts as an unusually strong reference point: people quietly change their effort, timing, or choices to land on one, or to avoid stopping one short of it, even when the round number carries no different real-world consequence than the number right next to it.

SAME DISTANCE FROM 1000 ONLY ONE SIDE GETS RETAKEN 990 1010 1000 A ROUND NUMBER, RIGHT BETWEEN THEM A SCORE OF 990 GETS THE WHOLE TEST RETAKEN TO FIX IT A SCORE OF 1010 IS SIMPLY KEPT, NO RETEST NEEDED Both scores are ten points from round. Only the lower one feels unfinished.
Likely mechanismA number ending in zero reads as a real target, worth extra effort in the final stretch to reach, or to avoid stopping one short of

The psychology. A score, an average, or a balance sitting one small step below a round number, a 990 instead of a 1,000, a $97 bill instead of $100, gets treated as unfinished business, while the exact same size of gap sitting just above a round number goes unnoticed. The round number itself becomes the thing a person judges themselves against, states out loud, or remembers afterwards, so crossing it feels like a distinct achievement and stopping one short of it feels like a distinct failure, even though the real underlying difference is trivial. People respond by quietly adjusting effort, timing, or behaviour specifically to land on the round number, or to avoid falling just short of it. Professional baseball offers an unusually clean, well-documented example of the same pattern: a batting average of .299 and .300 differ by a single hit across an entire season, functionally nothing, but .300 is the number a player's whole season gets remembered by.

The full write-up: study, numbers, and caveats
The psychology

A score, an average, or a balance sitting one small step below a round number, a 990 instead of a 1,000, a $97 bill instead of $100, gets treated as unfinished business, while the exact same size of gap sitting just above a round number goes unnoticed. The round number itself becomes the thing a person judges themselves against, states out loud, or remembers afterwards, so crossing it feels like a distinct achievement and stopping one short of it feels like a distinct failure, even though the real underlying difference is trivial. People respond by quietly adjusting effort, timing, or behaviour specifically to land on the round number, or to avoid falling just short of it.

Professional baseball offers an unusually clean, well-documented example of the same pattern: a batting average of .299 and .300 differ by a single hit across an entire season, functionally nothing, but .300 is the number a player's whole season gets remembered by.

Where it causes errors

A student who scores 990 on the SAT is measurably more likely to retake the entire exam than a student who scored 1000, even though the two scores are close to functionally identical to any admissions office reading the application. The round number, not the actual competitive difference between the two scores, is what triggers months of extra prep and a second test fee for a gain that mostly isn't there.

Where it can help

The same pull can be used honestly. A savings app that frames a balance as “on the way to $1,000” rather than just showing the raw number gives someone a round milestone worth a final push to reach, turning an otherwise arbitrary in-between balance into a target actually worth the extra effort, the identical mechanism, aimed at a goal the person already wanted.

Pope, D., & Simonsohn, U. (2011). “Round Numbers as Goals: Evidence From Baseball, SAT Takers, and the Lab.” Psychological Science, 22(1), 71–79
Real archival data from two unrelated domains, plus a controlled lab study confirming the mechanism.
StrengthThe same round-number discontinuity shows up independently in decades of professional baseball at-bat records and in national SAT retesting records, two populations with completely different stakes and incentives, and the authors then reproduce the mechanism directly in a controlled lab task, combining real-world external validity with lab-level causal evidence in one paper.
WeaknessThe baseball and SAT findings are both naturally occurring, archival data, not a randomised field experiment, so on their own they can't fully rule out every alternative structural explanation for the pattern near the threshold. That gap is exactly what the lab study is designed to close, but it does mean the field evidence alone is correlational.
Key findings
Verbatim (abstract)“Professional baseball players modify their behavior as the season is about to end, seeking to finish with a batting average just above rather than below .300.”
Verbatim“High school juniors were at least 10 to 20 percentage points more likely to retake the SAT if their total score ended in 90 (e.g., 1190) than if it ended in the most proximate 00 (e.g., 1200).”
ParaphrasedIn a controlled lab task, participants told their performance was just short of a round-number benchmark reported a stronger desire to keep working than participants told they were the same distance from a non-round benchmark, directly supporting round numbers acting as goals rather than the field patterns having some other explanation.
Also worth citing: Allen, E. J., Dechow, P. M., Pope, D. G., & Wu, G. (2017). “Reference-Dependent Preferences: Evidence From Marathon Runners.” Management Science, 63(6), 1657–1672. Across more than 9 million recorded marathon finishes, times bunch just under round-hour marks like 4:00:00, the same pull, showing up in pacing instead of hits or test scores.
See also: Salience, the round number is what makes that particular point on the scale stand out enough to become a reference point in the first place, and Goal Gradient Effect, the same “effort intensifies near the target” mechanism, here aimed at a round number instead of a fixed external goal.
58

Payment Transparency

Why does cash counted out by hand stay with you longer than a card tap for the same amount?

A payment's felt cost depends less on the price than on whether paying it makes you rehearse the amount and whether the money leaves you immediately. Skip both, and even a real payment barely registers as spending.

COUNTING OUT CASH GETS REMEMBERED TAPPING A CARD OFTEN DOESN'T CASH COUNTED OUT BY HAND LATER SPENDING FEELS HELD BACK CARD TAPPED, BILL COMES LATER LATER SPENDING GOES UNCHECKED Same price, paid either way. Only how well it gets remembered differs.
Likely mechanismA payment that's counted out and taken immediately is remembered; one that isn't, isn't

The psychology. Not every payment method leaves the same mark on memory. Soman's framework separates two properties most people lump together as “convenience”: rehearsal, whether paying requires you to actively write down or restate the amount, and immediacy, whether the money leaves your account the moment you pay or only later, when a bill arrives. Both properties independently affect how well a payment gets remembered, and a well-remembered payment is what normally makes the next purchase feel a little more constrained. Skip the rehearsal, skip the immediacy, and a real payment can pass almost unnoticed.

The full write-up: study, numbers, and caveats
The psychology

Not every payment method leaves the same mark on memory. Soman's framework separates two properties most people lump together as “convenience”: rehearsal, whether paying requires you to actively write down or restate the amount, and immediacy, whether the money leaves your account the moment you pay or only later, when a bill arrives. Both properties independently affect how well a payment gets remembered, and a well-remembered payment is what normally makes the next purchase feel a little more constrained. Skip the rehearsal, skip the immediacy, and a real payment can pass almost unnoticed.

Where it causes errors

A shopper who taps a stored credit card, with no amount to write and no balance falling in front of them, is measurably less likely to feel constrained by that purchase when a further one comes up minutes later, even though the money is just as gone. A checkout that defaults customers to a saved one-tap card, removing manual entry and any running total on screen, removes both levers that would otherwise remind a shopper what they've already spent.

Where it can help

The same two levers work in reverse as a genuine budgeting tool. Some banking apps deliberately reintroduce what a card payment would otherwise skip: a running balance shown at the moment of tap, or a weekly total broken down by category. It puts back the cues a card removed, so people can actually feel what they're spending rather than being pushed to spend less than they'd choose to.

Soman, D. (2001). “Effects of Payment Mechanism on Spending Behavior: The Role of Rehearsal and Immediacy of Payments.” Journal of Consumer Research, 27(4), 460–474
The first experiment compares credit-card to check payment directly; the second independently manipulates rehearsal and immediacy across four payment mechanisms to isolate which lever drives the effect.
StrengthThe second experiment varied rehearsal and immediacy independently, across four payment mechanisms in a within-subject design, rather than simply comparing cash to card as one bundled condition. That is what lets the paper claim they're two distinct, separable mechanisms rather than one vague “cards feel different” effect.
WeaknessLike much of this literature, the experiments measure stated purchase intention in constructed scenarios rather than tracking real, incentivised spending in the field, so the size of the effect on genuine spending, as opposed to reported willingness to buy, is less certain than a real-money field experiment would show.
Key findings
Verbatim (abstract)“Specifically, past payments strongly reduce purchase intention when the payment mechanism requires the consumer to write down the amount paid (rehearsal) and when the consumer’s wealth is depleted immediately rather than with a delay (immediacy).”
ParaphrasedCredit cards score low on both rehearsal and immediacy, since paying rarely requires writing the amount and the money isn't actually withdrawn until the bill is settled later. The paper offers this as a mechanism behind commonly observed credit-card overspending, distinct from credit limits or interest rates.
See also: Pain of Paying, the felt discomfort this principle's missing rehearsal and immediacy would otherwise trigger, and Mental Accounting, how a payment's mental category shapes whether it gets tracked at all. Also see The Credit Card Premium, where skipping both levers at once shows up as a real, measured jump in what people will pay.
59

The Credit Card Premium

Why did the same tickets fetch bids nearly twice as high the moment bidders were told they'd pay by credit card?

People are reliably willing to pay more for the identical item once they're told they'll pay by credit card instead of cash, a gap that shows up even in a real auction where the money at stake is genuine.

TOLD TO PAY WITH CREDIT, BIDDERS BID MORE THAN DOUBLE TICKET TOLD: WINNER PAYS BY CASH BID STAYS AT BASELINE TICKET TOLD: WINNER PAYS BY CREDIT BID MORE THAN DOUBLES Same tickets, same real auction, either way. Only the promised payment method changed the bid.
Likely mechanismBeing told you'll pay later, by card, blunts the felt cost of the price right now

The psychology. Credit takes Payment Transparency's two levers, rehearsal and immediacy, and removes both further than a debit card does: nothing is written down, and unlike debit, no money leaves the account at all until the statement is settled, sometimes weeks later. Prelec and Simester isolated exactly this in a real auction: bidders told only which payment method they'd use if they won bid substantially more under credit instructions than under cash instructions, for the literal same tickets. Because bidding higher had a real cost, and the researchers checked for and largely ruled out simple cash-on-hand constraints as the explanation, the size of the gap points to the felt cost of paying being blunted by the deferral itself, not just liquidity.

The full write-up: study, numbers, and caveats
The psychology

Credit takes Payment Transparency's two levers, rehearsal and immediacy, and removes both further than a debit card does: nothing is written down, and unlike debit, no money leaves the account at all until the statement is settled, sometimes weeks later. Prelec and Simester isolated exactly this in a real auction: bidders told only which payment method they'd use if they won bid substantially more under credit instructions than under cash instructions, for the literal same tickets. Because bidding higher had a real cost, and the researchers checked for and largely ruled out simple cash-on-hand constraints as the explanation, the size of the gap points to the felt cost of paying being blunted by the deferral itself, not just liquidity.

Where it causes errors

Any high-consideration purchase where a seller defaults customers towards paying by card, an auction, a big-ticket retail sale, an in-app purchase against a stored card, quietly inflates what people are willing to bid or pay beyond what they'd offer under cash terms, without anyone making an active, conscious decision to pay more.

Where it can help

This is mostly a consumer vulnerability rather than something a business can use honestly, but the same mechanism works protectively when someone applies it to themselves. Deliberately switching a specific temptation purchase, an auction, a big discretionary buy, to debit or cash reintroduces the immediacy that would otherwise be missing: a self-imposed commitment device, the same spirit as Mental Accounting's earmarked savings account and Precommitment Devices more generally.

Prelec, D., & Simester, D. (2001). “Always Leave Home Without It: A Further Investigation of the Credit-Card Effect on Willingness to Pay.” Marketing Letters, 12(1), 5–12
A real silent auction for real Boston Celtics tickets, with a genuine winner who had to actually pay.
Experiment Teardown
MBA students bidding in one real silent auction
Only which payment method they were told to use if they won differed
Told: pay cash
Bids for real Celtics tickets, cash instructed
Told: pay credit
Bids for the identical tickets, credit instructed
Credit-instructed bids averaged more than double the cash-instructed bids ($60.64 versus $28.51), for the same tickets in the same room.
StrengthA real auction for a real, desirable item, with a genuine winner who had to actually pay, rather than a hypothetical willingness-to-pay survey, so bids reflect a real financial commitment rather than a stated intention.
WeaknessThe sample was MBA students in one classroom auction for one category of item, event tickets, an unusually numerate population bidding in a single event. The paper's own second study, using a restaurant gift certificate of known value, found no significant premium at all in its cleanest replication condition, so the roughly two-fold Celtics-ticket magnitude may not generalise to every purchase type or population without further work.
Key findings
ParaphrasedBidders instructed to pay by credit card if they won bid an average of $60.64 for the Celtics tickets, more than double the $28.51 average among bidders instructed to pay cash, a 113% premium for the identical tickets in the same auction.
ParaphrasedThe researchers checked whether the gap could simply reflect cash-strapped bidders being unable to bid as high in the cash condition, and found the effect held up too broadly across the sample for a liquidity constraint alone to explain it.
Also worth citing: Runnemark, E., Hedman, J., & Xiao, X. (2015). “Do Consumers Pay More Using Debit Cards Than Cash? An Experiment.” Electronic Commerce Research and Applications, 14(5), 285–291. A separate, real-money experiment found a smaller but real gap between debit card and cash, suggesting the effect scales with how far a payment method defers and de-emphasises the cost, not just whether it's a card at all.
See also: Payment Transparency, the general mechanism this is the sharpest real-world case of, and Pain of Paying, the felt discomfort credit is unusually good at removing.
60

Pseudo-Set Framing

Why does grouping five loose nickels into one picture of a quarter double how often people bet them all?

People push to complete any group framed as a whole, a quarter's worth of nickels, a six-item gift, a four-step checklist, even when the grouping is arbitrary, costs money to finish, and carries no reward for finishing at all.

THE SAME BETS, FRAMED AS A SET, GOT TAKEN NEARLY TWICE AS OFTEN 5 NICKELS SHOWN AS SEPARATE COINS 16% ACCEPTED EVERY BET OFFERED 1 QUARTER SAME NICKELS, PICTURED AS ONE QUARTER 29% ACCEPTED EVERY BET OFFERED The bets and the odds never changed. Only the picture of one complete quarter did.
Likely mechanismAn unfinished "set", even one drawn up on the spot, feels incomplete in a way one leftover item doesn't, and finishing it removes that feeling

The psychology. Gestalt psychology holds that people perceive groups, not just individual items, and register a group as more or less complete based on how it looks, not just how much objectively remains. Barasz, John, Keenan, and Norton showed that simply drawing an arbitrary line around a handful of items, five nickels pictured as wedges of one quarter, a handful of tasks pictured as one checklist, makes the group read as a single thing with a completion state of its own. Leaving that pseudo-set unfinished feels like a real loss, distinct from whatever the individual items were worth on their own, and the pull to close it out survives even when people are told outright that the grouping is made up.

The full write-up: study, numbers, and caveats
The psychology

Gestalt psychology holds that people perceive groups, not just individual items, and register a group as more or less complete based on how it looks, not just how much objectively remains. Barasz, John, Keenan, and Norton showed that simply drawing an arbitrary line around a handful of items, five nickels pictured as wedges of one quarter, a handful of tasks pictured as one checklist, makes the group read as a single thing with a completion state of its own. Leaving that pseudo-set unfinished feels like a real loss, distinct from whatever the individual items were worth on their own, and the pull to close it out survives even when people are told outright that the grouping is made up.

Where it causes errors

The pull to finish a pseudo-set doesn't depend on the reward being worth it. A slot machine's "collect all 6 symbols" bonus, a shelf tag offering a discount at "any 4," or a subscription's onboarding checklist framed as one badge, all group items into an arbitrary set. That grouping alone can get someone to keep spending, gambling, or clicking well past the point a cold cost-benefit calculation would have stopped them. The set itself becomes the target, not the outcome it was meant to serve.

Where it can help

The identical mechanism, aimed at something someone already wants to finish, a course, a repayment plan, an account setup, turns a long undifferentiated task into a visible, closeable set of smaller ones. A four-part account setup framed as one thing to complete, rather than four separate steps with no shared frame, gives people a real, honest completion state to chase instead of an arbitrary one, with nothing fabricated about what's actually left to do.

Barasz, K., John, L. K., Keenan, E. A., & Norton, M. I. (2017). “Pseudo-Set Framing.” Journal of Experimental Psychology: General, 146(10), 1460–1477
Five studies plus real field data, spanning gambling, effort, giving, and purchase decisions, all built around the same arbitrary "set".
Experiment Teardown
Participants offered a series of small bets
Only how the same nickels were pictured differed
Separate nickels
5 nickels shown as separate, unconnected coins
Chose whether to accept each bet
Pictured as a quarter
Same 5 nickels, pictured as wedges of one quarter
Chose whether to accept each bet
Result Separate Quarter Swing
Accepted every bet offered 16% 29% +13

The bets and the odds never changed. Only whether the nickels looked like one complete quarter did, and that alone nearly doubled who gambled all the way through.

Internal validityRandom assignment to framing condition, with the underlying bets and expected value held identical across both groups.
External validityA small-stakes lab gamble with nickels; the paper's other four studies, on effort, giving, and purchasing, are what show the pull to complete a set isn't limited to gambling.
Experiment Teardown
Real donors on the Canadian Red Cross's 2016 online Holiday Campaign (N = 7,117)
Randomly shown one of three landing pages: cash-emphasis, gift-emphasis, or a six-item set
Cash-emphasis page
Cash option most prominent, gift items optional
Pseudo-set page
Encouraged to complete a six-item “Global Survival Kit”

The set, not the amount, became the target. Among donors who chose to give gifts, 21% shown the six-item “Global Survival Kit” chose all six items, compared with 3% on the cash-emphasis page and 5% on the ordinary gift page.

Internal validityA real field experiment with 7,117 donors randomly assigned across three live landing pages, not a lab simulation.
External validityCharitable giving carries its own norms (warm-glow, social signalling) separate from gambling or shopping, which is why the paper's other four lab studies, outside giving, are what confirm the mechanism itself.
StrengthFive separate studies plus real field donation data, spanning gambling, effort, giving, and purchasing, all point to the same mechanism, and it held up in conditions built specifically to rule out simpler explanations: no reward for completing the set, a real cost to finish it, and telling participants outright that the grouping was arbitrary.
WeaknessThe cleanest, most quantified comparison (Study 1) is still a small-stakes lab gamble with nickels. The paper leans on four further studies and real field data to show the pull to complete a set generalises beyond a coin-flip game, rather than one single large field trial reporting a hard revenue number.
Key findings
Verbatim“Significantly more participants in the pseudo-set condition accepted all four gambles (29%) than in the control condition (16%).”
Verbatim (abstract)“These effects persist in the absence of any reward, when a cost must be incurred, and after participants are explicitly informed of the arbitrariness of the set.”
See also: Goal Gradient Effect, a related but distinct completion pull: that one is about accelerating as the real distance to a goal shrinks, this one is about a group feeling incomplete regardless of distance, based only on how it's framed.
61

Smart Defaults (Save More Tomorrow)

Why did people who wouldn't take an immediate pay cut agree to save more from a raise they hadn't got yet?

A default set for someone's own future benefit, timed so that accepting it costs nothing anyone can actually feel today.

SAVING MORE FROM A RAISE NOBODY HAD RECEIVED YET BEFORE JOINED THE PLAN 3.5% OF PAY SAVED 4 RAISES LATER ABOUT 40 MONTHS ON THE PLAN 13.6% OF PAY SAVED Take-home pay still rose at every raise. Only how much of each raise went to savings changed.
Likely mechanismCommitting today to a raise you haven't got yet means the increase never has to feel like a loss

The psychology. Three separate biases usually work against saving more: present bias makes a felt pay cut today outweigh a benefit felt decades from now, loss aversion makes that felt cut sting harder than an equivalent future gain would please, and once a savings rate is set, inertia keeps it there, for better or worse. Save More Tomorrow routes around all three at once.

The full write-up: study, numbers, and caveats
The psychology

Three separate biases usually work against saving more: present bias makes a felt pay cut today outweigh a benefit felt decades from now, loss aversion makes that felt cut sting harder than an equivalent future gain would please, and once a savings rate is set, inertia keeps it there, for better or worse.

Save More Tomorrow routes around all three at once. The commitment happens now, but the increase doesn't take effect until a future, already-scheduled pay raise, so nothing about today's paycheck moves and present bias has nothing to resist. The raise still arrives; only part of it is redirected, so take-home pay keeps rising and there's no loss to feel. And once someone's enrolled, the same inertia that normally protects a low savings rate now protects a rising one instead.

Where it causes errors

The design only avoids the pay cut it's built to avoid if the escalation is genuinely tied to a real raise. Plenty of auto-escalation programmes raise the contribution rate on a fixed calendar date instead, whether or not a raise shows up that year.

During a pay freeze, or for anyone whose raise doesn't keep pace with inflation, the automatic increase becomes exactly the felt pay cut the original design was built to sidestep, and the same inertia that once protected a rising savings rate now quietly protects a shrinking paycheck nobody chose to shrink. Retirement-industry guidance on 401(k) auto-escalation flags this directly: tie the increase to an actual raise, or the mechanism turns into the loss it was designed to avoid.

Where it can help

Used as designed, this is one of the most consequential nudges ever deployed. At the manufacturer where it was first tested, employees who joined raised their average savings rate from 3.5% of pay to 13.6% over four pay raises, about 40 months, without a single point of that increase ever registering as a pay cut. Nobody's choice was removed: anyone could opt out at any raise, and some did. What changed was the moment the decision got made, today, about a raise that hadn't happened yet, instead of at the one moment, the raise itself, when loss aversion is primed to fight it.

Thaler, R. H., & Benartzi, S. (2004). “Save More Tomorrow™: Using Behavioral Economics to Increase Employee Saving.” Journal of Political Economy, 112(S1), 164–187
A real field intervention inside one mid-size manufacturer's actual payroll and 401(k) system, tracked across four real pay raises.
StrengthA genuine field intervention, not a lab study or a stated-intention survey: real payroll deductions, tracked through four actual pay raises over roughly 40 months at one real employer.
WeaknessThe employees who joined SMarT had already been filtered twice, first by not yet saving the recommended rate, then by declining an immediate contribution increase, so the group studied may have been unusually receptive to a delayed, painless version of the same ask. How the same design performs on a workforce never offered that first, harder ask wasn't tested here.
Key findings
ParaphrasedOf 162 employees who had just declined an immediate increase to their contribution rate, 78% agreed to the SMarT plan instead: an increase timed to their next four pay raises, capped at 14% of pay.
ParaphrasedEmployees who joined raised their average savings rate from 3.5% to 13.6% of pay over the course of four raises, roughly 40 months, and 78% of those who joined were still enrolled by the fourth raise.
A note on sourcing: the published paper sat behind this site's network restrictions, so both findings above are paraphrased from the abstract and secondary academic summaries, not quoted verbatim.
See also: Default Effect, the passive version of the same lever. That principle is about the power of whatever's pre-selected; this one is about when and against what a default gets set, so it works with someone's biases instead of against them.
See also: Good Friction, which removes friction from the enrolment step itself, a single “yes” to join. Smart Defaults is a different lever: not how hard the decision is to make, but when it's made and what it's measured against.
See also: Every Default Decides Who Pays for Doing Nothing, a full report placing this timed default on a wider spectrum alongside forced choice, opt-in, opt-out, and personalised defaults.
Also worth reading: The Biases Draining Your Super Could Also Fill It, where Australia's legislated Superannuation Guarantee increases and MySuper's lifecycle investment defaults apply this exact logic to retirement savings.
62

Precommitment Devices (the Ulysses Contract)

Why did people, given a completely free choice of when a deadline should fall, still choose to bind themselves to an earlier date that cost them money if they missed it?

A voluntary constraint someone places on their own future choices today, so a future self with less willpower can't undo it.

SAME PROOFREADING TASK, THREE WAYS TO SET THE DEADLINE EVENLY-SPACED, SET BY THE RESEARCHERS DAY 7 DAY 14 DAY 21 Fewest missed deadlines, most errors caught SELF-IMPOSED, SET BY THE PARTICIPANT DAY 12 DAY 17 DAY 20 Better than none, but trailed the imposed schedule NO INTERIM DEADLINE, ALL DUE DAY 21 DAY 21 Most missed deadlines, fewest errors caught People chose to bind themselves to a deadline. They just didn't set it as well as an outside schedule would have.
Likely mechanismAnticipating a future lapse in willpower, people pay a real cost now to remove the tempting choice later

The psychology. People are time-inconsistent: the version of you making a plan today weighs the future differently than the version of you actually living in it will. Present bias makes near-term comfort loom larger than a cost or benefit sitting weeks away, so a plan that relies on willpower alone keeps losing to whichever version of you shows up on the day itself. A precommitment device works by moving the cost forward, locking in a real, felt consequence today so the future self inherits a constraint instead of a fresh decision.

The full write-up: study, numbers, and caveats
The psychology

People are time-inconsistent: the version of you making a plan today weighs the future differently than the version of you actually living in it will. Present bias makes near-term comfort, skipping the gym, putting off the assignment, loom larger than a cost or benefit sitting weeks away, so a plan that relies on willpower alone keeps losing to whichever version of you shows up on the day itself.

A precommitment device works by moving the cost forward: instead of trusting a future self to make the harder choice, it locks in a real, felt consequence today, a forfeited deposit, a public pledge, a deadline with a grade penalty attached, so the future self inherits a constraint instead of a decision.

Where it causes errors

A precommitment device only binds if the cost is real and the timing is right, and people are inconsistent judges of their own future selves on both counts. Commitment-contract services like StickK, where someone pledges money to a cause they'd hate funding if they miss a goal, only work when the stake and the deadline are set as strictly as an outside party would set them.

Left to design their own constraint, people tend to set stakes too low or deadlines too close to the finish line to actually bind, the same optimism about future self-control that caused the procrastination in the first place quietly weakens the device meant to fix it.

Where it can help

Used well, this is one of the few genuinely self-directed levers on this site: a savings account earmarked and named for a real goal, a subscription to a habit app with a real financial stake attached, or simply setting a costly, public deadline for a piece of work, all work by giving a future self less room to quietly slide. None of it requires deceiving anyone else, the person setting the constraint and the person bound by it are the same person, choosing in advance to make the easier path harder to take.

Ariely, D., & Wertenbroch, K. (2002). “Procrastination, Deadlines, and Performance: Self-Control by Precommitment.” Psychological Science, 13(3), 219–224
A real MIT Sloan executive-education class turning in graded term papers, paired with a controlled lab experiment paying real participants to proofread for real errors.
StrengthThe finding shows up twice, in a genuine field setting (a real class, real grades, a real 1%-per-day late penalty) and in a fully randomised lab experiment isolating the same three deadline structures under controlled, paid incentives, so it isn't resting on either a single uncontrolled field correlation or a lab-only artifact.
WeaknessHyndman and Bisin's 2026 direct replication of the controlled lab experiment, run with a different, real subject pool, found the deadline-structure manipulation had a negligible effect on the same performance measures, casting real doubt on how reliably that half of the original result generalises.
Key findings
Verbatim“The grades in the no-choice section (M = 88.76) were higher than the grades in the free-choice section (M = 85.67), t(97) = 3.03, p = .003.”
Verbatim (abstract)“People have self-control problems, they recognize them, and they try to control them by self-imposing costly deadlines. These deadlines help people control procrastination, but they are not as effective as some externally imposed deadlines in improving task performance.”
ParaphrasedA 2026 direct replication of the same lab design, run with a different subject pool, found the deadline-structure manipulation had a negligible effect on the three original performance measures, and evenly spaced externally imposed deadlines did not stand out as more effective at reducing procrastination.
A note on sourcing: Hyndman and Bisin's 2026 replication sat behind this site's network restrictions, so that finding above is paraphrased from its abstract and secondary academic summaries, not quoted verbatim. The original 2002 paper's findings above are now quoted directly from its full text.
See also: Mental Accounting, whose earmarked savings account is one concrete, everyday version of this same lever.
See also: The Credit Card Premium, where deliberately switching a temptation purchase to cash or debit is the same self-imposed constraint applied to a single spending decision.
See also: Present Bias, the underlying reason this device is needed at all: the same person genuinely chooses differently once a reward becomes immediate, which is exactly what a precommitment device is built to route around.
Also worth citing: Ashraf, N., Karlan, D., & Yin, W. (2006). “Tying Odysseus to the Mast: Evidence from a Commitment Savings Product in the Philippines.” Quarterly Journal of Economics, 121(2), 635–672, a randomised field trial of a real commitment savings account, restricting a client's own access to their money until a self-chosen goal or date. Average balances rose 81 percentage points more for the group offered the account than for a control group after twelve months, evidence this same lever works on money specifically, not only on deadlines.
Decoded on The Science Behind: Why does GoalSaver's 4.75% bonus feel like money you'd be losing, not money you haven't earned yet?, where a bank's bonus-interest condition does the same job as a formal commitment account, without the paperwork.
Also decoded on The Science Behind: How does Up Bank's behavioural design pay off for both the bank and its customers?, where Locked Savers attaches a real 3-hour delay, plus an optional social cost, to a future withdrawal.
63

Crossmodal Correspondence (the coffee mug effect)

Why did the identical café latte taste sweeter poured into a blue mug than a white one?

The brain treats taste as one blended sense, not five separate ones, so an unrelated touch or colour cue gets folded straight into how something is judged to taste, even though nothing about the substance itself changed.

THE SAME LATTE TASTES DIFFERENT POURED INTO A DIFFERENT COLOURED MUG POURED INTO A WHITE MUG RATED LESS SWEET, MORE INTENSE Same coffee, same temperature POURED INTO A BLUE MUG RATED SWEETER, LESS INTENSE The identical latte, poured seconds apart Nothing in the cup changed. Only the colour holding it did.
Likely mechanismA touch or colour cue gets folded straight into the taste itself, not judged as a separate, later step

The psychology. Flavour isn't assembled from taste alone: the brain integrates taste with sight, touch, sound, and smell into a single experience well before any conscious judgment happens, so a cue from an entirely different sense can shift what a person reports tasting without them ever noticing a separate influence. This is a different kind of error to a belief quietly colouring a belief; it's one sense quietly colouring another.

The full write-up: study, numbers, and caveats
The psychology

Flavour isn't assembled from taste alone: the brain integrates taste with sight, touch, sound, and smell into a single experience well before any conscious judgment happens, so a cue from an entirely different sense can shift what a person reports tasting without them ever noticing a separate influence. This is a different kind of error to a belief quietly colouring a belief; it's one sense quietly colouring another, which is exactly why it's so hard to correct for: nobody can taste more carefully to route around a cue that's already been folded into the taste itself.

Where it causes errors

A product can be marked down in perceived quality for a reason that has nothing to do with what it actually is: a coffee served in a flimsy-feeling cup, a snack in packaging whose colour visually clashes with the food itself, can get rated as lower quality than a chemically identical version served or packaged differently. The substance never changed; only a cue sitting in a completely different sense did.

Where it can help

This is one of the more genuinely benign principles on this list, because the underlying substance is real either way. A cup or package designed to accentuate a flavour that's actually present, drawing attention to a coffee's own real sweetness rather than inventing one, is amplifying something true. The line it can't cross without becoming manipulation is using the same cue to disguise an actual defect: a genuinely stale or bitter product hidden behind cues that suggest freshness or sweetness it doesn't have.

Van Doorn, G. H., Wuillemin, D., & Spence, C. (2014). “Does the Colour of the Mug Influence the Taste of the Coffee?” Flavour, 3, Article 10
A real café latte, poured from the same jug, served to the same tasters in a white mug, a blue mug, and a clear glass.
StrengthUsed a real beverage tasted directly, not a described or imagined scenario, with every cup made under identical machine settings so mug colour was the one relevant difference between groups.
WeaknessEach experiment was between-subjects, so no single taster compared mugs directly: Experiment 1 split 18 volunteers six to a mug, and Experiment 2 split 36 volunteers twelve to a mug, leaving each condition's sample small despite the replication.
Key findings
Verbatim (abstract)“The coffee was rated as less sweet in the white mug as compared to the transparent and blue mugs.”
Verbatim (abstract)“In experiment 1, the white mug enhanced the rated ‘intensity’ of the coffee flavour relative to the transparent mug.”
ParaphrasedThe authors proposed colour contrast against the coffee's own brown colour as the likely driver for the white mug's effect. They explicitly ruled out the opposite idea, that blue would soften the contrast and taste sweeter because it's brown's complementary colour, since ratings for the blue mug and the clear glass never differed significantly in either experiment.
Also worth citing: Ichimura, F., Motoki, K., Matsushita, K., & Ariga, A. (2023), “The Tactile Thickness of the Lip and Weight of a Glass Can Modulate Sensory Perception of Tea Beverage,” Food and Humanity, 1, 180–187, found the same kind of shift through touch instead of sight: identical tea tasted sweeter poured from a glass with a thick, heavy lip than from one with a thin, light lip.
See also: Halo Effect, related but distinct: that's one already-known trait colouring belief about a different, unrelated trait of the same source; this is a cue from a different sense changing the direct sensory experience itself, with no belief or inference in between.
64

Meaningless Differentiation

Why did an identical jacket get rated as better quality once its down filling carried a made-up “Alpine Class” label?

A visible, specific-sounding detail can be read as proof of a real, unstated benefit, even when the detail itself changes nothing: buyers assume a company wouldn't have added it for no reason.

THE SAME JACKET RATED HIGHER FOR A LABEL THAT MEANS NOTHING PLAIN DOWN FILLING RATED AS ORDINARY QUALITY IDENTICAL FILLING, LABELLED “ALPINE CLASS” RATED AS SUPERIOR QUALITY “Alpine Class” wasn't attached to any real, verifiable standard.
Likely mechanismIf a company bothered to add a feature, buyers assume it must mean something, even when it plainly doesn't

The psychology. When two products are otherwise hard to tell apart, a specific-sounding, hard-to-verify attribute gets treated as evidence in its own right: buyers reason that a company wouldn't bother adding a named, particular-looking detail unless it stood for something real, and substitute that inference for the harder work of actually checking whether it does.

The full write-up: study, numbers, and caveats
The psychology

When two products are otherwise hard to tell apart, a specific-sounding, hard-to-verify attribute gets treated as evidence in its own right: buyers reason that a company wouldn't bother adding a named, particular-looking detail unless it stood for something real, and substitute that inference for the harder work of actually checking whether it does. The attribute doesn't need to be explained or justified for the inference to fire, its specificity alone is what reads as meaningful.

Where it causes errors

A real product can be preferred over a functionally identical, or even better, competitor purely because a company invented a specific-sounding grade, an in-house certification with no external body behind it, a numbered rating with no published standard, and a buyer has no easy way to tell a decorative label apart from one backed by something real. The harder a category is to evaluate directly, the more room this has to work.

Where it can help

A visible, specific detail is a legitimate way to signal a real difference a buyer would otherwise have to take entirely on faith, provided the underlying claim is actually true: a garment genuinely woven more slowly on a genuinely scarcer machine, a material genuinely sourced to a stricter, real standard. The line is whether the label points at something that would still be true if nobody ever asked about it.

Carpenter, G. S., Glazer, R., & Nakamoto, K. (1994). “Meaningful Brands from Meaningless Differentiation: The Dependence on Irrelevant Attributes.” Journal of Marketing Research, 31(3), 339–350
A set of controlled experiments testing whether an invented, functionally irrelevant product attribute could still shift how buyers rated quality.
StrengthDirectly manipulated whether the irrelevant attribute was present and how it was framed across multiple experiments, letting the authors isolate exactly when the meaningless-attribute inference held and when it broke down, rather than relying on a single one-off comparison.
WeaknessBoth experiments rated hypothetical brands, not real purchases: 99 master's students in the first experiment and a separate group of undergraduates in the second placed marks on a preference scale for fictional products, so the size of the effect on an actual buying decision, with real money changing hands, wasn't directly tested.
Key findings
Verbatim“The average rating for the target brand with regular down fill in the subjective condition was 3.1; among those exposed to the same brand that included alpine class fill, the average brand rating climbed to 9.1.”
ParaphrasedThe effect replicated across three unrelated product categories, down jackets, pasta, and CD players, each given its own fictional irrelevant attribute (alpine class fill, authentic Milanese style, a studio-designed signal processing system), suggesting the inference isn't limited to one kind of product.
ParaphrasedThe effect weakened once the same attribute was explicitly and clearly labelled as irrelevant to the product's actual performance, suggesting the inference depends on the attribute's meaning staying ambiguous, not on buyers being deceived outright about what it was.
See also: Signpost Effect, related but distinct: that reframes an already-meaningful fact to activate a goal that was sitting dormant; this takes a fact with no inherent meaning at all and reads it as meaningful purely because it's there.
See also: Halo Effect, related but distinct: that extends one already-established trait to a different, unrelated trait of the same source; this manufactures the first impression out of an attribute that was never actually informative about anything.
65

Licensing Effect (moral licensing)

Why did shoppers who'd just imagined volunteering choose the $50 designer jeans over an equally priced vacuum cleaner?

Completing one virtuous or responsible act gives people a felt permission to indulge right afterwards, even when the two choices have nothing practical to do with each other.

THE SAME $50 CHOICE GOES DIFFERENTLY DEPENDING ON WHAT CAME FIRST IMAGINES VOLUNTEERING FIRST CHOOSES THE DESIGNER JEANS Feels licensed by the earlier virtuous act NO PRIOR COMMITMENT CHOOSES THE VACUUM CLEANER Same price, same options, nothing to license The only real difference between them: which choice came first.
Likely mechanismA virtuous act temporarily boosts the sense of being a good person, and that boost gets spent on an indulgence next

The psychology. Doing something that reflects well on the self, a donation, a workout, choosing the healthy option, creates a temporary boost to a person's self-concept. That boost behaves like a credit balance: having just proven they're a good or responsible person, someone feels licensed to spend some of that credit on a choice a plain, context-free evaluation wouldn't have permitted, without it registering as a contradiction, because the two acts get treated as morally fungible even when they're practically unrelated.

The full write-up: study, numbers, and caveats
The psychology

Doing something that reflects well on the self, a donation, a workout, choosing the healthy option, creates a temporary boost to a person's self-concept. That boost behaves like a credit balance: having just proven they're a good or responsible person, someone feels licensed to spend some of that credit on a choice a plain, context-free evaluation wouldn't have permitted, without it registering as a contradiction, because the two acts get treated as morally fungible even when they're practically unrelated.

Where it causes errors

“Buy one, give one” and similar cause-linked campaigns can quietly work this lever on the buyer, not just the recipient: a customer who's just funded a donation can feel more entitled to the indulgent version of what they were buying anyway. The effect can also cut the other way and reduce the good deed itself: in the cited study, people who'd first imagined a virtuous commitment went on to donate less, when they did donate at all, than people given no such prior commitment, the licence didn't just permit indulgence, it took a bite out of the generosity it followed.

Where it can help

The honest use of this isn't tricking anyone into feeling licensed, it's noticing when a business's own design is accidentally handing out licences it didn't mean to. Bundling a virtuous choice and an indulgent one into a single decision, rather than presenting them as two separate moments, removes the gap the licensing effect needs to operate in. A genuinely earned credential, a real completed workout, a real paid-down balance, can still be a fair, motivating waypoint, as long as it's not later used to wave through a choice that undoes the point of the first one.

Khan, U., & Dhar, R. (2006). “Licensing Effect in Consumer Choice.” Journal of Marketing Research, 43(2), 259–266
Five lab studies, mostly asking participants to imagine a prior virtuous commitment before choosing between a luxury and a plain option priced identically.
StrengthRan the effect across five separate studies varying both the virtuous act and the subsequent choice, and tested why it happens directly: a mediation analysis pointed to a temporary self-concept boost, and the effect shrank when the virtuous act was attributed to external pressure rather than free choice, a real boundary condition, not just a repeated correlation.
WeaknessThe virtuous choice in most of the five studies was imagined rather than actually performed, participants pictured committing to volunteer, they didn't really do it, leaving open how strongly the effect holds when the earlier good deed is real rather than hypothetical.
Key findings
Verbatim“Significantly more people in the license condition chose the designer jeans (57.4%) than in the control condition (27.7%).”
ParaphrasedThe same pattern replicated with a different pair of choices: 56.5% of participants in the license condition chose the more expensive, hedonic sunglasses, versus 27.7% in the control condition. In a separate study using a real donation instead of a hypothetical choice, among participants who did donate afterwards, average donations were lower in the licensing condition ($1.20) than in the control condition ($1.70).
Paraphrased62% of participants chose the designer jeans when the earlier virtuous act was freely chosen, versus 40% in the control condition, but that gap disappeared (45% versus 40%, not significant) when the virtuous act was instead framed as a penalty for a driving violation. A separate mediation analysis supported a temporary boost to self-concept, not general mood, as the actual driver.
Also worth citing: the broader “moral credentialing” finding this consumer-choice version builds on originates in social psychology, Monin, B., & Miller, D. T. (2001), “Moral Credentials and the Expression of Prejudice,” Journal of Personality and Social Psychology, 81(1), 33–43, though a 2026 registered replication report found no reliable evidence for that original, differently-scoped effect, worth knowing before leaning on the older paper specifically.
See also: Present Bias, related but distinct: this is about which choice gets permitted after an earlier virtuous act; that's about why an immediate temptation wins even with no virtuous act anywhere in the picture.
See also: Behavioural Economics vs. Behavioural Science vs. Psychology vs. UX, on why the applied practitioner name for this effect comes from Khan and Dhar's economics paper, not Monin and Miller's earlier psychology paper above.
Tested as an experiment: Does directing a tax refund to pay down a credit card increase spending right after?, taking the same mechanism from an already-earned reward onto a real windfall and a real debt.
66

Present Bias (the present self vs. the future self)

Why did the same person pick the fruit for a week from now, then pick the junk food the moment it actually arrived?

People weigh a reward available right now far more heavily than the identical reward available later, so a choice made calmly in advance and the choice actually made in the moment can flatly contradict each other, even though nothing about the options changed.

THE SAME PERSON CHOOSES DIFFERENTLY DEPENDING ON HOW FAR AWAY “NOW” IS CHOOSES A WEEK IN ADVANCE PICKS THE FRUIT Both options equally delayed, easy to compare calmly CHOOSES IN THE MOMENT, HUNGRY PICKS THE JUNK FOOD Same person, same original choice, now immediate The only thing that changed: how far away “now” was.
Likely mechanismThe value of a reward drops steeply the moment it's delayed at all, so “now” beats “later” out of proportion to the actual wait

The psychology. Standard theory assumes a person discounts a future reward at a constant rate, valuing next week's version almost as much as this week's. Real choice doesn't work that way: the drop in value from “right now” to “delayed at all” is disproportionately steep, then flattens out further into the future. That shape lets the same person genuinely prefer the healthy option a week out, when both choices are equally delayed, and then reverse that exact preference the instant “now” becomes one of the options.

The full write-up: study, numbers, and caveats
The psychology

Standard theory assumes a person discounts a future reward at a constant rate, valuing next week's version almost as much as this week's. Real choice doesn't work that way: the drop in value from “right now” to “delayed at all” is disproportionately steep, then flattens out further into the future.

That shape lets the same person genuinely prefer the healthy option a week out, when both choices are equally delayed, and then reverse that exact preference the instant “now” becomes one of the options. That's not because they changed their mind. It's because being in the room with immediate temptation triggers a completely different, present-favouring calculation than picturing the same choice in advance.

Where it causes errors

Someone signs up for a gym membership or a savings plan fully confident in their disciplined future self, then the actual moment to act arrives and an immediate competing priority wins instead, not through a change of values, but through the same timing effect the cited study documents directly. The study's own term for the gap is an “empathy gap”: a calm, well-fed decision-maker genuinely can't predict what a hungry, in-the-moment version of themselves will choose.

Where it can help

This is the direct justification for structuring a choice so it's made once, calmly, in advance, rather than re-litigated at the moment of maximum temptation, the specific tool for doing that is Precommitment Devices. Short of a binding commitment, simply defaulting to the advance choice unless a customer deliberately overrides it in the moment already removes some of the timing advantage the present self would otherwise get for free.

Read, D., & van Leeuwen, B. (1998). “Predicting Hunger: The Effects of Appetite and Delay on Choice.” Organizational Behavior and Human Decision Processes, 76(2), 189–205
A real two-stage snack-choice study: participants chose between a healthy and unhealthy snack to be delivered a week later, then were offered the chance to change that choice the moment it actually arrived, either hungry or freshly fed.
StrengthCompared the same people's advance choice against their own in-the-moment choice for an identical future reward, isolating the timing of the decision itself as the variable, and manipulated real hunger state at the moment of choice to test the mechanism directly rather than just observing the reversal.
WeaknessThe choice was a low-stakes snack, tested among 200 real employees at their actual workplaces in Amsterdam rather than in a lab, so whether the same size reversal holds for higher-stakes financial or health decisions with real, larger consequences wasn't directly tested here.
Key findings
Verbatim (abstract)“participants were dynamically inconsistent: they chose far more unhealthy snacks for immediate choice than for advance choice.”
Verbatim“participants who were hungry during immediate choice (those in conditions HH and SH) chose more unhealthy snacks than did those who were satisfied (HS and SS).”
ParaphrasedThe same individuals were measurably inconsistent across the two time points, what the authors call dynamic inconsistency, confirming the reversal was a real change in one person's revealed preference, not just a difference between two separate groups.
See also: Precommitment Devices, related but distinct: that's the tool that removes the future self's ability to relitigate a choice; this is the underlying bias that makes relitigating so likely to go the present self's way.
See also: Licensing Effect, related but distinct: that's about which choice gets permitted after an earlier virtuous act; this is about why immediate temptation wins even with no virtuous act anywhere in the picture.
Also worth reading: The Biases Draining Your Super Could Also Fill It, on the 3.05 million Australians who withdrew $37.8 billion of retirement savings early during the pandemic.
67

Sunk Cost Fallacy

Why does money already spent decide how much more gets spent?

Money, time, or effort already sunk into something keeps pulling people towards continuing it, even once continuing is the worse choice going forward.

PAYING FULL PRICE FOR THE SEASON MEANT SEEING MORE OF IT FULL PRICE SEASON TICKETS, PAID IN FULL 4.11 PLAYS ATTENDED DISCOUNTED SAME SEASON, RANDOM DISCOUNT ~3.3 PLAYS ATTENDED The season of plays was identical either way. Only how much each person had already paid to see it differed.
Likely mechanismStopping after spending money on something feels like admitting that money was wasted, so people keep going instead

The psychology. Standard economic theory says only future costs and benefits should matter to a decision; money already spent is gone regardless of what happens next, and a purely rational choice would ignore it completely. Real decisions don't work that way. Continuing to invest in something, a project, a purchase, a course of action, tracks how much has already gone into it, because stopping would mean admitting that earlier spending bought nothing, and that admission itself feels like a cost.

The full write-up: study, numbers, and caveats
The psychology

Standard economic theory says only future costs and benefits should matter to a decision; money already spent is gone regardless of what happens next, and a purely rational choice would ignore it completely. Real decisions don't work that way. Continuing to invest in something, a project, a purchase, a course of action, tracks how much has already gone into it, because stopping would mean admitting that earlier spending bought nothing, and that admission itself feels like a cost.

Where it causes errors

A renovation that's blown its budget keeps getting more budget, because stopping now would mean the money already spent bought an unfinished house instead of a finished one. A failing internal project gets another quarter of headcount because killing it would make the last four quarters look wasted, not because the new quarter is actually expected to succeed where the last four didn't. A gym membership someone stopped enjoying gets renewed because of what the first year cost, a cost the renewal itself does nothing to recover.

Where it can help

Awareness of the effect is what makes a genuine "kill" decision possible: a business that sets a stop-loss review or a kill criterion before a project starts, in advance of any money being spent, gives itself a rule that isn't contaminated by sunk cost once the money is actually on the table. The honest, decision-improving move is building that checkpoint early, not resisting the pull to keep going once it's already engaged.

Arkes, H. R., & Blumer, C. (1985). “The Psychology of Sunk Cost.” Organizational Behavior and Human Decision Processes, 35(1), 124–140
The paper that named the effect, combining hypothetical decision scenarios with a real field experiment on actual theatergoers' behaviour.
StrengthThe Ohio University theatre experiment randomly assigned real season-ticket buyers to pay full price or a discounted price for the identical season, then tracked their actual attendance, a real behavioural measure with random assignment, not a hypothetical scenario asking what people say they'd do.
WeaknessSeveral of the paper's other findings, including the widely cited "radar-blank airplane" scenario, are hypothetical vignettes asking what someone would choose, not real money on the line, so those specific numbers describe stated intentions rather than observed behaviour.
Key findings
ParaphrasedIn the hypothetical airplane scenario, 85% of participants chose to continue funding a doomed project once told $10 million had already been invested in it, compared with just 17% choosing to continue when the identical remaining decision was described with no mention of prior investment.
Verbatim“The no-discount group used significantly more tickets (4.11) than both the $2 discount group (3.32) and the $7 discount group (3.29).”
Verbatim (abstract)“Evidence that the psychological justification for this behavior is predicated on the desire not to appear wasteful is presented.”
See also: Endowment Effect, a related but distinct pull: that's about overvaluing something because it's already owned, this is about continuing to invest because of what's already been spent, whether or not anything is owned yet.
68

Temporal Reframing

Why does the same yearly cost feel smaller as a daily one?

Restating a cost as a small recurring amount instead of one aggregate total changes what it gets compared against, and makes the identical spend feel easier to say yes to.

THE SAME CHARITABLE ASK GOT MORE YESES FRAMED BY THE DAY $300/YEAR ASKED FOR AS ONE YEARLY TOTAL 30% AGREED 85¢/DAY ASKED FOR AS A DAILY AMOUNT 52% AGREED Roughly the same yearly total, asked for two different ways. The daily-coin framing got nearly twice the yes rate.
Likely mechanismA cost framed as a daily amount gets compared to small daily expenses, not the big ones it's actually replacing

The psychology. Deciding whether a price is worth paying involves comparing it against some other expense already in mind. Restating an aggregate cost as a small recurring one changes which expenses come to mind for comparison: a "$300 a year" framing calls up other big, infrequent purchases, while an "85 cents a day" framing calls up small, everyday ones, coffee, a snack, spare change. The recurring frame wins because it's being judged against a cheaper set of rivals, not because the underlying maths changed.

The full write-up: study, numbers, and caveats
The psychology

Deciding whether a price is worth paying involves comparing it against some other expense already in mind. Restating an aggregate cost as a small recurring one changes which expenses come to mind for comparison: a "$300 a year" framing calls up other big, infrequent purchases, while an "85 cents a day" framing calls up small, everyday ones, coffee, a snack, spare change. The recurring frame wins because it's being judged against a cheaper set of rivals, not because the underlying maths changed.

Where it causes errors

A subscription, an extended warranty, or an insurance add-on quoted as "just a few dollars a day" can secure agreement it wouldn't get quoted as its real annual total, letting someone commit to an ongoing cost they'd have hesitated over if they'd seen the aggregate figure first. The payment schedule doesn't change; only which number gets shown, and which smaller expenses that number gets measured against, does.

Where it can help

The identical reframe can make a genuinely beneficial habit feel achievable instead of daunting: a savings app that shows a retirement contribution as "the price of a coffee a day" isn't hiding the total, it's making a number that would otherwise feel too large to start suddenly feel small enough to begin. The honest version discloses the real annual total alongside the daily figure; the manipulative version shows only the smaller number and hopes nobody does the multiplication.

Gourville, J. T. (1998). “Pennies-a-Day: The Effect of Temporal Reframing on Transaction Evaluation.” Journal of Consumer Research, 24(4), 395–408
A series of lab studies testing the "pennies-a-day" (PAD) framing across a range of product and donation categories, and the two-step comparison process behind it.
StrengthThe paper doesn't just show the PAD effect exists, it tests the proposed two-step mechanism directly (which comparison expenses get retrieved, and how that retrieval changes the transaction evaluation), giving a specific causal account rather than a described correlation.
WeaknessThese are lab studies measuring stated willingness to agree or donate, not real long-term subscription or donation behaviour, so whether the reframe also changes how people feel about the cost once it has actually been recurring for months isn't something this design can answer.
Key findings
Verbatim“the percentage of subjects agreeing to donate was significantly higher under the PAD than under the aggregate framing (52 vs. 30 percent; X2(1) = 4.66, p < .05).”
Verbatim (abstract)“The PAD framing of a target transaction is shown to systematically foster the retrieval and consideration of small ongoing expenses as the standard of comparison, whereas an aggregate framing of that same transaction is shown to foster the retrieval and consideration of large infrequent expenses. This difference in retrieval is shown to significantly influence subsequent transaction evaluation and compliance.”
ParaphrasedGeneral support for the pennies-a-day effect held across a range of product categories tested in the paper's lab studies, not just the charity example.
See also: Chunking, a related but distinct move: that's about breaking one large ask into several smaller ones over time, this is about restating the same single total in smaller-feeling units without changing the schedule at all.
69

Bundling

Why does a discount stated on the bundle feel bigger than the same discount split across its parts?

The identical dollar saving carries more weight in how good a deal feels when it's stated directly on a combined bundle price than when the same amount is spread across separate item-level discounts.

THE SAME $20 SAVED FELT BIGGER STATED ON THE BUNDLE, NOT THE ITEMS $10 + $10 TWO ITEMS, EACH DISCOUNTED SEPARATELY FELT LIKE A SMALLER DEAL $20 OFF THE BUNDLE OF BOTH, DISCOUNTED ONCE FELT LIKE A BIGGER DEAL Twenty dollars saved, either way. Stated as one bundle discount, the same saving carried more weight.
Likely mechanismA single discount stated on the combined price registers as one bigger gain than the identical amount split into smaller pieces

The psychology. Perceived savings in a bundle come from two separate sources people track differently: what each item would have cost bought alone, and any additional saving offered directly on the bundle price itself. The second kind carries more weight than the first, dollar for dollar, because it registers as one direct, unambiguous gain rather than several smaller ones a buyer has to notice and add up themselves.

The full write-up: study, numbers, and caveats
The psychology

Perceived savings in a bundle come from two separate sources people track differently: what each item would have cost bought alone, and any additional saving offered directly on the bundle price itself. The second kind carries more weight than the first, dollar for dollar, because it registers as one direct, unambiguous gain rather than several smaller ones a buyer has to notice and add up themselves.

Where it causes errors

A retailer can make an identical total discount feel far more generous by moving it from the item level to the bundle level, "save an extra $20 when you buy both together" instead of "$10 off each," without the total saved changing at all. A buyer can end up choosing a bundle over the better-fitting individual items purely because the bundle's saving was easier to feel, not because the bundle itself was the better deal.

Where it can help

When a seller genuinely is offering combined savings, stating the discount once, directly on the bundle, communicates the real value more clearly than scattering several small per-item price cuts a buyer might not notice or bother adding up. The honest version states an accurate combined saving; the manipulative version inflates the bundle discount relative to what the items were actually worth apart.

Yadav, M. S., & Monroe, K. B. (1993). “How Buyers Perceive Savings in a Bundle Price: An Examination of a Bundle's Transaction Value.” Journal of Marketing Research, 30(3), 350–358
A study separating a bundle's perceived value into two distinct components, savings on the individual items and additional savings on the bundle itself, and testing how each one contributes.
StrengthIsolating two types of "savings" that get blended together in most bundle-pricing discussions, and testing their separate and interacting effects, gives a more precise account than a general claim that bundling increases perceived value.
Weakness252 undergraduate students at a state university rated a printed advertisement for a garment bag and a pullman suitcase; like most classic bundle-pricing studies of this era, they evaluated a described pricing scenario rather than an actual purchase with real money changing hands, a design typical of consumer-judgment research from this period.
Key findings
Verbatim (abstract)“Such perceptions of overall bundle savings may consist of two separate perceptions of savings, each with a different relative influence: (1) perceived savings on the individual items if purchased separately and (2) perceived additional savings on the bundle.”
Verbatim (abstract)“Results of an experiment indicate that additional savings offered directly on the bundle have a greater relative impact on buyers' perceptions of transaction value than savings offered on the bundle's individual items.”
Verbatim (abstract)“The effect of each saving is also influenced by the magnitude of the other saving.”
See also: Mental Accounting, a related but distinct habit: that's about which account a cost or saving gets filed under, this is about whether a saving gets stated once as a single gain or scattered across several smaller ones.
70

Clawback (Loss-Framed Incentives)

Why did factory teams work harder to avoid losing a bonus they'd already been paid than to earn the exact same bonus by hitting a target?

The identical reward moves people more when it's framed as something already theirs and at risk of being taken back than when the same reward is framed as something still to be earned.

SAME BONUS, TOLD TWO DIFFERENT WAYS ONE VERSION FELT LIKE SOMETHING TO PROTECT $ TOLD: HIT THE TARGET, EARN THE BONUS TARGET HIT, BONUS PAID AS PROMISED $ TOLD: IT'S ALREADY YOURS, DON'T LOSE IT PRODUCTIVITY ROSE ABOUT 1% MORE Same dollar bonus, same production target, in both versions. Only the framing changed: protecting it beat earning it fresh.
Likely mechanismA reward framed as already yours and losable motivates harder than the same reward only promised ahead

The psychology. A bonus is normally judged against a reference point of having nothing extra: hit the target and you gain something you didn't have. Loss aversion means a loss against your reference point hurts noticeably more than an equivalent gain pleases, so anything that shifts the reference point upward, even just by saying the money's already yours, turns the same future outcome into a potential loss instead of a potential gain. Missing the target no longer feels like failing to earn a bonus; it feels like losing money you already had.

The full write-up: study, numbers, and caveats
The psychology

A bonus is normally judged against a reference point of having nothing extra: hit the target and you gain something you didn't have. Loss aversion means a loss against your reference point hurts noticeably more than an equivalent gain pleases, so anything that shifts the reference point upward, even just by saying the money's already yours, turns the same future outcome into a potential loss instead of a potential gain. Missing the target no longer feels like failing to earn a bonus; it feels like losing money you already had.

Where it causes errors

The same mechanism turns punitive when the "already given" framing isn't matched by an already-given reality. Signing and relocation bonus clawback clauses, common in tech and finance hiring, require an employee to repay money they already spent if they leave before a set date, often years out.

Executive-compensation clawback rules the SEC finalised in 2022 under the Dodd-Frank Act work the same way at a larger scale: incentive pay already paid out gets recovered from a current or former executive after an accounting restatement, sometimes years after they earned it. Both are real, disclosed policies, but a reward that can be taken back years after the fact reads less like motivation and more like a debt that was never really settled.

Where it can help

Used honestly, the same lever helps rather than exploits. Two real Australian banks apply an almost identical loss frame to bonus interest: CommBank's GoalSaver, and Savers from Up, an app-only neobank. Both structure the rate so a customer earns it by default each month and gets warned, in plain language, before an easy lapse costs them the difference.

Nothing here is hidden or punitive: the condition is disclosed upfront, the warning arrives with enough time to act on it, and the "loss" being protected is a rate the customer already qualifies for, not one invented to manufacture urgency. The same reference-point shift that makes a clawed-back executive bonus feel like a betrayal makes a timely, honest warning feel like the bank is on the customer's side.

Hossain, T., & List, J. A. (2012). “The Behavioralist Visits the Factory: Increasing Productivity Using Simple Framing Manipulations.” Management Science, 58(12), 2151–2167
A real field experiment inside an operating Chinese high-tech manufacturing plant, run on actual production teams and their real weekly bonus pay, not a lab simulation or a survey.
StrengthA genuine field experiment inside a real, operating factory in Nanjing, China, run over roughly six months (July 2008 to January 2009) with 165 workers and real RMB 80 bonuses (about USD 11.72, more than 20% of a week's pay for the lowest-paid workers) tied to real measured output, not a hypothetical scenario or a stated-preference survey.
WeaknessThe effect was measured at one Chinese high-tech manufacturing site, with its own labour market, culture, and task type; whether the same loss-frame edge holds at different bonus sizes, industries, or for individual white-collar work rather than team piece-work wasn't tested here.
Key findings
Verbatim (abstract)“The magnitude of the effect is roughly 1%: that is, total team productivity is enhanced by 1% purely due to the framing manipulation.”
Verbatim (abstract)“Conditional incentives framed as both ‘losses’ and ‘gains’ increase productivity for both individuals and teams… neither the framing nor the incentive effect lose their importance over time; rather the effects are observed over the entire sample period.”
See also: Loss Aversion, the broader phenomenon this is one specific, deliberately engineered application of: a reward reframed as a possession at risk, not a prospect still to win.
Decoded on The Science Behind: CBA's GoalSaver and Up Bank both apply this same loss frame honestly to bonus interest, warning a customer before an easy lapse costs them a rate they'd already qualify for.
71

Disclosure Backfire (Sunlight That Doesn't Disinfect)

Why did warning people that their advisor had a real financial incentive to inflate the advice make that advice more biased, not less?

Disclosing a conflict of interest is meant to let the person receiving advice discount it appropriately. It can instead free the person giving the advice to lean into their bias further, while the listener still doesn't discount enough to make up the difference.

SAYING THE INCENTIVE OUT LOUD MADE THE ADVICE MORE BIASED ADVISOR SAYS NOTHING ABOUT BEING PAID TO ADVISE HIGH ESTIMATE LANDS A LITTLE HIGH ADVISOR TELLS THE ESTIMATOR THEY'RE PAID TO ADVISE HIGH ESTIMATE LANDS EVEN HIGHER Naming the incentive out loud didn't fix the bias. It grew, and the estimator still didn't discount enough to catch it.
Likely mechanismAdmitting a bias aloud can feel like licence to act on it more, not a check that reins it in

The psychology. Conflicted advice usually comes from a real incentive: a commission, a referral fee, a stake in the outcome. Requiring the advisor to disclose that incentive is meant to let the person receiving advice discount it appropriately, the same way a reader discounts a sponsored review. Two things get in the way. Moral licensing lets a disclosed advisor feel they've already done their duty, freeing them to lean into the bias rather than correct for it. And the listener rarely discounts enough to cancel that increase out, partly because rejecting disclosed advice can feel like accusing the advisor of dishonesty to their face.

The full write-up: study, numbers, and caveats
The psychology

Conflicted advice usually comes from a real incentive: a commission, a referral fee, a stake in the outcome. Requiring the advisor to disclose that incentive is meant to let the person receiving advice discount it appropriately, the same way a reader discounts a sponsored review. Two things get in the way. First, moral licensing: having disclosed the conflict, an advisor can feel they've already done their duty, which frees them to lean into the bias rather than correct for it, since the listener has, in theory, been warned.

Second, the listener rarely discounts enough to cancel that increase out, partly because rejecting disclosed advice can feel like accusing the advisor of dishonesty to their face, a social cost most people would rather avoid than a purely financial calculation would predict. The two effects compound: the advice gets more biased, and the correction for it doesn't grow to match.

Where it causes errors

A financial advisor paid on commission is legally required, in most markets, to disclose that fact before recommending a product. The disclosure is real and the products themselves are usually legitimate. What it doesn't reliably do is make the recommendation less commission-driven.

An advisor who has disclosed the conflict has, in a real sense, already done what the rule asked of them. A client who has just been told the advisor earns more if they buy still has to weigh the discomfort of pushing back against an advisor sitting across the desk from them. The rule creates the appearance of a safeguard without necessarily changing the underlying incentive it was built to neutralise.

Where it can help

The fix that actually works isn't better wording on the disclosure itself. It's removing the incentive the disclosure was trying to flag in the first place. A fee-only financial advisor, paid a flat rate regardless of what they recommend, can disclose exactly how they're compensated without triggering the same backfire, because there's no percentage-based upside left for them to feel licensed to lean into. The honest use of this principle is diagnostic, not persuasive: if disclosing a conflict is the primary fix on offer, that's a signal the underlying incentive hasn't actually been dealt with yet.

Cain, D. M., Loewenstein, G., & Moore, D. A. (2005). “The Dirt on Coming Clean: Perverse Effects of Disclosing Conflicts of Interest.” The Journal of Legal Studies, 34(1), 1–25
A set of controlled lab experiments built around a real coin-jar estimation task, where advisors who could see the actual value gave advice to estimators who couldn't, with a real, sometimes-disclosed financial incentive for the advisor to advise high.
StrengthA genuine controlled experiment, not a survey or a hypothetical scenario: 147 Carnegie Mellon University undergraduates were randomly assigned to advisor and estimator roles judging real jars of coins, with a real financial incentive to inflate advice built into the advisor's own payment structure, not just described to participants.
WeaknessThe task's biased advisors weren't licensed professionals with a reputation or a compliance department to answer to, and the stakes, while real, were the modest dollar amounts typical of a lab task; whether the same-sized backfire holds for a professional advisor with more on the line wasn't tested here.
Key findings
Verbatim (abstract)“First, people generally do not discount advice from biased advisors as much as they should, even when advisors’ conflicts of interest are disclosed. Second, disclosure can increase the bias in advice because it leads advisors to feel morally licensed and strategically encouraged to exaggerate their advice even further.”
Verbatim“estimators earned less money when conflicts of interest were disclosed than when they were not, and advisors made more money with disclosure than without disclosure.”
Also worth citing: Loewenstein, G., Sah, S., & Cain, D. M. (2012). “The Unintended Consequences of Conflict of Interest Disclosure.” JAMA, 307(7), 669–670, a follow-up perspective piece by two of the same authors arguing the same dynamic helps explain why conflict-of-interest disclosure rules in medicine and finance haven't reliably improved outcomes for the patients and clients they're meant to protect.
See also: The Fine Print Nobody Reads, a full report on why standard disclosure fails this way so often, and what research shows actually works instead.
72

Propinquity Effect

Why did two neighbours sharing a hallway end up close friends four times more often than two who lived at opposite ends of that same hallway, with nothing else about them different?

People who cross paths often, simply because they live or work near each other, become more familiar with each other than they would otherwise, and that familiarity gets read as trust or fit, even in a decision meant to be judged on merit alone.

SAME HALLWAY, SAME BUILDING, SAME MONTH MOVED IN 19 FT APART pass each other daily 41% named each other a close friend DOORS 89 FEET APART, OPPOSITE ENDS OF THE HALL rarely cross paths 10% named each other a close friend
Likely mechanismRepeated, incidental contact from physical closeness breeds familiarity, and familiarity gets mistaken for a signal of trust or fit

The psychology. Two neighbours don't need to like each other to become close, they just need to keep running into each other. A doorway on the way to the mailbox, a hallway shared on the way out each morning, these create dozens of small, low-stakes, unplanned contacts that a friendship can grow out of, with no decision on either side to seek the other out. The bias hiding inside that mechanism is what makes it worth naming: the resulting familiarity feels like it was earned through genuine compatibility, not manufactured by an apartment floor plan neither person chose.

The full write-up: study, numbers, and caveats
The psychology

Two neighbours don't need to like each other to become close, they just need to keep running into each other. A doorway on the way to the mailbox, a hallway shared on the way out each morning, these create dozens of small, low-stakes, unplanned contacts that a friendship can grow out of, with no decision on either side to seek the other out.

What matters isn't raw physical distance so much as "functional distance", how often two people's actual paths cross given the layout they share, which is why a resident living next to a stairwell or mailbox everyone passes gets named as a friend more often than their straight-line distance from others would predict. The bias hiding inside that mechanism is what makes it worth naming: the resulting familiarity feels like it was earned through genuine compatibility, not manufactured by an apartment floor plan neither person chose.

The same shortcut runs on judgments that have nothing to do with friendship. A scout, a manager, or an evaluator who happens to cross paths with a candidate more often develops a felt sense of "knowing" that candidate, and that felt sense gets folded into a judgment that's supposed to be about merit alone.

Where it causes errors

Major League Baseball's amateur draft is run by each team's own scouting director, evaluating thousands of prospects against a shared pool of statistics, scouting reports, and game film. A 2022 study following the 2000–2019 drafts (roughly 30,000 players drafted from a pool of over a million) found that a player is measurably more likely to get drafted by a given team the closer he lives to that team's scouting director, holding his actual on-field skill constant using machine learning to flexibly control for it.

The players who benefited from that geographic closeness turned out to be worse bets, not better ones: they were 38% less likely to ever appear in a single MLB game than a similarly-drafted player without the proximity advantage, and the effect was strongest in the draft's later rounds, exactly where a scouting director has the most personal discretion and the least outside scrutiny. Nobody in that process decided to favour nearby players on purpose. The mechanism is the same one running in a housing complex, just wearing a different job.

Where it can help

The same lever runs in reverse when it's used on purpose instead of by accident. A new hire who's placed near the colleagues they'll actually need to collaborate with, rather than wherever a desk happened to be free, builds the cross-team familiarity that makes asking a quick question or flagging a real problem feel easy instead of awkward.

A mentorship programme that puts a junior and senior colleague on the same regular meeting, rather than leaving the relationship to form on its own, is deliberately engineering the same repeated, low-stakes contact a shared mailbox creates by accident. Used this way, the goal isn't to bias a decision, it's to build the familiarity that makes an already-good working relationship easier to reach.

Festinger, L., Schachter, S., & Back, K. (1950). Social Pressures in Informal Groups: A Study of Human Factors in Housing. Harper & Brothers
A real field study, not a lab experiment: married WWII veterans attending MIT on the GI Bill were assigned, essentially at random, to units in the newly built Westgate and Westgate West housing complexes, and later asked to name their three closest friends within the complex.
StrengthRandom assignment to housing units rules out the obvious alternative explanation, that people simply chose to live near others they already liked. Whatever friendship pattern emerged had to come from the physical layout itself, not from residents sorting themselves by compatibility first.
WeaknessThe entire sample was one highly specific population, married graduate-student veterans in a single housing complex in 1946, and friendship was measured by asking couples to name a short list of names, a self-report that leans on memory and social comfort as much as on an objective count of who actually spent time with whom.
Key findings
Paraphrased41% of next-door neighbours, roughly 19 feet apart, named each other as one of their three closest friends, compared with 22% of residents two doors apart and just 10% of residents at opposite ends of the same hallway, roughly 89 feet away.
ParaphrasedAbout 65% of all the close friendships residents named were with someone living in the same building, even though residents had just as much opportunity to befriend people in the complex's other buildings.
ParaphrasedResidents whose apartment sat next to a shared stairwell or mailbox, a spot everyone in the building had to pass, were named as a friend more often than their straight-line distance from others would predict, showing that how often paths actually crossed mattered more than raw physical distance alone.
A note on sourcing: the original 1950 monograph isn't available as searchable full text on this site's network, so all three findings above are paraphrased from the book's long-standing, widely-replicated figures as reported across independent academic summaries, not quoted verbatim from the original page.
Also worth citing: Ahmadi, M., Durst, N., Lachman, J., List, J. A., List, M., List, N., & Vayalinkal, A. (2022). “Nothing Propinks Like Propinquity: Using Machine Learning to Estimate the Effects of Spatial Proximity in the Major League Baseball Draft.” NBER Working Paper No. 30786, a modern corroborating study covering roughly 30,000 players drafted from 2000 to 2019, with player skill estimated by a machine-learning model built for the Chicago White Sox. A player is 7.1% more likely to be drafted by a given team for every 1,000 km he lives closer to that team's scouting director, controlling for skill, and that effect climbs to 11.8% in rounds 16 and later, where a director has the most personal discretion. Players who benefit from that proximity are 38% less likely to ever appear in an MLB game than a similarly drafted player without it, and their initial signing bonuses run 12% to 25% higher, conditional on draft order. Still a working paper rather than a peer-reviewed publication at time of writing.
See also: How do you prove distance biased a decision when nobody can run a controlled trial on real MLB scouts?, a full breakdown of how the draft study argues its case without being able to randomly assign where a scout lives.
Special report: Proximity Gets Mistaken for Merit reads the housing study and the draft study together, seventy years apart, as one argument.
73

Availability Heuristic

Why did one attack on planes get hundreds of people killed on the highway instead?

When something is easy to picture or recall, a recent headline, a vivid story, a memory of your own, it feels more common and more likely than it actually is, simply because it comes to mind so easily.

MOST PEOPLE GUESSED THE FIRST LETTER WAS COMMONER. IT ISN'T. EXAMPLES CAME TO MIND EASILY 105 OF 152 GUESSED FIRST LETTER THE REAL COUNT IN ENGLISH TEXT THIRD LETTER IS ABOUT 2X COMMONER The words that came to mind first felt like the honest count. The dictionary disagreed, by about two to one.
Likely mechanismEasily recalled examples feel more common than they really are

The psychology. The mind doesn't keep a running tally of how often something actually happens. Asked to judge how common or how risky something is, it substitutes an easier question: how quickly do examples come to mind? A vivid story, a recent headline, or a memorable personal experience surfaces instantly, and that speed of recall gets mistaken for a real measure of frequency.

The full write-up: study, numbers, and caveats
The psychology

The mind doesn't keep a running tally of how often something actually happens. Asked to judge how common or how risky something is, it substitutes an easier question: how quickly do examples come to mind? A vivid story, a recent headline, or a memorable personal experience surfaces instantly, and that speed of recall gets mistaken for a real measure of frequency.

That substitution breaks down exactly where memorability and actual frequency pull apart. Dramatic, unusual, or emotionally loaded events (a plane crash, a shark attack, a fraud story a friend told you) are disproportionately easy to recall, precisely because they're rare and vivid, not because they're common. The more available an example feels, the more people overestimate how often the thing it represents actually happens.

Where it causes errors

After the September 11 attacks, many Americans who would normally have flown chose to drive instead. It felt like the rational choice: the vivid, horrifying memory of hijacked planes made flying feel like the greater risk. Economist Gerd Gigerenzer compared U.S. Department of Transportation crash data for October through December 2001 against prior years. He found an estimated 353 additional road deaths in that single quarter, more than the roughly 250 people killed on the four hijacked flights themselves.

A later reanalysis by researchers Michael Sivak and Michael Flannagan put the shift even higher, closer to 1,000 additional traffic deaths. Flying had not become more dangerous per mile travelled. Driving was, and remains, far riskier. But the vivid, easily recalled image of a hijacked plane made that risk feel real in a way an ordinary highway death never does.

Where it can help

The countermeasure is to stop relying on memory alone for anything that actually matters. A 19-item checklist built by the World Health Organisation was tested across eight hospitals on four continents, from a rural clinic in Tanzania to a teaching hospital in Toronto. It forces a surgical team to explicitly confirm steps that are easy to skip in the moment, precisely because they rarely go wrong: the patient's identity, the surgical site, whether antibiotics were given on time.

Across roughly 8,000 patients, major complications fell from 11.0% to 7.0%, and in-hospital deaths fell from 1.5% to 0.8%. The checklist doesn't make anyone smarter or more careful. It just replaces whatever happens to come to mind in the moment with a fixed list that doesn't care what's vivid or recent.

Tversky, A., & Kahneman, D. (1973). “Availability: A Heuristic for Judging Frequency and Probability.” Cognitive Psychology, 5(2), 207–232
The paper that introduced the availability heuristic as a distinct judgment shortcut, alongside representativeness and anchoring, in Kahneman and Tversky's broader research programme on heuristics and biases.
StrengthThe letter-position experiment isolates the mechanism cleanly. The same 152 subjects judged five different letters (K, L, N, R, V), and the actual frequency of each letter in real English text was independently countable, giving a real, checkable answer to compare the guesses against.
WeaknessThe task is an artificial laboratory judgment about letter frequency, disconnected from any real-world risk or purchase decision. The paper alone doesn't establish how large the same effect is when something more consequential is being judged. That evidence comes from later work, like the traffic-fatality analysis cited below.
Key findings
Verbatim“Among the 152 subjects, 105 judged the first position to be more likely for a majority of the letters, and 47 judged the third position to be more likely for a majority of the letters… The median estimated ratio was 2:1 for each of the five letters. These results were obtained despite the fact that all letters were more frequent in the third position.”
Verbatim“It is certainly easier to think of words that start with a K than of words where K is in the third position. If the judgment of frequency is mediated by assessed availability, then words that start with K should be judged more frequent.”
Verbatim (abstract)“This paper explores a judgmental heuristic in which a person evaluates the frequency of classes or the probability of events by availability, i.e., by the ease with which relevant instances come to mind. In general, availability is correlated with ecological frequency, but it is also affected by other factors. Consequently, the reliance on the availability heuristic leads to systematic biases.”
Also worth citing: Gigerenzer, G. (2004). “Dread Risk, September 11, and Fatal Traffic Accidents.” Psychological Science, 15(4), 286–287, the real-world traffic-fatality analysis referenced above; a later reanalysis, Sivak, M., & Flannagan, M. J. (2004). “Consequences for Road Traffic Fatalities of the Reduction in Flying Following September 11, 2001.” Transportation Research Part F, 7(4–5), 301–305, found an even larger shift.
Also worth citing: Haynes, A. B., Weiser, T. G., Berry, W. R., et al. (2009). “A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population.” New England Journal of Medicine, 360(5), 491–499, the WHO Surgical Safety Checklist trial referenced above.
74

Diversification Heuristic

Why does splitting your money evenly depend on how many options are on the menu, not what's in them?

Asked to divide money or attention across several unfamiliar options, most people default to splitting it evenly across however many there are, a rule that has nothing to do with what each option actually contains.

SPLIT $1,000 ACROSS A 10-OPTION MENU EVENLY, WHATEVER GETS TAPPED 4 TAP 4 OF THE 10 OPTIONS $250 IN EACH 5 TAP 5 OF THE 10 OPTIONS $200 IN EACH The investor's real view of risk never changed. Only how many rows got tapped did.
Likely mechanismWeighing what each option actually holds is hard, so people just split evenly by count instead

The psychology. Weighing what each option actually contains, its real risk, its expected return, how much it overlaps with everything else on the list, takes effort most people don't have on hand in the moment. Instead of doing that comparison, the mind substitutes an easier question: how many options are there, so what's one divided by that? Splitting evenly feels like sensible diversification without ever checking what's actually inside each option.

The full write-up: study, numbers, and caveats
The psychology

Weighing what each option actually contains, its real risk, its expected return, how much it overlaps with everything else on the list, takes effort most people don't have on hand in the moment. Instead of doing that comparison, the mind substitutes an easier question: how many options are there, so what's one divided by that? Splitting evenly feels like sensible diversification without ever checking what's actually inside each option or how similar it is to the others.

This is why the rule is called naive diversification: it diversifies by count, not by content. Two menus holding the exact same underlying assets, one built from three funds and one built from six, get split three ways or six ways by someone using this rule, producing two different final portfolios from the identical set of building blocks, purely because of how many boxes the menu happened to draw around them.

Where it causes errors

CommSec Pocket's own "Invest in" menu lists ten flat options side by side: Aussie Top 200, Aussie Dividends, Aussie Sustainability, Aussie Corporate Bonds, Global 100, Diversified Equities, Emerging Markets, Health Wise, Sustainability Leaders, and Tech Savvy, with no default and no indication of how much to put in each. A first-time investor who wants "a spread" and taps four or five of them is likely to split evenly across whichever ones they happened to pick, never noticing that Diversified Equities already holds hundreds of companies across several markets, while Tech Savvy, Health Wise, and Emerging Markets are concentrated, higher-volatility single-theme bets, or that Aussie Top 200 and Aussie Dividends likely already share many of the same large ASX companies.

Benartzi and Thaler found the identical pattern in real employer retirement plans: the share of an employee's contributions held in stock funds tracked closely with the share of stock funds among the total funds the plan happened to offer. A plan that listed mostly stock funds ended up with mostly-stock employee portfolios, and a plan that listed mostly bond funds ended up bond-heavy, almost independent of what the employees actually said about their own risk tolerance.

Where it can help

Knowing that people will roughly divide evenly across whatever's offered turns the menu itself into a lever. A plan sponsor or app designer can curate a smaller, genuinely balanced set of options, or lead with one already-diversified default, and get a sensible outcome "by accident" from someone who would otherwise have equal-weighted a messier list. The menu's composition is itself a design decision with real consequences, not a neutral, complete list to be as long as possible.

Spotted in the wild

CommSec Pocket's real “Invest in” screen: ten flat options, no default, no grouping, and nothing telling a new investor how much of their money should go where.

CommSec Pocket's Invest in menu, showing ten flat investment options (Aussie Top 200, Aussie Dividends, Aussie Sustainability, Aussie Corporate Bonds, Global 100, Diversified Equities, Emerging Markets, Health Wise, Sustainability Leaders, Tech Savvy) with the full list outlined

CommSec Pocket, “Invest in” menu. Ten options, equal visual weight, no indication of how they overlap or how concentrated each one is. Captured 2026-08-28.

This screenshot shows a real, current menu, not a claim that any particular investor actually equal-weighted it. The evidence that people tend to do exactly that with menus shaped like this is the cited study below.

Benartzi, S., & Thaler, R. H. (2001). “Naive Diversification Strategies in Defined Contribution Saving Plans.” American Economic Review, 91(1), 79–98
The paper that named the “1/n” naive-diversification strategy in retirement saving.
StrengthCombines real field data from actual employer retirement plans, revealed behaviour with real money at stake, not a survey, with a controlled hypothetical-menu experiment that manipulated the number and composition of funds offered, letting the same 1/n pattern show up under both real stakes and clean experimental control.
WeaknessThe field-data correlation, more stock funds offered tracking more stock held, is open to a competing explanation: employers who choose to offer more stock funds might also differ in ways, workforce demographics, the employer's own risk philosophy, that independently push toward more stock holdings. The companion experiment addresses this concern but doesn't fully rule it out for the field data on its own.
Key findings
Verbatim (abstract)“We show that some investors follow the ‘1/n strategy’: they divide their contributions evenly across the funds offered in the plan.”
Verbatim (abstract)“Consistent with this naive notion of diversification, we find that the proportion invested in stocks depends strongly on the proportion of stock funds in the plan.”
ParaphrasedIn a companion hypothetical-menu experiment, University of California employees offered a choice among four fixed-income funds and one stock fund allocated an average of 43% of their contributions to equities, while a separate group offered one fixed-income fund and four stock funds, the same mix TWA's pilot plan actually uses, allocated 68%, a 25-point gap the authors found statistically significant even though a rational, mean-variance-optimizing investor would only have moved a few points.
Run this one live: Live Sessions: Diversification Heuristic, a real room tapping funds from CommSec Pocket's own flat 10-option menu versus a curated 3-option one, instead of the CommSec analogy alone.
Also worth reading: The Biases Draining Your Super Could Also Fill It, applying this same naive “1/n” pattern to a superannuation fund's own investment menu.
75

Ambiguity Aversion

Why would someone turn down a bet with unknown odds, even when it could easily beat a worse bet whose odds are known?

Given a choice between a risk with known odds and an otherwise identical risk with hidden odds, most people pick the known one, even when the hidden odds could be just as good or better.

THE URN WITH A HIDDEN SPLIT GETS FEWER BETS AT THE SAME ODDS RED50 BALLS BLACK50 BALLS URN A: KNOWN 50 RED, 50 BLACK MOST PEOPLE BET ON URN A ? URN B: SAME 100 BALLS, SPLIT HIDDEN FEWER PEOPLE BET ON URN B Same prize, same 100 balls, only the split is hidden. Urn B loses bettors at exactly the same odds as Urn A.
Likely mechanismAn unknown probability feels more threatening than a bad known one, so the unknown option gets avoided

The psychology. Risk and uncertainty aren't the same problem for the brain, even when they carry an identical expected outcome. Offered a choice between a bet with known odds and an otherwise identical bet with hidden odds, most people pick the known bet, and they do it even for bets they'd normally consider a bad deal. People aren't reacting to bad odds. They're reacting to not being told the odds at all.

The full write-up: study, numbers, and caveats
The psychology

Risk and uncertainty aren't the same problem for the brain, even when they carry an identical expected outcome. Offered a choice between a bet with known odds and an otherwise identical bet with hidden odds, most people pick the known bet, and they do it even for bets they'd normally consider a bad deal. People aren't reacting to bad odds. They're reacting to not being told the odds at all.

This asymmetry breaks a basic rule of rational choice. If someone won't bet on red from an urn with a hidden mix, they should be willing to bet on black instead, since between the two bets, one has to be at least a 50/50 shot. Daniel Ellsberg's 1961 demonstration showed people decline both bets on the ambiguous urn, a pattern no single, consistent probability judgment can produce. Something else is driving the choice: an aversion to ambiguity itself, on top of and separate from ordinary risk aversion.

Where it causes errors

Ambiguity aversion helps explain a well-documented puzzle in household investing: many people hold far less of their portfolio in stocks than any sensible risk calculation would suggest, and what stock they do hold skews heavily toward their own employer or their own country's market, the market whose rules feel familiar rather than hidden. A large US household survey found ambiguity-averse respondents were less likely to hold stocks at all, held a smaller share of their assets in stocks when they did, and held more of their own employer's stock specifically, exactly the pattern this principle predicts: familiar risk gets tolerated, unfamiliar risk gets avoided, independent of which one actually pays more on average.

Where it can help

A business that discloses the same policy plainly everywhere it appears removes an unnecessary source of ambiguity aversion: a customer who can see exactly what they're agreeing to doesn't need to imagine a worse hidden version of it. The same logic applies to any decision genuinely worth de-risking, an insurance product with a clear, simple contract, a warranty with one stated term rather than several conditional ones. Reducing genuine ambiguity is what actually lowers the aversion, not just reassuring someone that everything is fine.

Spotted in the wild

Real ambiguity doesn't always look hidden. Sometimes it's two different real numbers, both plainly published, with nothing telling the reader which one actually applies to them.

Apple's Returns & Refunds page, showing a seven-day return window for the Pick-up Point Return Policy and a 14 calendar day window for the Apple Store Returns Policy, both highlighted

Apple, Shopping Help: Returns & Refunds. Two return windows for two return methods on the same order: seven days for a Pick-up Point drop-off, 14 calendar days for an Apple Store return, both real, both currently published, with nothing on the page explaining why they differ. Captured 2026-08-28.

This is Apple's own published policy, not a claim that the difference confuses every reader or changes anyone's behaviour. It's a real instance of the kind of unresolved, unexplained variation that keeps a choice ambiguous rather than merely risky, the mechanism the cited study below actually tests.

Ellsberg, D. (1961). “Risk, Ambiguity, and the Savage Axioms.” Quarterly Journal of Economics, 75(4), 643–669
The paper that first distinguished ambiguity (unknown probabilities) from ordinary risk (known probabilities) as things people respond to differently, illustrated with a thought experiment about drawing coloured balls from two urns.
StrengthEllsberg's urn structure holds every other feature of the bet fixed, the same prize, the same total number of balls, the same type of bet, so the only real difference between the two urns is whether the colour split is known or hidden, isolating ambiguity from every other reason a bet might be avoided.
WeaknessEllsberg's own demonstration wasn't a controlled experiment. It was posed informally to a small group of colleagues at Harvard and RAND, without random assignment, a fixed sample, or reported percentages, more a philosophical challenge to expected-utility theory than a data-gathering study. The pattern he described was only confirmed under real experimental conditions in later work.
Key findings
ParaphrasedEllsberg described two urns, each holding 100 balls. Urn A is known to hold exactly 50 red and 50 black balls. Urn B holds red and black balls in some unknown ratio. Offered a bet that pays out for drawing red, most people he spoke with preferred betting on Urn A. Offered the identical bet for drawing black instead, most preferred Urn A again, a pattern that can't be reconciled with any single, fixed probability estimate for Urn B's contents.
Verbatim“What is at issue might be called the ambiguity of this information, a quality depending on the amount, type, reliability and ‘unanimity’ of information, and giving rise to one's degree of ‘confidence’ in an estimate of relative likelihoods.”
A note on sourcing: Ellsberg's original demonstration was an informal thought experiment described in his own paper, not a controlled study with a defined sample size or reported percentages; he reports typical responses (“if you are in the majority…”) rather than tabulated results.
Also worth citing: Dimmock, S. G., Kouwenberg, R., Mitchell, O. S., & Peijnenburg, K. (2016). “Ambiguity Aversion and Household Portfolio Choice Puzzles: Empirical Evidence.” Journal of Financial Economics, 119(3), 559–577, the real, quantified measurement behind the investing example above: a representative US household survey that measured each respondent's own ambiguity aversion directly, using questions built on Ellsberg's urn structure, then tested it against their real portfolio data. Ambiguity-averse respondents were less likely to hold any stocks at all, held less in stocks when they did, held more of their own employer's stock specifically, and were more likely to sell out of stocks during the 2008 financial crisis.
See also: Risk Aversion, the more familiar half of this picture: even with known odds, people weight a loss more heavily than a same-sized gain. Ambiguity aversion is what happens one level up, when the odds themselves aren't known.
76

Risk Aversion

Why do most people take a guaranteed $40 over a coin flip that pays $0 or $100?

Given a choice between a certain amount and a gamble worth the same or more on average, most people take the certain amount, giving up expected value in exchange for certainty.

THE SAME BET GETS SAFER CHOICES ONCE THE MONEY IS REAL SAFE40¢ SURE RISKY50/50, 80¢ LOW STAKES: A FEW CENTS EITHER WAY MOST PEOPLE PICK THE RISKY BET SAFE$40 SURE RISKY50/50, $80 HIGH STAKES: THE SAME BET, REAL MONEY MOST PEOPLE SWITCH TO THE SAFE BET Same coin flip, same payoff ratio, only the stakes changed. Raising the amount at risk pushes most people toward the sure thing.
Likely mechanismEach extra dollar of gain matters a little less than the last, so a sure amount can beat a gamble worth more on average

The psychology. Classical economics predicts people should judge a gamble purely by its expected value, the probability-weighted average of what it pays. Real choices depart from that in a specific way: each additional dollar of gain is worth a little less than the last one was, so a fixed amount already in hand can outweigh a gamble with a higher average payout.

The full write-up: study, numbers, and caveats
The psychology

Classical economics predicts people should judge a gamble purely by its expected value, the probability-weighted average of what it pays. Real choices depart from that in a specific way: each additional dollar of gain is worth a little less than the last one was, so a fixed amount already in hand can outweigh a gamble with a higher average payout, purely because of how the value of money itself curves as amounts grow. This is diminishing marginal utility, and it's a different mechanism from loss aversion's asymmetry between losses and gains: risk aversion shows up even in choices with no possible loss at all, just a sure gain against a bigger, less certain one.

How strongly it shows up depends heavily on how real the money is. Charles Holt and Susan Laury ran the same menu of paired lottery choices at both trivial payouts and payouts scaled up to serious money, and found people got measurably more risk averse as the stakes moved from small to large and from hypothetical to real. Simply asking someone to imagine a bigger number on paper doesn't reproduce the same shift; actual money on the table does.

Where it causes errors

Risk aversion on its own is a coherent preference, not a mistake. It becomes one when it's applied to the wrong timeframe: a retirement saver who avoids stocks entirely because the year-to-year swings feel threatening, even on a decades-long horizon where the swings mostly wash out, is applying a short-horizon risk judgment to a long-horizon decision. The aversion itself is real; the timeframe it's being applied to is where the error sits.

Where it can help

Understanding risk aversion cuts both ways, and the honest version of it is building products that reduce real risk, not just the feeling of it. A stated return policy is the clearest version. A real, published window to send back an unwanted purchase turns an otherwise irreversible bet into a reversible one, directly lowering the risk of buying something unseen or untried.

A 0%-interest instalment plan reduces a different kind of risk: the risk of the payment itself. If the terms are exactly as stated, spreading a known total cost over fixed payments removes the risk of triggering interest, not just its appearance. Either mechanism turns exploitative the moment it encourages a purchase someone couldn't otherwise afford or return, or hides a shortened window or a cost in the fine print.

Spotted in the wild

Two live examples of businesses reducing the real risk of a purchase, not just how risky it feels, on the exact pages where that risk aversion would otherwise kick in.

Apple's Returns & Refunds page, showing a seven-day return window for the Pick-up Point Return Policy and a 14 calendar day window for the Apple Store Returns Policy, both highlighted

Apple, Shopping Help: Returns & Refunds. A published 14-day return window for Apple Store purchases (seven days for a Pick-up Point drop-off) gives a shopper a real, stated right to undo the purchase, before they've even paid. Captured 2026-08-28.

An Apple.com iPhone product page banner reading Pay for your new iPhone over time with Afterpay Pay Monthly, 0% interest and no monthly fees for up to 24 months, highlighted

Apple.com, iPhone product page. Afterpay's “Pay Monthly” financing offer, 0% interest and no monthly fees for up to 24 months, sits directly above the phone line-up, reducing the payment risk of the same purchase instead. Captured 2026-08-28.

Both are the retailer's own advertised terms, not independent verification of every order's actual outcome. A shopper still has to check the specific terms that apply to them. What the two screenshots show is the tactic itself, in two different forms. A return window removes the risk of being stuck with the wrong product. An instalment plan removes the risk of an unexpected interest cost. Either can lower the resistance risk aversion would otherwise create.

Holt, C. A., & Laury, S. K. (2002). “Risk Aversion and Incentive Effects.” American Economic Review, 92(5), 1644–1655
The paper that popularised the paired lottery-choice task now widely used to measure an individual's own risk aversion directly, by finding the exact point where they switch from the safer bet to the riskier one.
StrengthEvery subject faced the identical menu of ten paired lottery choices, a safer, lower-variance pair against a riskier, higher-variance pair, at every step, so the only thing that varied within a subject was the specific probabilities on offer, giving a clean read on exactly where each person's own risk tolerance sat.
WeaknessThe sample was 212 university students across three US universities (Georgia State, the University of Miami, and the University of Central Florida), split across five separate payoff-scale treatments, with only 19 and 18 subjects respectively in the two highest, most expensive treatments, so the higher-stakes results in particular rest on a fairly small number of people.
Key findings
Verbatim“With real payoffs, risk aversion increases sharply when payoffs are scaled up by factors of 20, 50, and 90.”
Verbatim“Subjects facing hypothetical choices cannot imagine how they would actually behave under high-incentive conditions.”
See also: Loss Aversion, a related but distinct asymmetry: risk aversion is about the curve of value itself, present even in a choice with no possible loss, while loss aversion is specifically about losses hurting more than same-sized gains help.
See also: Ambiguity Aversion, the same caution taken one step further. Even once the odds are known, as they are in every choice above, most people still prefer the safer option. Hide the odds entirely, and the aversion gets stronger still.
Also worth reading: The Biases Draining Your Super Could Also Fill It, a real case of exactly this: a short-horizon risk judgment applied to a decades-long retirement balance during the 2020 market crash.
The formal name for this gap: Ecological validity: does the task even resemble the real thing?, on Five Ways an Experiment Can Be Right and Still Wrong, using this exact study as the case.
77

Positional Concern (Rank Beats the Number)

Why did most people in a real survey choose to earn half as much money, as long as it meant beating everyone around them?

Given a choice between an absolute amount of money and a smaller amount that still ranks above everyone else's, many people pick the smaller amount: how much they have relative to others can matter as much as, or more than, how much they actually have.

OFFERED MORE MONEY, BUT RANKED BEHIND MOST PICKED LESS MONEY, RANKED AHEAD YOU$100,000 THEM$200,000 WORLD A: DOUBLE THE MONEY, HALF AS FAR AHEAD REJECTED BY MOST RESPONDENTS YOU$50,000 THEM$25,000 WORLD B: HALF THE MONEY, TWICE AS FAR AHEAD CHOSEN BY 56% OF RESPONDENTS Half the income, in exchange for outranking everyone around them. Prices were held constant across both worlds. Only the ranking changed.
Likely mechanismBeating other people's amount can matter more than the actual amount received

The psychology. People don't evaluate an amount of money in isolation. They evaluate it against a reference group, and that comparison can carry real weight of its own, separate from the number itself. Solnick and Hemenway tested how much weight directly, asking 257 Harvard School of Public Health students, faculty and staff to choose between two hypothetical worlds with prices held constant: earning $50,000 a year while everyone else earns $25,000, or earning $100,000 a year while everyone else earns $200,000.

The full write-up: study, numbers, and caveats
The psychology

People don't evaluate an amount of money in isolation. They evaluate it against a reference group, and that comparison can carry real weight of its own, separate from the number itself. Solnick and Hemenway tested how much weight directly, asking 257 Harvard School of Public Health students, faculty and staff to choose between two hypothetical worlds with prices held constant: earning $50,000 a year while everyone else earns $25,000, or earning $100,000 a year while everyone else earns $200,000.

56% chose the first option, half the absolute income, in exchange for ranking above their peers instead of behind them. The same pattern showed up across other goods the survey tested: education, physical attractiveness, a supervisor's praise. The strength of it varied by domain, strongest for attractiveness and a supervisor's approval, and weakest for vacation time, where most respondents cared more about the actual number of days off than about beating anyone else's.

Where it causes errors

Positional concern can turn a genuinely good outcome into a bad one, purely through the frame it's compared against. A raise that objectively improves someone's life can register as a loss the moment they learn a peer got a bigger one, the objective gain overridden by the relative slippage.

Pay-transparency rollouts run into this directly: making salaries visible doesn't just reveal fairness problems, it manufactures new positional comparisons that didn't exist while the numbers were private, for employees whose own pay never actually changed.

Where it can help

The same mechanism can be used honestly, without manufacturing a comparison that wasn't there before. Showing someone their real progress against a relevant peer group, a savings rate against similar savers, a fitness goal against people at the same starting point, can motivate action a purely absolute number doesn't.

The honest version compares against a real, representative reference group and states the comparison plainly. The dishonest version cherry-picks a reference group designed to make the number look better or worse than it actually is.

Solnick, S. J., & Hemenway, D. (1998). “Is More Always Better? A Survey on Positional Concerns.” Journal of Economic Behavior & Organization, 37(3), 373–383
A survey of 257 faculty, students, and staff at the Harvard School of Public Health, asking respondents to choose between a lower-absolute/higher-relative outcome and a higher-absolute/lower-relative one, across income and several other domains.
StrengthThe core question held prices and lifestyle costs constant across both hypothetical worlds, isolating relative standing as the only real difference between the two options, and the same comparison was replicated across multiple domains beyond income (education, attractiveness, vacation time), not just asked once.
WeaknessEvery choice was a hypothetical survey question, not a real financial decision with real money at stake, the same cheap-talk gap this site's own Reading the Research content flags: what people say they'd choose isn't guaranteed to match what they'd actually do if their own real income were on the line.
Key findings
Verbatim (abstract)“Half of the respondents preferred to have 50% less real income but high relative income.”
Verbatim (abstract)“Concerns about position were strongest for attractiveness and supervisor's praise and weakest for vacation time.”
See also: Social Norm, a related but distinct comparison effect: Social Norm is about matching your own behaviour to what others are doing, Positional Concern is about how a payoff itself feels once you know what others got.
See also: Cheap Talk & Hypothetical Bias, the same stated-vs-real gap this study's own hypothetical design runs into.
78

Warm-Glow Giving (Giving Without Being Watched)

In a game where nobody could ever find out what you chose, why did most people still hand over almost a third of the money?

Even with complete, unaccountable control over money and zero consequence for keeping all of it, most people don't. Giving something away tends to feel rewarding on its own, separate from anything the giver could get back.

MOST DICTATORS GAVE SOMETHING AWAY EVEN THOUGH KEEPING IT ALL WAS RISK-FREE KEPT$7.20 GIVEN$2.80 STANDARD DICTATOR GAME: $10 TO SPLIT ALONE 28% GIVEN AWAY, ON AVERAGE KEPTNEARLY ALL GIVENNEAR $0 SAME GAME, BUT NOW FULLY ANONYMOUS GENEROSITY LARGELY DISAPPEARED The money could be kept in full, with zero chance of being found out. Giving still happened anyway, until anonymity became complete.
Likely mechanismGiving itself feels rewarding, even when the giver could never be seen or repaid

The psychology. A dictator game hands one person an amount of money and total, unaccountable power over how much, if any, to give to a second person who has no say in the matter and no way to retaliate or reward. Engel's 2011 meta-analysis pooled 616 treatments and 20,813 reconstructed individual dictator decisions from 131 published studies, the largest single look at how people actually behave when told, in effect, no one can stop you from keeping it all.

The full write-up: study, numbers, and caveats
The psychology

A dictator game hands one person an amount of money and total, unaccountable power over how much, if any, to give to a second person who has no say in the matter and no way to retaliate or reward. Engel's 2011 meta-analysis pooled 616 treatments and 20,813 reconstructed individual dictator decisions from 131 published studies, the largest single look at how people actually behave when told, in effect, no one can stop you from keeping it all.

Nearly two-thirds gave something anyway, and the average dictator handed over about 28% of the pot. Andreoni's term for this is warm-glow giving: a private, personal reward that comes from the act of giving itself, distinct from caring what happens to the person on the other end.

Where it causes errors

Warm glow rewards the feeling of having given, not the outcome the gift actually produces, so it can pull a decision toward whichever option feels most personally satisfying rather than whichever does the most good. Checkout round-up prompts, the “round up your total for charity” button many retailers and ride-share apps now show, work on exactly this mechanism: the act of tapping yes supplies the reward, whether or not the shopper has any idea which cause the change is going to or how effectively it's used.

Where it can help

The same mechanism, aimed honestly, can enable real generosity toward someone the giver will never meet and gets nothing back from. Caffè sospeso, the Naples tradition of pre-paying for a stranger's coffee, works because the giver never learns who receives it and asks for nothing in return, the exact shape of the dictator game itself, just built into a coffee bar instead of a lab.

Engel, C. (2011). “Dictator Games: A Meta Study.” Experimental Economics, 14(4), 583–610
A meta-analysis pooling 616 dictator-game treatments and 20,813 reconstructed individual decisions across 131 published studies, using regression to isolate the effect of specific design choices (anonymity, framing, recipient identity) while controlling for the others.
StrengthPooling over 20,000 real decisions across a hundred-plus independent studies overcomes the small-sample problem any single dictator-game experiment has on its own, and the regression approach isolates which specific design features (not just “dictator games in general”) actually move giving up or down.
WeaknessAveraging across a hundred-plus studies blends together very different protocols, different stake sizes, student versus non-student samples, face-to-face versus fully anonymous conditions, into one headline number. The 28% average obscures how much the studies actually disagree, from close to zero in the most anonymous designs to well above that in the least.
Key findings
Verbatim“If one calculates the grand mean from all reported or constructed means per 616 treatments, dictators on average give 28.35% of the pie.”
Verbatim“16.74% choose the equal split.”
ParaphrasedGuaranteeing anonymity from the experimenter too, the double-blind design Hoffman, McCabe and colleagues introduced in 1994, made little difference on its own across Engel's pooled data. Only once the analysis controlled for whether the game was one-shot or repeated did double-blind designs show a weak, small reduction in generosity, suggesting full experimenter anonymity matters only in more complex designs, not as a blanket effect.
Also worth citing: Hoffman, E., McCabe, K., & Smith, V. (1996). “Social Distance and Other-Regarding Behavior in Dictator Games.” American Economic Review, 86(3), 653–660. A related moderator: generosity fell further as the social distance between dictator and recipient grew, even without the full anonymity manipulation above.
See also: Reciprocity, a related but distinct mechanism: Reciprocity is triggered by receiving something first, Warm-Glow Giving happens with no prior gift, no future interaction, and no possibility of being thanked.
79

Bulletproof Glass Effect

Why can a company's own privacy promise make customers trust it less, not more?

A prominent, detailed privacy or security notice can backfire. Instead of feeling reassured, people infer that anything worth explaining this much protection against must be a real threat, and trust drops rather than rises.

A DETAILED PRIVACY NOTICE MADE CUSTOMERS TRUST THE STORE LESS TRUSTBASELINE BUYBASELINE CHECKOUT PAGE, NO PRIVACY NOTICE SHOWN TRUST AND PURCHASE INTENT: BASELINE TRUSTLOWER BUYLOWER SAME CHECKOUT, PROMINENT PRIVACY NOTICE ADDED TRUST AND PURCHASE INTENT: LOWER The notice described only real, existing protections, nothing alarming. Explaining the protection in this much detail was itself the problem.
Likely mechanismExplaining a protection in detail implies the danger it guards against is real

The psychology. Most managers expect a visible privacy or security notice to make customers feel safer. Brough, Norton, Sciarappa and John tested that assumption directly, combining a content analysis of real companies' privacy notices, a survey of managers, a field experiment with real customers at checkout, and five further online experiments.

The full write-up: study, numbers, and caveats
The psychology

Most managers expect a visible privacy or security notice to make customers feel safer. Brough, Norton, Sciarappa and John tested that assumption directly, combining a content analysis of real companies' privacy notices, a survey of managers, a field experiment with real customers at checkout, and five further online experiments.

They found the opposite of what most managers predicted. A prominent, detailed notice made customers trust the retailer less and buy less, even when the notice described only protections the company genuinely had in place. The authors named it after the same logic as bulletproof glass on a house: a visitor doesn't feel safer seeing it, they wonder what the house is expecting to be protected from.

Where it causes errors

A team redesigning a checkout page to look more trustworthy can do real damage with a well-intentioned security badge or a long, detailed data-protection paragraph. The added detail reads as a threat signal, not a reassurance, and can quietly suppress conversions the team never connects back to the change that caused it.

Where it can help

The finding doesn't mean privacy information should disappear, only that its shape matters. A brief, benevolence-framed statement, plainly saying the company protects customer data because it respects their privacy, reassures without itemising specific defences in a way that implies a specific danger. A company can still meet its disclosure obligations while choosing language that doesn't do its own scaring.

Brough, A. R., Norton, D. A., Sciarappa, S. L., & John, L. K. (2022). “The Bulletproof Glass Effect: Unintended Consequences of Privacy Notices.” Journal of Marketing Research, 59(4), 739–754
A multi-method study: a content analysis of publicly traded companies' actual privacy notices, a survey of managers' expectations, a field experiment with real customers at a real checkout, and five online experiments isolating the mechanism.
StrengthThe field experiment tested the effect on real customers making real purchase decisions, not just hypothetical scenarios, and the multi-method design (content analysis, manager survey, field test, online experiments) checked the same finding from several independent angles rather than resting on one study alone.
WeaknessThe field experiment's effect was real but modest in absolute terms: enrollment at Borrowell, a financial technology firm whose sign-up already asked for income and credit-report access, dropped from 41.48% to 39.66%, a 1.82-percentage-point difference, when the privacy notice was made more prominent. That the effect held in a context already involving sensitive financial data undercuts the idea that data sensitivity is a natural exemption; the boundary condition the studies actually identify is prior trust, not data category, since Study 4 found the effect reversed once customers already had reason to distrust the company.
Key findings
Verbatim (abstract)“formal privacy notices undermined consumer trust and decreased purchase interest even when they emphasized objective protection (Studies 2, 3, and 5) or omitted any mention of potentially concerning data practices (Study 6).”
Verbatim (abstract)“These unintended consequences did not occur, however, when consumers had an a priori reason to be distrustful (Study 4) or when benevolence cues were added to privacy notices (Studies 5 and 6).”
80

Implementation Intentions (If-Then Planning)

Why did simply deciding when and where you'd do something more than double the odds you actually did it?

Deciding in advance exactly when, where and how you'll act on a goal, a specific if-then plan rather than a general intention, makes people far more likely to actually follow through.

THE SAME GOAL, ONLY ONE GROUP PLANNED WHEN AND WHERE COMPLETED32% GOAL ONLY: “I'LL WRITE THE REPORT” MOST NEVER WROTE IT COMPLETED71% PLUS: “WHEN X, THEN I'LL WRITE IT” MORE THAN DOUBLE THE COMPLETION RATE The goal was identical: write about Christmas Eve, mail it within 48 hours. Only one group decided in advance exactly when and where.
Likely mechanismA specific plan hands control of the action to a cue, so it fires automatically

The psychology. A general intention states a goal but not the moment of action, leaving the decision of when to actually do it to be made later, in the moment, when hesitation, forgetting or a competing demand can crowd it out. Gollwitzer and Brandstätter tested a sharper alternative: an implementation intention, a plan in the form “when situation Y arises, I will do X,” that links a specific future moment directly to the intended response, before the moment itself arrives.

The full write-up: study, numbers, and caveats
The psychology

A general intention states a goal but not the moment of action, leaving the decision of when to actually do it to be made later, in the moment, when hesitation, forgetting or a competing demand can crowd it out. Gollwitzer and Brandstätter tested a sharper alternative: an implementation intention, a plan in the form “when situation Y arises, I will do X,” that links a specific future moment directly to the intended response, before the moment itself arrives.

College students were asked to write a report on how they'd spent Christmas Eve and mail it within 48 hours of Christmas. Half the group also specified exactly when and where they would sit down and write it. 71% of that group completed the report on time, versus 32% of the group given only the general goal, more than double, from the same intention, the same deadline and the same task.

Where it causes errors

A specific plan solves the problem of forgetting or hesitating to start. It doesn't solve a problem the plan was never built to fix: not having the actual skill, resources or ability the goal requires. Forming a detailed if-then plan can feel like the hard part is handled, when the real barrier was never about remembering to begin.

A separate research program tested this directly and found implementation intentions helped with easy goals and with difficult goals whose main obstacle really was getting started, but showed no benefit for a certain class of genuinely difficult goals, ones the plan alone couldn't actually make achievable.

Where it can help

This is one of the more directly useful findings on this site: naming the exact when and where for something you already want to do is a small, low-cost, low-pressure way to close the gap between intending and doing. A savings transfer set for the moment a paycheck lands, a specific gym slot rather than “exercise more,” a fixed time to review a subscription before it renews, are all the same plan structure applied honestly, to something the person already decided they wanted.

Gollwitzer, P. M., & Brandstätter, V. (1997). “Implementation Intentions and Effective Goal Pursuit.” Journal of Personality and Social Psychology, 73(1), 186–199
Three studies: a correlational study of college students' own difficult goals; an experiment assigning the same difficult goal (writing and mailing a Christmas Eve report within 48 hours) to everyone, with half the group additionally specifying exactly when and where they would act; and a lab study testing whether people who had formed an implementation intention actually seized the specified opportunity the moment it appeared.
StrengthThe experimental study held the goal, the deadline and the task completely constant across both groups, isolating the plan itself, not a different or easier goal, as the only explanation for the gap in completion rates.
WeaknessThe task was a one-off, low-stakes report for a psychology study, not a recurring real-world habit with competing demands, temptations or a cost to failure. Whether the same size effect holds for harder, ongoing goals (exercise, saving, medication) is a separate empirical question this specific study doesn't answer on its own.
Key findings
Verbatim“When the goal intention of Study 2 was furnished with implementation intentions, however, completion rate drastically increased from 32% to 71%.”
Verbatim (abstract)“In correlational Study 1, difficult goal intentions were completed about 3 times more often when participants had furnished them with implementation intentions.”
Also worth citing: Gollwitzer, P. M., & Sheeran, P. (2006). “Implementation Intentions and Goal Achievement: A Meta-Analysis of Effects and Processes.” Advances in Experimental Social Psychology, 38, 69–119. A meta-analysis across dozens of studies, confirming implementation intentions work through several distinct processes, not just one, including goal shielding: naming a specific obstacle in advance and planning a response to it protects the original goal from that obstacle when it actually shows up.
Also worth citing: Dewitte, S., Verguts, T., & Lens, W. (2003). “Implementation Intentions Do Not Enhance All Types of Goals: The Moderating Role of Goal Difficulty.” Current Psychology, 22, 73–89. The corroborating study behind the “Where it causes errors” section above.
Also worth citing: Nickerson, D. W., & Rogers, T. (2010). “Do You Have a Voting Plan? Implementation Intentions, Voter Turnout, and Organic Plan Making.” Psychological Science, 21(2), 194–199. A real field experiment, 287,228 voters ahead of the 2008 US presidential election: a call that helped people form a specific voting plan (what time, coming from where, doing what beforehand) raised turnout by 4.1 percentage points overall, and by 9.1 points in single-voter households, a genuine, ethical, large-scale use of this exact mechanism.
See also: Reminder Fatigue, the real cost of pushing this mechanism too far: a well-timed reminder helps follow-through, but each additional one carries a measurable risk of losing the relationship entirely.
81

Reminder Fatigue

Why does a reminder that works today also make people cut off contact for good?

A repeated reminder raises compliance the moment it's sent. Each additional one also raises the odds the recipient opts out of all future contact, and the two effects get tracked separately, so the win is easy to see and the cost is easy to miss.

A REMINDER PAYS THE CHARITY BUT THE COST FALLS ON THE NON-GIVER NET BENEFIT TO THE CHARITY $0.18 ON AVERAGE, PER REMINDER SENT WELFARE LOST BY EACH NON-GIVER $2.35 AN ECONOMIST'S ESTIMATE, PER REMINDER The charity nets a little. The people who don't give lose thirteen times as much.
Likely mechanismRepeating an ask reads as ignoring an earlier no, until people cut ties entirely

The psychology. A reminder does double duty. It restates the specific ask, and it also tells the recipient something about the relationship: that the sender is willing to make contact again after an unanswered first attempt. Below some threshold, that reads as helpful. Past it, repetition itself starts to read as the sender not respecting an earlier silence, and the response isn't a smaller no to this one ask. It's a decision to close the channel for good.

The full write-up: study, numbers, and caveats
The psychology

A reminder does double duty. It restates the specific ask, and it also tells the recipient something about the relationship: that the sender is willing to make contact again after an unanswered first attempt. Below some threshold, that reads as helpful, a nudge someone might genuinely have wanted.

Past it, repetition itself starts to read as the sender not respecting an earlier silence or decline. The response isn't a smaller no to this one ask. It's a decision to close the channel for good, a different and much larger cost than simply lowering the odds of any single future ask.

Where it causes errors

A campaign that measures only the metric it's nudging, this quarter's donations, this month's click-through, can look like an unambiguous win. It can also be quietly costing a business its highest-value, longest-tenured relationships, the ones with the most left to lose by leaving for good. The same trap applies to appointment reminders, renewal notices, and app push notifications: each additional nudge can lift the immediate number while raising the odds this specific person opts out of ever being reachable again.

Where it can help

The same mechanism argues for real restraint, not cleverer wording. Cap how often a specific ask repeats, retire a reminder once it's been ignored a set number of times, and reserve repetition for asks that are genuinely time-sensitive. Don't default to it just because it measurably works in the short run. A well-designed reminder system counts a permanent opt-out as a real cost of every message sent, weighed against the value of that one extra nudge before the next one goes out.

Damgaard, M. T., & Gravert, C. (2018). “The Hidden Costs of Nudging: Experimental Evidence from Reminders in Fundraising.” Journal of Public Economics, 157, 15–26
A large-scale field experiment with a real charity's own donor list and reminder programme.
StrengthA genuine field experiment on a charity's actual donors, not a lab simulation. The authors also went further than reporting two separate percentages. They structurally estimated a welfare model translating both the donation gain and the unsubscribe cost into a single comparable unit, dollars, instead of leaving the reader to informally weigh one against the other.
WeaknessThe setting is one charity's email fundraising programme. Whether the same size effect holds for a different channel, a push notification, a phone call, a text message, wasn't directly tested here. Each of those carries its own cost of sending and its own cost of being ignored, which could change the tradeoff entirely.
Key findings
Verbatim (abstract)“we find that reminders increase donations, but they also substantially increase unsubscriptions from the mailing list.”
Verbatim (abstract)“reminders are welfare diminishing for the potential donors as non-givers incur a welfare loss of $2.35 for every reminder. The net benefit of every reminder to the charity is $0.18.”
Verbatim“the Low Frequency treatment reduces the unsubscription rate from 0.49% to 0.30%.”
Also worth citing: Ancker, J. S., Edwards, A., Nosal, S., Hauser, D., Mauer, E., & Kaushal, R. (2017). “Effects of Workload, Work Complexity, and Repeated Alerts on Alert Fatigue in a Clinical Decision Support System.” BMC Medical Informatics and Decision Making, 17, Article 36. A completely different domain, real clinical alerts inside electronic health records, found the identical shape of effect: the likelihood a clinician acted on a reminder dropped by roughly 30% for each additional reminder received in the same encounter.
See also: Not Enough Choice, a related but distinct mechanism: that's resistance to a restricted set of options; this is resistance to being contacted again after an earlier silence or decline went unanswered.
See also: Implementation Intentions, the mechanism this principle is a caution against overusing: a well-timed reminder or plan-prompt genuinely helps follow-through, but each additional one carries a real, measurable cost this site's own reminder-call and confirmation-screen experiments are built to respect.
Decoded on Reading the Research: Backfires Have Six Different Causes, where this exact study appears as “The Charity Reminder Boomerang,” the sixth case in a roundup of interventions that looked like wins on the metric being watched.
82

Biased Assimilation

Why did the same two conflicting studies leave two people who disagreed even further apart than before?

Shown identical, mixed evidence, people don't move toward agreement. They accept whatever confirms what they already believed, pick apart whatever contradicts it, and end up more certain of their original view than when they started, on both sides at once.

READERS GRADE THE EVIDENCE BY WHICH SIDE THEY'RE ON BELIEVES IT DETERS MURDER RATES LOWER RATES HIGHER CALLED CONVINCING CALLED FLAWED MORE SURE IT DETERS MURDER DOUBTS IT DETERS MURDER RATES LOWER RATES HIGHER CALLED FLAWED CALLED CONVINCING MORE SURE IT DOESN'T Same two studies, read by both sides. Everyone left more certain than before.
Likely mechanismEvidence that confirms a belief gets accepted; evidence that contradicts it gets picked apart

The psychology. People don't apply one consistent standard when judging whether a piece of evidence is any good. Evidence that lines up with what they already believe gets accepted close to at face value. Evidence that contradicts it gets read with far more scrutiny: the sample was too small, the method was flawed, the comparison wasn't fair. The same argument passes or fails a credibility test depending on which side it happens to support.

The full write-up: study, numbers, and caveats
The psychology

People don't apply one consistent standard when judging whether a piece of evidence is any good. Evidence that lines up with what they already believe gets accepted close to at face value. Evidence that contradicts it gets read with far more scrutiny: the sample was too small, the method was flawed, the comparison wasn't fair. The same argument passes or fails a credibility test depending on which side it happens to support.

That double standard compounds once someone reads a full mix of evidence, some supporting their view, some against it. They keep the confirming half and discount the disconfirming half, so the same balanced set of facts leaves both sides more sure of their original position than before they read anything. Nobody needs biased information for this to happen. Genuinely mixed evidence is enough on its own.

Where it causes errors

The death penalty debate the original study used is still going, decades later, for exactly this reason. Research on whether it actually deters murder has reached genuinely mixed conclusions, some studies finding a measurable effect, some finding none, some finding the opposite. Each side of the public debate keeps citing whichever studies confirm what they already believed and dismissing the rest as bad research, so the same body of evidence that should be narrowing the disagreement has done nothing to move it in decades.

The same pattern shows up anywhere an ambiguous result gets treated as proof. Two people can look at the same sales dashboard, survey, or test result and each walk away more convinced their own prior read was right, because each one graded the noisy parts of the data by the study's own rule: convincing if it agrees, flawed if it doesn't.

Where it can help

The fix works by removing the thing biased assimilation actually needs: the ability to tell, before judging a result, which side it supports. A properly randomised experiment with one clear control group makes that much harder to do, because there's one predetermined comparison everyone agreed to in advance, rather than a pile of ambiguous evidence each side can sift through afterward. Blind peer review, where a reviewer judges a paper without knowing whose theory it supports, runs on the same principle.

Lord, C. G., Ross, L., & Lepper, M. R. (1979). “Biased Assimilation and Attitude Polarization: The Effects of Prior Theories on Subsequently Considered Evidence.” Journal of Personality and Social Psychology, 37(11), 2098–2109
A lab experiment giving real undergraduate supporters and opponents of the death penalty two conflicting fictitious research studies to evaluate.
StrengthThe two fictitious studies used two different real research designs, a comparison across states with and without the death penalty and a before-and-after comparison within one state that adopted it, and the researchers deliberately swapped which conclusion went with which design across participants. That rules out the simplest alternative explanation, that one method just looked more rigorous, since either method could carry either conclusion.
WeaknessThe subjects were undergraduates responding to two short fictitious study summaries on one hot-button issue, inside a single lab session. Whether the same size of effect holds for people weighing real, higher-stakes evidence over months, an investment thesis, a policy rollout, a medical diagnosis, wasn't tested here.
Key findings
Paraphrased48 students who identified as either strong supporters or strong opponents of the death penalty each read two fictitious research summaries: one concluding the death penalty lowers murder rates, one concluding it doesn't. Each side rated whichever study matched their existing belief as more convincing and better conducted than the one that contradicted it.
ParaphrasedThe two fictitious studies used two genuinely different real-world research designs, a cross-state comparison and a within-state before-and-after comparison, and which conclusion was attached to which design was swapped across participants, so the pattern couldn't be explained by one design simply looking more credible on its own.
ParaphrasedAfter reading the identical pair of mixed studies, both supporters and opponents of the death penalty rated themselves as more confident in their original position than before they'd read anything, even though the two studies, taken together, offered no net evidence either way.
A note on sourcing: this site's network access blocked every attempt to reach the primary paper itself (university PDF mirrors and academic databases all returned network-policy errors), so every finding above is paraphrased from independently corroborating secondary sources describing the study's own design and results, not quoted verbatim from the original paper.
See also: Availability Heuristic, a related but distinct judgment error: that's misjudging how common something is because examples of it are easy to recall; this is misjudging how convincing evidence is because of which side it happens to support.
Run this one live: Live Sessions: Biased Assimilation, seeding two groups with opposite priors, showing both the identical result, then watching the room's ratings and excuses sort themselves by belief instead of evidence.
83

Probability Weighting (why a tiny chance of a huge prize outweighs its real odds)

Why does a 1-in-a-million chance to win feel worth chasing?

People don't weigh probabilities the way arithmetic says they should. A very small chance of a big prize gets treated as far more likely than it really is, while a very high chance of the same outcome gets treated as slightly less certain than it really is.

THE SAME OVERWEIGHTED SMALL CHANCE PULLS PEOPLE TWO DIFFERENT WAYS A TINY CHANCE OF WINNING BIG FEELS WORTH CHASING The overweighted small chance of a gain A TINY CHANCE OF LOSING BIG WORTH INSURING AGAINST The same overweighted small chance, now a loss Same tiny probability, weighted the same way each time. Only whether it's a gain or a loss decides which way it pulls.
Likely mechanismThe mind turns a tiny real chance into a felt chance that's far bigger than the actual odds warrant

The psychology. Standard decision theory says a 1% chance should carry exactly 1% of the weight in a choice, and a 99% chance exactly 99%. Real judgment doesn't track probability on that straight line. Low probabilities get inflated well past their true size, and high probabilities get discounted slightly below theirs, so the felt difference between “impossible” and “a tiny chance” looms far larger than the felt difference between “very likely” and “certain,” even when the actual gap in both cases is the same few percentage points.

The full write-up: study, numbers, and caveats
The psychology

Standard decision theory says a 1% chance should carry exactly 1% of the weight in a choice, and a 99% chance exactly 99%. Real judgment doesn't track probability on that straight line. Low probabilities get inflated well past their true size, and high probabilities get discounted slightly below theirs, so the felt difference between “impossible” and “a tiny chance” looms far larger than the felt difference between “very likely” and “certain,” even when the actual gap in both cases is the same few percentage points.

That distortion runs in both directions off the same tiny probability. Pointed at a gain, it's what makes a long-shot prize feel worth chasing. Pointed at a loss, the identical-sized sliver of risk is what makes an unlikely disaster feel worth insuring against, which is exactly why the same study that explains lottery tickets also explains why people buy travel insurance for a flight that's very unlikely to be cancelled.

The size of the distortion tracks how the prize actually feels, not just its odds. Later research found people overweight a small chance of something emotionally vivid, a kiss from a favourite celebrity, an electric shock, even more than they overweight an equally small chance of a plain cash amount. The more a prize makes someone feel something before they've won anything, the harder the real odds are to weigh correctly.

Where it causes errors

A promotion that bundles a routine, everyday purchase with “a chance to win” a huge prize doesn't need to state real odds for the effect to work, and often doesn't. Almost any small number reads as similarly enticing once it's framed as a chance to win, so the actual likelihood barely matters to how attractive the whole deal feels. The same overweighting sells extended warranties on cheap appliances, add-on insurance for low-risk events, and scratch-card-style loyalty promotions whose true expected value is a fraction of a cent.

Where it can help

The same overweighting can be pointed at something the person and the business both actually want, and the upside isn't only strategic. A real field experiment found that simply holding a lottery-style ticket measurably lifts a person's happiness in the hours before the draw, whether or not it wins anything, so the hope itself is a genuine, if modest, good, not a trick played on whoever's holding it.

A charity or a health service offering entry into a prize draw for completing a blood donation, a survey, or a screening can lift participation for less than the cost of paying everyone a small guaranteed fee, because the felt value of a small chance at a big reward beats the felt value of a certain small one, even when the certain payment is worth more on paper. Naming the real odds honestly, rather than hiding them, is what keeps this use in the encouraging column rather than the exploitative one.

Kahneman, D., & Tversky, A. (1979). “Prospect Theory: An Analysis of Decision under Risk.” Econometrica, 47(2), 263–291
One of the most cited papers in economics: a series of choice problems, each pitting a gamble against an equivalent sure outcome, run across groups of real respondents to map how people actually weigh probability rather than assuming they weigh it the way expected-utility theory says they should.
StrengthThe same distortion, overweighting small probabilities, underweighting large ones, showed up again and again across dozens of differently worded choice problems and multiple respondent groups, rather than resting on one single question, giving the pattern real robustness within the study itself.
WeaknessEvery choice was a hypothetical gamble on paper with no real money actually paid out to respondents, so what people said they'd choose isn't guaranteed to match what they'd do with real stakes on the table, a gap later research using real-money incentives has tested directly.
Key findings
Verbatim (abstract)“people underweight outcomes that are merely probable in comparison with outcomes that are obtained with certainty. This tendency, called the certainty effect, contributes to risk aversion in choices involving sure gains and to risk seeking in choices involving sure losses.”
ParaphrasedThe paper's decision-weighting function shows small probabilities carrying more weight in a choice than their true value justifies, which the authors note helps explain why both long-shot gambling and insurance against rare losses feel more compelling than their real odds alone would predict.
ParaphrasedPeople also simplify choices between multiple risky options by ignoring whatever the options share and focusing only on what differs between them, an “isolation effect” that can produce different preferences for the objectively identical choice depending only on how it's presented.
See also: Illusion of Control, related but distinct: that's false confidence in a genuinely random outcome once skill-like cues are added to it; this is misjudging the raw odds themselves, before any skill framing enters the picture at all.
Also worth citing: Rottenstreich, Y., & Hsee, C. K. (2001). “Money, Kisses, and Electric Shocks: On the Affective Psychology of Risk.” Psychological Science, 12(3), 185–190, the study behind the emotionally vivid prize comparison above: the same small probability gets weighted even more heavily once the outcome itself carries real emotional charge.
Also worth citing: Burger, M. J., Hendriks, M., Pleeging, E., & van Ours, J. C. (2020). “The Joy of Lottery Play: Evidence from a Field Experiment.” Experimental Economics, 23(4), 1235–1256, the field experiment behind the anticipatory-happiness finding above: buying the ticket itself, not just winning, is where the measured wellbeing gain shows up.
Also worth reading: The Lottery You Can't Lose, a real product built on this same overweighting of small probabilities, pointed at a savings goal instead of a wager.
Decoded on The Science Behind: Why does a $2.50 fries deal come with a chance to win free food for life?, a real fast-food app promotion built on exactly this mechanism.
Try this one live: Does a chance at 0% increase home loan applications more than the guaranteed discount alone?, an experiment blueprint testing the same mechanism at a much higher stake than a savings rate or a fries deal.