Every one starts from a mechanism already covered on this site, a real field session or a real cited study. Browse by product area, or jump straight to any blueprint by name.
Not run yet: full blueprints. Every entry here starts from a mechanism already covered on this site, a real field session or a real cited study, and works out what it would take to actually test the next step: the theory, an assumed baseline, control vs. treatment, the sample size needed to trust the result, and the metrics that would prove it one way or the other. Search by name, or filter by the kind of product you'd run it in.
The mechanism. Mental Accounting says money doesn't sit in one undifferentiated pool in people's heads. It gets sorted into labelled accounts, and a labelled account is protected in a way an unlabelled one isn't. Salience does the actual labelling work, and not by being more memorable in the abstract, but by contrast: a name you can picture (“Save for Phone”) stands out against the generic, interchangeable rows around it (“Savings Pocket 2,” “Everyday Account”), much like one bold number jumping off an otherwise plain bill.
Already tested.The high school field session tested the labelling half of this live, with paper envelopes and a room of Year 8/9 students, and watched $30 accumulate in the “save” envelopes that nobody consciously decided to save.
This experiment. It tests the contrast half directly: does a name that stands out in a list of accounts protect a balance as well as a name that stood out on a hand-labelled envelope did?
Hypothesis
Prompting a customer to name a new savings pocket at the moment they open it, rather than defaulting to an auto-generated label, will reduce the share of that pocket's balance withdrawn to zero within 90 days, without adding enough friction to meaningfully hurt onboarding completion.
Assumed baseline
Illustrative, not real product data. Assume 30% of customers who open a second savings pocket withdraw its full balance within 90 days, the generically-labelled pocket behaving less like “savings” and more like a second everyday-spend account with extra taps.
Control vs. treatment
The only thing that changes between the two arms is who chooses the label. Everything else about opening a pocket stays identical.
Control · auto-named pocket
Savings Pocket 2
A new pocket, automatically named. Nothing to fill in, one tap and it's created.
Treatment · named at creation
What's this pocket for?
Give it a name you'll picture later, it's the easiest way to keep track of what this money's already spoken for.
PhoneHolidayEmergency
Sample size to actually trust this
Baseline30%90-day full withdrawal
Target24%−6pp, ~20% relative
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm856customers
Total sample1,7122 arms
Assumed volume800/wknew second pockets
Est. run time~3 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number. Swap in your own weekly pocket-creation rate and the run time follows.
Outcome variables
✓Primary: % of pocket balance withdrawn to $0 within 90 days (lower is the win)
✓Secondary: average days-to-first-withdrawal: does naming delay the first dip, not just reduce the total
✓Guardrail: pocket-creation completion rate: the naming step must not become a reason people abandon the flow
✓Guardrail: time-to-complete onboarding, and the % who skip the naming prompt where a skip option exists
The mechanisms behind this:Mental Accounting & Salience, the full study-and-citation write-up for each is on the Principles page.
Experiment blueprintDigital & Product Experiment
Retail Banking · Credit & LendingAlso applies to: BNPL, personal loans, any business selling more than one way to finance the same purchaseNot yet run
Does tradeoff transparency work between products, not just within one?
CBA proved this comparing credit cards to each other. Does it hold comparing credit to other ways to pay?
Builds on Tradeoff Transparency, the real Buell & Choi (2025) field experiment run inside Commonwealth Bank of Australia's credit card funnel
Theory
The study. Buell and Choi ran a real field experiment inside CBA's credit card acquisition funnel: showing each card's drawbacks as prominently as its perks, a “Good and the Bad” page, instead of benefits-only marketing, left take-up statistically unchanged, but improved which card people picked. Customers who saw the downsides went on to spend 9.9% more, cancel 20.5% less, and miss fewer payments. The full teardown, with the real numbers, is on this site's Tradeoff Transparency page.
The gap. That experiment tested tradeoff transparency within one product category: comparing credit cards to each other. It didn't test the comparison one level up: credit card against the other ways the same customer could finance the same purchase, like a personal loan or a Buy Now, Pay Later plan, today sold on entirely separate pages, by entirely separate teams, with no drawbacks ever shown against each other at all.
Hypothesis
Showing a customer's actual financing options, credit card, personal loan, BNPL, side by side, with each one's real drawbacks named as plainly as its benefits, will shift some customers towards the option that actually fits how they intend to use it. Take-up itself shouldn't move: CBA's within-category version changed which card people picked without changing how many applied, and this is that same mechanism working one level up.
Assumed baseline
Illustrative, not disclosed CBA data. Assume 8% of applicants who land in today's credit-card-only funnel end up choosing a personal loan or BNPL plan instead, entirely on their own initiative. The funnel never surfaces either as an alternative, so this only happens when someone already knew to go looking for one.
Ethical guardrail
The line. Making a financing alternative visible for the first time isn't a neutral change, so the line is worth stating before the design, not after. This isn't built to talk anyone out of a credit card. For plenty of customers, revolving credit is the right tool.
The restraint. The design only surfaces the other real options before the commitment: inform the decision, don't engineer it, exactly the restraint the original CBA study held to. See Not Testing Is Still a Bet for the fuller case for that line.
Control vs. treatment
Control · one product at a time
Low Rate Card
12.99% p.a., no annual fee. No rewards, no interest-free days on cash advances.
Treatment · compare how you'd actually pay
Three ways to finance this
Credit Card: goodReusable, ongoing credit line
Credit Card: badInterest compounds if not paid in full
Personal Loan: goodFixed end date, usually lower rate
Personal Loan: badCan't redraw once it's repaid
BNPL: goodNo interest if paid on time
BNPL: badLate fees stack fast, builds no credit history
The card stays the first, largest option in both arms. The treatment only adds the comparison, it doesn't hide or de-rank anything.
Sample size to actually trust this
Baseline8%choose loan/BNPL unprompted
Target14%+6pp, ~75% relative
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm423applicants
Total sample8462 arms
Assumed volume2,400/monew financing applicants
Est. run time~3 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption: swap in a real monthly applicant count and the timeline follows.
Outcome variables
✓Primary: among applicants who switch to a loan or BNPL plan, missed-payment or default rate in the first 6 months, benchmarked against that product's normal book. The actual fit test: does an informed switch perform better, or just happen more often
✓Secondary: % of applicants selecting a personal loan or BNPL plan instead of a credit card. The volume shift the primary outcome needs to be interpreted against
✓Secondary: among applicants who still choose the card, average interest paid in the first 6 months: does staying on-card get cheaper too when the choice was more informed, mirroring CBA's own within-category result
✓Guardrail: overall financing conversion rate across all three products combined. Must not drop as a side effect of showing the comparison
✓Guardrail: support-contact rate referencing the comparison screen, as an early read on confusion
The mechanism behind this:Tradeoff Transparency, the full CBA study-and-citation write-up, including the real result table, is on the Principles page.
Experiment blueprintDigital & Product Experiment
Financial Services · Service & SupportAlso applies to: telco, airlines, delivery/logistics, any service business handling failures and complaintsNot yet run
Does a costly, specific apology beat a generic one after a service failure?
List found a costly apology paid for itself once, and backfired the moment it repeated. How do you design for both?
The study. Halperin, Ho, List and Muir ran this at real scale on 1.5 million Uber riders who'd just had a late ride: a $5 coupon apology paid for itself, restoring future spend that a generic “sorry” alone didn't: a costly signal doing work a free one can't, the mechanism this site documents as Reciprocity.
The sharper finding. Their paper also found a result worth designing around from the start: sending the identical coupon again after a second bad ride backfired. Riders spent less than if Uber had said nothing at all. The authors' own explanation: an apology reads as a promise, and reusing the exact same one a second time reads as that promise already broken. The full teardown, with the real numbers, is in the Voltage Effect field session.
Hypothesis
After a service failure, a costly and specific apology, a fee credit plus a plain explanation of what happened, will outperform a generic scripted apology on post-incident satisfaction. Because the same research found that reusing an apology verbatim can backfire on a repeat failure, the design should also test, as a secondary branch rather than the headline result, whether varying the apology's content and framing protects that advantage where a fixed script wouldn't.
Assumed baseline
Illustrative, not real support-desk data. Assume post-incident CSAT (1–5) currently averages 3.1 after a generic apology message, standard deviation 1.1, typical for a scripted, one-size response with no real cost attached.
Control vs. treatment
Both arms fire at the same moment, right after a qualifying failure is logged. Only the cost and the specificity of the response change.
Control · generic script
Sorry about that
We're sorry for the inconvenience. Thanks for your patience.
Treatment · costly & specific
Your transfer arrived 40 minutes late
A processing delay on our end. We've refunded the $3 fee and added a $5 credit. Sorry for the wait.
Sample size to actually trust this
Baseline mean3.1CSAT, 1–5 scale
Target mean3.4+0.3, ~10% relative
Assumed SD1.1illustrative
Significanceα = 0.05two-sided
Sample per arm212incidents
Total sample4242 arms
Assumed volume150/wkqualifying incidents
Est. run time~3 wksincl. buffer
Two-sample means test, equal allocation: n = 2 × (zα/2 + zβ)² × σ² / (μ₁−μ₂)², power 80%. This covers the headline test only: the repeat-failure sub-branch below draws from a smaller slice of this same sample, not an additional recruited group.
Outcome variables
✓Primary: post-incident CSAT (1–5), collected immediately after resolution
✓Secondary: 90-day retention for affected customers, matching the real spend-based outcome List's team used
✓Guardrail: cost per incident (average credit issued) weighed against retained-revenue estimate. The whole point of List's finding is that this pays for itself
✓Guardrail: support re-contact rate: the test of whether a specific explanation actually resolves the question, rather than only sounding reassuring
Sub-branch: if the failure repeats
Whichever customers in the headline sample have a second qualifying failure within the test window fall into this branch automatically. It isn't a separately recruited group, and it isn't the result the experiment is powered to confirm. Because who gets a second failure isn't random, treat this cut as exploratory: a read worth having, not a policy to ship on its own.
Sub-branch: same script again vs. varied
Control · same script again
Sorry about that
We're sorry for the inconvenience. Thanks for your patience.
Treatment · varied, not repeated
This shouldn't have happened twice
We've escalated the underlying issue to our transfers team, and changed how it's handled going forward.
The mechanism under test here isn't whether to apologise again. It's whether the repeat apology reads as a new response to a new problem, or the same promise, broken again.
Outcome for this branch: the same post-incident CSAT measure as the headline test, read only within this subset, plus, as a directional check rather than a powered comparison, whether “varied” avoids the spend-drop the Uber study found when the identical coupon repeated.
Retail Banking · Credit & LendingAlso applies to: insurance, wealth management, business banking, any multi-channel business launching one standout digital capability for the first timeNot yet run
Does a great digital home loan lift home loan trust everywhere else, too?
Why would a customer who never opens the new digital home loan still be more likely to sign one at a branch?
Builds on Halo Effect · Thorndike (1920), Journal of Applied Psychology
Theory
The study. Thorndike's finding, in one sentence: a rater who forms a strong impression on one visible trait unconsciously extends that same judgment to other traits of the same target they haven't actually evaluated, see Halo Effect for the full study. Thorndike, and later Nisbett & Wilson, tested that at the level of a single person's traits: one instructor's warmth colouring judgment of his unchanged accent and mannerisms.
The gap. This experiment tests the identical mechanism one level up: a single brand's channels, not a single person's traits. The bank in this scenario has never offered a digital home loan before; every application has gone through a broker or a branch appointment. It's about to launch a fully digital application: pre-approval in an afternoon, no paperwork mailed anywhere. Nothing else about the bank changes on launch day: the branch process, the broker relationships, the underwriting, the call centre, all identical to the day before.
The prediction. If the halo effect operates here the way it operates on a single rated person, simply being aware the bank now has one excellent digital channel should lift trust, and further lift application intent, in the channels that didn't change at all.
Hypothesis
The hypothesis. Customers exposed to messaging about the bank's new digital home loan capability, regardless of whether they ever open it, will show higher intent to get a home loan with the bank through any channel, digital or not, than a matched group with no exposure. That's the headline, general claim: mere awareness of one excellent channel lifts trust in the whole relationship.
What's separate. Whether customers who go on to actually use the digital channel show an even larger lift is tested separately, as a secondary, non-randomised read (see Sub-branch), not folded into this headline number.
Assumed baseline
Illustrative, not real/disclosed data. Assume 6% of existing customers who are in the market for a home loan in a given quarter say they'd seriously consider getting it with this bank, through any channel, before the digital product launches or any awareness messaging goes out.
Ethical guardrail
The risk. A halo is, by definition, trust extended to something nobody actually checked. That's the whole risk here: if awareness of the digital channel lifts intent to apply through the branch or a broker, and those channels haven't been resourced to match the promise the digital product makes, the bank ends up converting people on a trust it hasn't earned in the channel they actually use.
The guardrail. This experiment is not licence to under-invest in non-digital application support because the digital halo is doing the persuading, see Not Testing Is Still a Bet. The guardrails below exist specifically to catch that failure mode before it reaches a customer.
Control vs. treatment
Both arms see the same everyday banking app. Only the presence of a banner about the new digital home loan capability changes, who taps it, or applies through any channel afterwards, is observed, not assigned.
Control · no exposure
Accounts
Everyday Account · $2,140.50 Savings · $8,760.00
Treatment · aware of digital home loan
Get pre-approved this afternoon
Our new home loan application: entirely digital, no paperwork mailed anywhere.
Sample size to actually trust this
Baseline6%channel-agnostic consideration
Target9%+3pp, ~50% relative
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm1,205customers
Total sample2,4102 arms
Assumed volume500/wkeligible customers entering test
Est. run time~5 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². This covers the headline aware-vs-not comparison only: the sub-branch below draws from a smaller, self-selected slice of the treatment arm and isn't independently powered.
Outcome variables
✓Primary: % who apply for, or seriously consider, a home loan with the bank via any channel within 6 months of the awareness window, channel-agnostic conversion, the actual halo test
✓Secondary: the same metric split by channel of application (digital vs. branch/broker/phone), the lift showing up in channels that never changed is the halo signature; a lift confined to the digital channel alone is not
✓Secondary: unprompted, survey-measured trust/consideration score for a home loan specifically with this bank, independent of whether an application was ever started
✓Guardrail: branch/broker application abandonment rate and time-to-decision for the treatment group: a customer drawn in by the digital halo must not land in a slower, un-improved process than the impression promised
✓Guardrail: complaint rate and post-application CSAT for non-digital applicants in the treatment group, compared to control, catches the halo luring people into a channel that then under-delivers
Sub-branch: does actually using it deepen the halo further?
Whichever treatment-arm customers go on to actually start or complete a digital home loan application fall into this branch automatically. They aren't randomly assigned to “use it,” they choose to, so this comparison is self-selected and exploratory, not a third randomised arm. Treat it as a directional read on dose-response, not a number to plan a rollout around.
Sub-branch: saw the banner vs. actually applied digitally
Aware · saw the banner, didn't apply
Get pre-approved this afternoon
Seen, not opened. The customer never starts the digital application.
Used · started or completed a digital application
Pre-approval in progress
Direct experience of the excellent trait, not just awareness of it.
The mechanism under test here is the same one Thorndike's raters showed for a single trait: exposure alone produced some halo, but direct, first-hand experience of the excellent trait produced more.
Outcome for this branch: the same channel-agnostic home loan intent measure as the headline test, read only within the treatment arm, comparing those who engaged with the digital product against those who were merely exposed to it.
The mechanism behind this:Halo Effect, Thorndike (1920) and Nisbett & Wilson (1977), full study and citations on the Principles page.
Experiment blueprintDigital & Product Experiment
Retail Banking · Credit & LendingAlso applies to: personal loans, credit cards, any product priced with a percentage rate customers compare at a glanceNot yet run
Does a home loan rate of 5.99% get remembered as a 5% loan, not almost 6%?
A 5.99% home loan rate is really almost 6%. Does it get compared like a 5% one instead?
Builds on Left Digit Bias · Lacetera, Pope & Sydnor (2012), American Economic Review
Theory
The study. Lacetera, Pope and Sydnor's finding, in one sentence: a number's leftmost digit dominates how its magnitude gets judged, so a real difference that crosses a leading-digit boundary, 39,999 miles to 40,000, $2.99 to $3.00, produces a disproportionately large shift in perceived value, even though the actual gap is tiny. Their evidence is 22 million real used-car sales; Thomas & Morwitz's lab work found the identical pattern in ordinary prices.
The gap. Neither tested an interest rate. This experiment moves the same mechanism onto a home loan's headline number. A rate advertised at 5.99% p.a. and one advertised at 6.00% p.a. cost a borrower almost exactly the same in real repayments, but the leading digit differs: 5.99% reads as “5% something,” while 6.00% reads as “6%” outright.
The prediction. If the mechanism holds here the way it holds for a car's odometer, 5.99% should be judged, and chosen, as meaningfully cheaper than 6.00%, disproportionate to the real cost difference between them.
Hypothesis
A home loan advertised at 5.99% p.a. will be perceived as, and will convert to more started applications than, an otherwise identical loan advertised at 6.00% p.a., even though the two rates are set close enough that the real difference in monthly repayment is trivial.
Assumed baseline
Illustrative, not real/disclosed data. Assume 4% of visitors to the rate page who are shown a 6.00% p.a. headline rate go on to start an application within the same session.
Ethical guardrail
Already standard. This experiment doesn't invent a new practice. .99 rate framing is already standard, legal advertising across the lending industry, and the real cost stays disclosed via the comparison rate regardless of which headline number is shown.
What's being tested. How much that convention alone moves behaviour, so the bank can decide, with real numbers, whether to keep leaning on it as-is, cap how far it leans, or counter it by surfacing the comparison rate more prominently. See Not Testing Is Still a Bet for why measuring this honestly beats assuming it's fine because everyone already does it.
Control vs. treatment
The comparison rate, the standardised, all-costs-included figure regulators require, is identical in both arms. Only the headline rate's leading digit changes.
Control · 6.00% p.a.
6.00% p.a.
Variable rate home loan. Comparison rate 6.04% p.a.
Treatment · 5.99% p.a.
5.99% p.a.
Variable rate home loan. Comparison rate 6.04% p.a.
Same loan, same real cost, same comparison rate. The only thing that moves is which side of the 6% boundary the headline number sits on.
Sample size to actually trust this
Baseline4%started an application
Target5.5%+1.5pp, ~38% relative
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm3,150rate-page visitors
Total sample6,3002 arms
Assumed volume900/wkeligible visitors entering test
✓Primary: % of rate-page visitors who start a home loan application within the same session
✓Secondary: intercept survey asking visitors to rate the offer as cheap / about right / expensive, checking whether the perception itself shifts, not just the click
✓Guardrail: application-to-approval drop-off once the comparison rate is unavoidably shown. The headline number must not be pulling in applicants who bail once they see the real cost
✓Guardrail: complaint or dispute rate referencing the rate not matching expectations, a spike here would mean the framing misled people beyond the ordinary digit effect
The mechanism behind this:Left Digit Bias, Lacetera, Pope & Sydnor (2012) and Thomas & Morwitz (2005), full study and citations on the Principles page.
Experiment blueprintDigital & Product Experiment
Retail Banking · Credit & LendingAlso applies to: personal loans, car loans, any credit product weighing a prize-linked promotion on top of an already competitive rateNot yet run
Does a chance at 0% increase home loan applications more than the guaranteed discount alone?
A genuinely good guaranteed rate already exists. Does bolting on a small, real chance at 0% for a year pull in more applicants than the guaranteed rate alone?
Builds on Probability Weighting · Kahneman & Tversky (1979), Econometrica · Rottenstreich & Hsee (2001), Psychological Science · Burger, Hendriks, Pleeging & van Ours (2020), Experimental Economics
Theory
The studies. Bankwest's real Interesting Rates promotion pays every Easy Saver customer a genuinely competitive 5.75% p.a. everyday rate, then draws 50 winners a month for an 11.50% p.a. rate on top. Probability Weighting is the direct mechanism: a small, real chance of the better rate gets weighted in people's minds far past its true odds.
Two further findings shape how hard that weighting bites. Rottenstreich and Hsee (2001) found people overweight a small chance of something emotionally vivid even more than an equally small chance of a plain cash amount of the same value. That's exactly why a fast-food deal advertises “free food for life” instead of its cash value. A year of 0% interest on a real mortgage is at least as vivid a prize as free food, arguably more so: it converts directly into a large, concrete, easy-to-picture sum. Burger et al. (2020) found a second, separate effect: simply holding a lottery-style entry measurably lifts happiness before the draw resolves, win or not. A home loan application is normally a stressful, multi-week wait. A live entry gives an applicant something hopeful to hold onto during it, a genuinely different benefit from the conversion lift Probability Weighting alone would predict.
The gap. Classic prize-linked savings research, like Filiz-Ozbay, Guryan, Hyndman, Kearney and Ozbay's 2015 study, tests a real substitution: part of a saver's own guaranteed return is swapped for lottery expected value. Bankwest's promotion doesn't do that; its 5.75% rate exists whether or not anyone enters the draw. A home loan version pairing an already-competitive guaranteed discount with a separate chance at 0% sits further still from that substitution: no part of the applicant's real rate funds the prize. This tests Probability Weighting's own pull, not the prize-linked-savings trade-off.
This experiment. It tests whether advertising a genuinely competitive guaranteed discount alongside a separate, real chance at 0% p.a. for 12 months increases application starts on a new home loan rate page, against advertising the identical guaranteed discount alone, with that guaranteed rate held exactly constant between arms. It also checks a second, distinct prediction: that simply holding a live entry changes how applicants feel during the underwriting wait, not just whether they start applying.
Hypothesis
Visitors shown the guaranteed discount plus a chance at 0% p.a. for 12 months start more applications within 7 days than visitors shown the identical guaranteed discount with no draw attached, even though the guaranteed rate on offer never changes between arms. Independent of that conversion lift, applicants in the treatment arm report less stress and higher satisfaction during the underwriting wait than applicants in the control arm, the same anticipatory-happiness effect Burger et al. found in lottery ticket-holders, whether or not they end up winning the draw.
Assumed baseline
Illustrative, not real/disclosed data. Assume 8% of visitors to the new home loan rate page start an application within 7 days when the page advertises the guaranteed discount alone.
Ethical guardrail
The guardrail. The guaranteed discount has to be genuinely competitive against the rate the bank would otherwise offer that same applicant, not a rate quietly held back elsewhere to help fund the draw, and it must stay identical in both arms so the draw is the only variable under test. The odds, the number of winners, and the exact prize, 0% p.a. for 12 months on a capped loan amount, need disclosing as plainly as the guaranteed rate itself.
What's being tested. Whether a real, cited effect helps a genuinely competitive offer reach more of the people it's meant for, not whether a draw's excitement can be tuned to pull applicants into borrowing more than they can service. Loan size and serviceability assessment run identically in both arms, and a guardrail metric below checks that directly. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms are new visitors to the home loan rate page, randomly assigned. The guaranteed discount rate is identical in both arms; only whether the page also advertises a chance at 0% differs.
Control · guaranteed discount only
3.00% p.a. fixed, first year
A genuinely competitive fixed rate for your first 12 months, reverting to our standard variable rate after.
Treatment · guaranteed discount + prize draw
3.00% p.a. fixed, first year
The same fixed rate as always. Every eligible application this month is entered into a draw for 0% p.a. on that first year instead.
Same guaranteed rate either way: the draw is a separate chance at 0%, not a discount taken off the 3.00% on offer.
Same guaranteed rate, same eligibility, same application form in both arms. Only whether the page also advertises the draw differs.
Sample size to actually trust this
Baseline8%started an application, 7 days
Target10.5%+2.5pp assumed lift
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm2,102rate-page visitors
Total sample4,2042 arms
Assumed volume3,000/moeligible rate-page visitors
Est. run time~6 wksto enrol both arms
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: application-start rate within 7 days of viewing the home loan rate page, treatment versus control
✓Secondary: application completion rate among those who start, confirming the draw pulls in applicants who finish, not just curious clicks
✓Secondary: applicant-reported stress and satisfaction during the underwriting wait, a short post-decision survey testing whether simply holding a live entry lifts mood independent of the draw's outcome
✓Guardrail: share of started applications later reduced or declined at the serviceability-assessment stage must not rise in the treatment arm
✓Guardrail: support or complaint contacts referencing confusion about the draw's odds or eligibility must stay within the same range as ordinary rate-page queries
The mechanism behind this:Probability Weighting, Kahneman & Tversky (1979), together with the affect-rich-prize and anticipatory-happiness findings cited on that same page (Rottenstreich & Hsee, 2001; Burger et al., 2020). See also The Lottery You Can't Lose for the classic prize-linked-savings trade-off this experiment deliberately doesn't reproduce, and Why do you have to be “into double denim” to win a better savings rate? for the real campaign this concept extends into lending.
Experiment blueprintDigital & Product Experiment
Retail Banking · Fees & PricingAlso applies to: insurance, telecom, subscription and utility pricing, any recurring price a real business has to raiseNot yet run
Does explaining why a fee went up change whether customers think it's fair?
Kahneman, Knetsch and Thaler proved this with a hypothetical snow shovel. Does it hold for a real bank fee notice?
Builds on Perceived Fairness · Kahneman, Knetsch & Thaler (1986), American Economic Review
Theory
The study. Kahneman, Knetsch and Thaler's finding, in one sentence: an identical price increase is judged fair when it's framed as protecting the seller's existing profit from a real cost increase, and unfair when it looks like exploiting a shift in demand or bargaining power. The number itself doesn't decide the verdict, the story behind it does. Their evidence is a telephone survey of Toronto and Vancouver residents judging realistic vignettes, including a hardware store that raised snow-shovel prices the morning after a storm (rated unfair by roughly 82% of respondents) and a landlord charging existing versus new tenants differently for the same unit.
The gap. Those vignettes measured stated opinions about a hypothetical price change, not real behaviour in response to a real one, and none of them tested a financial-services fee or rate notice specifically, where the cost being passed on (funding costs, compliance costs) is far less visible to a customer than a snowstorm is.
The prediction. If the same dual-entitlement logic holds here, a bank that names the real cost driver behind a fee or rate increase should see the change land as fairer, and, unlike KKT's survey, this experiment can check whether that shows up in real behaviour, not just a nicer opinion: fewer support contacts, fewer customers leaving.
Hypothesis
Customers who receive a fee or interest-rate increase notice that names a real, verifiable cost driver will be less likely to contact support to question or dispute the change, and more likely to rate it as fair, than customers who receive an otherwise identical notice stating only the new number, even though the actual new fee or rate is unchanged between the two groups.
Assumed baseline
Illustrative, not real/disclosed data. Assume 18% of customers who receive a bare fee-increase notice, new amount only, no explanation, contact support to question or dispute the change within 30 days.
Ethical guardrail
The guardrail. This mechanism only works honestly when the disclosed cost driver is real and verifiable, a genuine rise in funding costs, compliance costs, or similar, that the bank could substantiate if asked. Citing a vague or invented cost story to manufacture a feeling of fairness would be deceptive framing, not disclosure, and is explicitly out of scope for this experiment.
What's being tested. Only whether naming a true cost driver changes how a genuine, already-decided fee change lands, not whether that framing can be used to justify an increase that wouldn't otherwise be defensible on its own. See Not Testing Is Still a Bet for why measuring this honestly matters more than assuming good framing is harmless.
Control vs. treatment
The new fee and its effective date are identical in both arms. Only the presence of a real, named cost driver changes.
Control · number only
Your account fee is changing
From 1 October, your monthly account fee will be $12, up from $8.
Treatment · number + real reason
Your account fee is changing
From 1 October, your monthly account fee will be $12, up from $8.
Why this is changing: our cost of providing fee-free access across the shared ATM network has risen. This new fee reflects that increase, nothing else about your account is changing.
Same new fee, same effective date. The only difference is whether a real, named cost sits behind the number or not.
Sample size to actually trust this
Baseline18%contacted support about the change
Target13%−5pp, ~28% relative
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm820customers receiving the notice
Total sample1,6402 arms
Assumed volume300/wkcustomers receiving a fee or rate-change notice
✓Primary: % of affected customers who contact support to question or dispute the fee change within 30 days of the notice
✓Secondary: a one-question follow-up survey to a subsample, asking how fair the change felt (1–5), checking whether the perception itself shifts, not just the support-contact behaviour
✓Secondary: 60-day account or product retention rate among affected customers, versus a matched unaffected group
✓Guardrail: rate of complaints escalated to a formal dispute or ombudsman referral, naming a reason must not itself invite more scrutiny than a bare notice would have
The mechanism behind this:Perceived Fairness (Dual Entitlement), Kahneman, Knetsch & Thaler (1986), full study and citations on the Principles page.
Experiment blueprintDigital & Product Experiment
Retail Banking · Savings & DepositsAlso applies to: budgeting apps, neobanks, any card issuer with an onboarding flow and in-app messagingNot yet run
Does telling customers the payment-transparency heuristic change what they actually spend?
Prelec and Simester proved credit blunts the felt cost of paying. Does telling a customer that, at the right moment, blunt it back?
The mechanism. Payment Transparency says a purchase's felt cost depends on whether the payment method makes you rehearse the amount and whether the money leaves immediately. The Credit Card Premium is the sharpest real case of it: bidders told they'd pay by credit bid close to double what cash-instructed bidders bid, for the identical tickets, because credit strips out both cues.
Already suggested, not yet tested. Both principle articles' own “Where it can help” sections argue the same two cues can be put back deliberately, a running balance shown at the moment of tap, a receipt that totals spend by category, but neither study tested whether simply telling someone the mechanism, in plain language, changes what they actually spend. That's a different, weaker intervention than redesigning the interface itself: it relies on a customer remembering and acting on a fact, not a changed screen.
This experiment. It tests whether education alone, no interface change, no restricted product, moves real spending, and whether it matters when that education arrives: repeated, close to the moment of use, or delivered once, before any spending has happened at all.
Hypothesis
Customers told the payment-transparency heuristic in plain language spend less on their lowest-transparency linked payment method (a stored or linked credit card) over the following months than customers told nothing. The reduction is larger and more durable for customers who receive the heuristic repeatedly, as a contextual in-app message near the moment of spending, than for customers who receive it once, as onboarding content before they've made a single transaction, because the underlying mechanism, rehearsal and immediacy, is about what's salient at the moment of paying, not what was once explained in the abstract.
Assumed baseline
Illustrative, not real/disclosed data. Assume customers in the control arm spend an average of $780 a month on their lowest-transparency linked payment method in the 90 days after opening a new debit account, with a standard deviation of $360 reflecting how unevenly card spending is actually distributed across a customer base.
Ethical guardrail
The guardrail. The message has to state the real mechanism honestly, something like “paying by credit doesn't feel like spending right away, which is part of why people spend more on it,” not a shame-based or fear-based framing, and it has to apply the same way regardless of a customer's income or existing balance, so it reads as financial literacy, not a judgment about who's trusted to spend.
What's being tested. Whether informing someone of a real, cited mechanism changes their own behaviour once they know it, not whether the message can be tuned to manipulate spending in either direction. See Not Testing Is Still a Bet for why measuring this honestly matters more than assuming a well-meaning message is automatically harmless.
Control vs. two treatments
All three arms are new debit account holders in their first 90 days, randomly assigned at account opening. Only the presence, content, and timing of the heuristic message differ.
Control · no messaging
Your account
Standard account home screen. No spend-heuristic messaging is shown at any point.
Treatment A · in-app, contextual
You spent $64 on your credit card today
Paying by credit doesn't feel like spending right away, that's part of why credit purchases add up faster than debit or cash.
Try this: check your running balance before your next purchase.
Treatment B · onboarding, once
Before you start: how you pay changes what you spend
Research shows people tend to spend least with cash, more with debit, and most with credit, because credit doesn't feel like spending right away.
One-time note: you'll see this once, during setup. It won't be repeated.
Same account, same product, same starting balance. Only whether the heuristic is ever shown, and how often, differs between arms.
Sample size to actually trust this
Baseline$780avg. monthly credit spend, control
Target$730−$50, the smaller expected effect (Treatment B)
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm815new debit account holders
Total sample2,4453 arms
Assumed volume900/monew debit account sign-ups
Est. run time~3 moto enrol, plus each customer's own 90-day window
Two-sample continuous-outcome test, equal allocation: n = 2 × (zα/2 + zβ)² × σ² / (μ₁−μ₂)². Sized against Treatment B's smaller expected effect, so Treatment A, expected to move spending further, is over-powered by comparison.
Outcome variables
✓Primary: average monthly spend on the customer's lowest-transparency linked payment method over the 90 days after enrolment, each treatment versus control
✓Secondary: share of total spend that shifts from credit to debit or cash over the same 90 days, whether the mix moves, not just the credit total
✓Secondary: message engagement: tap-through rate on Treatment A's in-app card, and completion rate on Treatment B's onboarding screen, confirming both arms actually saw the message before comparing outcomes
✓Guardrail: rate of missed or late credit card repayments among treated customers must not rise, the goal is awareness, not anxiety that backfires into worse account management
✓Guardrail: 90-day app engagement and account retention must not fall relative to control, the messaging should inform, not annoy people into disengaging
Retail Banking · Savings & DepositsAlso applies to: telco, insurance, and subscription apps with a multi-step onboarding or setup flowNot yet run
Does framing account setup as one set to complete finish more setups?
Barasz, John, Keenan, and Norton proved an arbitrary group of items feels incomplete and pulls people to finish it. Does drawing that same line around a bank's own setup steps get more customers to actually finish setting up?
Builds on Pseudo-Set Framing · Barasz, John, Keenan & Norton (2017), Journal of Experimental Psychology: General
Theory
The study. Pseudo-Set Framing showed that grouping arbitrary items into a visible "set" makes people push to complete it, even when the reward doesn't change, even when finishing costs something, and even when people are told the grouping is made up. The effect held across five separate studies, gambling, effort, giving, and purchasing, not just one narrow task.
The gap. None of the paper's studies tested a real product's own multi-step setup flow, the exact place a bank, telco, or insurer already asks customers to complete several small, genuinely useful tasks (verify ID, link a funding source, add a beneficiary, set a goal) with no shared visual frame tying them together.
This experiment. It tests whether picturing those same required steps as wedges of one circle to complete, rather than an ordinary checklist, increases how many customers finish all of them, without changing what's actually required or adding a single extra step.
Hypothesis
New customers shown their remaining setup steps as wedges of one "complete your profile" circle finish all required steps within 14 days at a higher rate than customers shown the identical steps as a plain checklist, even though the number of steps, the effort each takes, and the reward for finishing are all unchanged between arms.
Assumed baseline
Illustrative, not real/disclosed data. Assume 41% of new customers complete all 4 required setup steps (ID verification, linked funding source, a named beneficiary, a savings goal) within 14 days of opening an account under the existing plain-checklist flow.
Ethical guardrail
The guardrail. Every wedge in the circle has to correspond to a step that's genuinely required or genuinely useful to the customer, none added purely to round the set out or make the shape more compelling. The circle can't be blocked or nagged into completion with dark patterns, delayed access to core banking functions, repeated interstitials, that go beyond what the ordinary checklist arm also uses.
What's being tested. Whether a real, cited completion mechanism helps customers finish setup steps they already need to do, not whether the visual can be tuned to pressure people into rushing through steps they'd otherwise have skipped for a good reason. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms are new account holders in their first setup session, randomly assigned at account opening. Only how the same 4 remaining steps are pictured differs; the steps, their order, and the reward for finishing are identical.
Control · plain checklist
Finish setting up
1 of 4 steps done: ID verified. Link a funding source, add a beneficiary, and set a savings goal to finish.
Treatment · one circle, four wedges
Your profile: 1 of 4 pieces in place
The same 4 steps, pictured as wedges of one circle. ID verified. Link a funding source, add a beneficiary, and set a savings goal to complete it.
Same steps, same reward: nothing required here changes, only how it's pictured.
Same account, same 4 requirements, same completion reward. Only whether the remaining steps are pictured as a checklist or as wedges of one circle differs between arms.
Sample size to actually trust this
Baseline41%complete all 4 steps, control
Target51%+10pt assumed lift, treatment
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm386new account holders
Total sample7722 arms
Assumed volume1,200/monew account sign-ups
Est. run time~1 moto enrol, plus each customer's 14-day window
Two-sample proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². At 1,200 new accounts a month split evenly, 772 total is reached inside roughly a month of enrolment.
Outcome variables
✓Primary: share of new customers who complete all 4 required setup steps within 14 days of account opening, treatment versus control
✓Secondary: median time to completion among customers who do finish, whether the pull to close the set also speeds up completion, not just raises the rate
✓Secondary: view rate on the profile/checklist screen itself, confirming customers in both arms actually saw their remaining steps before comparing completion
✓Guardrail: rate of steps completed with invalid or placeholder data (a fake beneficiary, an unreachable funding source) must not rise in the treatment arm, the goal is real completion, not gaming the shape
✓Guardrail: setup-related support contacts must not increase relative to control, the visual should clarify what's left, not create confusion or pressure
The mechanism behind this:Pseudo-Set Framing, full study and citation on the Principles page. See also Goal Gradient Effect, the related but distinct pull of getting objectively closer to a finish line, versus this experiment's test of a group simply looking unfinished.
Experiment blueprintDigital & Product Experiment
Retail Banking · Savings & DepositsAlso applies to: neobanks, fintech savings apps, any product paying a small variable bonus rateNot yet run
Does a prize-linked feature beat an equal-value rate bonus for new deposits?
A lab study found people defer more for a lottery-style payout than a guaranteed one of identical expected value. Does the same swap move real deposits in a real banking app?
Builds on The Lottery You Can't Lose · Filiz-Ozbay, Guryan, Hyndman, Kearney & Ozbay (2015), Journal of Public Economics; Kearney, Tufano, Guryan & Hurst (2010), NBER
Theory
The studies. A controlled lab experiment found participants offered a lottery-style bonus deferred payment more often than participants offered a guaranteed bonus of the identical expected value, with the pull strongest among low-balance participants. Michigan's real-world Save to Win programme moved $8.56 million into new accounts in its first year using the same structure.
The gap. Neither study is a live digital-product test: the lab experiment paid a one-off lab bonus, and Save to Win is a whole account type, not a feature layered onto an existing one. No test yet isolates the payout-structure swap alone, inside one already-existing digital savings product, with everything else held constant.
This experiment. It tests whether replacing part of an existing savings feature's guaranteed bonus rate with a prize-draw of equal expected value, inside the same app, for the same customers, increases net new deposits, without changing anything else about the product.
Hypothesis
Customers offered a prize-draw bonus structure, funded from the identical pool that would otherwise pay a guaranteed rate bump, deposit more into the dedicated savings feature over 60 days than customers offered the guaranteed rate bump directly, even though the total expected payout is the same in both arms. The gap is largest among customers with the lowest starting balances, matching the lab study's own finding.
Assumed baseline
Illustrative, not real/disclosed data. Assume customers in the guaranteed-rate control arm deposit an average of $310 in net new funds into the dedicated savings feature over 60 days, with a standard deviation of $410 reflecting how unevenly savings deposits are actually distributed.
Ethical guardrail
The guardrail. The prize pool must be funded from the exact amount that would otherwise be paid as guaranteed interest, never as an added cost that makes the product worse in expectation, and the deposit itself must never be at risk regardless of the draw's outcome. Odds and expected value have to be disclosed as plainly as an interest rate would be, not buried under excitement-focused copy.
What's being tested. Whether a real, cited payout-structure effect helps customers save more of their own money, not whether the framing can be tuned to pressure the specific low-balance customers it appeals to most into depositing more than they can afford. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms are existing customers newly opting into a dedicated savings feature, randomly assigned at opt-in. The expected payout is identical in both arms; only whether it's paid as a guaranteed rate or a prize draw differs.
Control · guaranteed rate bonus
Savings Boost: +0.4% bonus rate
Deposits into this feature earn an extra 0.4% on top of your standard rate, paid monthly, guaranteed.
Treatment · prize-draw bonus
Savings Draw: monthly prize entries
Every $25 in this feature earns one entry into a monthly prize draw. Your deposit is never at risk, win or not.
Same value, different shape: the prize pool pays out the same total, on average, as the +0.4% rate.
Same underlying account, same expected payout, same funding pool. Only whether that payout arrives as a guaranteed rate or a prize draw differs between arms.
Sample size to actually trust this
Baseline$310avg. net new deposit, 60 days, control
Target$370+$60 assumed lift, treatment
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm733customers opting into the feature
Total sample1,4662 arms
Assumed volume1,000/monew feature opt-ins
Est. run time~2 moto enrol, plus each customer's 60-day window
Two-sample continuous-outcome test, equal allocation: n = 2 × (zα/2 + zβ)² × σ² / (μ₁−μ₂)². At 1,000 opt-ins a month split evenly, 1,466 total is reached in roughly 3 months of enrolment.
Outcome variables
✓Primary: net new deposits into the dedicated savings feature over 60 days, treatment versus control
✓Secondary: feature opt-in rate itself, confirming both arms activate the feature at similar rates before comparing what happens after
✓Secondary: balance retention in the 30 days after the first prize draw resolves, whether engagement fades for customers who didn't win
✓Guardrail: rate of overdrafts or missed payments elsewhere in the account must not rise among treatment customers, the feature shouldn't pull in money people can't actually spare
✓Guardrail: support contacts about how the draw works or whether a customer won must stay within the same range as ordinary account queries, confirming the mechanic was clearly understood
The research behind this:The Lottery You Can't Lose, full citations and the Save to Win field data. See also Loss Aversion, the asymmetry a no-loss lottery structure is specifically built to avoid triggering.
Experiment blueprintDigital & Product Experiment
Retail Banking · Credit & LendingAlso applies to: insurance quotes, account switching, any lead form long enough that people abandon partway throughNot yet run
Does showing progress through a home loan appointment booking increase completion?
A pre-filled loyalty card sped people up with the identical distance left to go. Does simply showing progress do that, or does how the remaining steps are chunked matter just as much?
Builds on Goal Gradient Effect · Kivetz, Urminsky & Zheng (2006), Journal of Marketing Research
Theory
The study. Kivetz, Urminsky and Zheng found effort and motivation increase the closer someone perceives themselves to be to a finish line, and that perceived proximity alone, not real distance, drives it: a 12-stamp card with 2 stamps already filled in accelerated purchasing identically to a genuinely shorter 10-stamp card, even though both needed exactly 10 more purchases. The principle's own “where it can help” case is a genuine progress bar showing someone how close they actually are to finishing a form.
The gap. The principle proposes this; nobody has tested it. A home loan appointment booking flow is exactly the kind of long, multi-step form (contact details, loan and property details, a preferred meeting type, a time slot) where abandonment is common and a completion cue has never been isolated as the only variable.
This experiment. It tests two different applications of that mechanism on the same flow: showing an honest step count at the flow's natural, coarse chunking, and re-chunking the identical real work into more, still-genuine steps so a customer sits further through the count at the same real point, the exact lever the original study's 12-stamp card used against a 10-stamp one.
Hypothesis
Visitors shown a real step-progress indicator, at either chunking, complete and confirm a booking at a higher rate than visitors shown the identical flow with no progress indicator, even though the fields required and the appointment options offered are unchanged in every arm. The study's mechanism runs on how close someone feels, not how far they actually have left, so splitting the same real work into more granular genuine steps, a customer sitting further through the count at an identical real point, produces a larger lift than an honest step count at the flow's natural, coarser chunking.
Assumed baseline
Illustrative, not real/disclosed data. Assume 34% of visitors who start the appointment booking flow complete all 4 steps and confirm a booking under the existing no-progress-indicator flow.
Ethical guardrail
The guardrail. The step count and progress bar have to reflect the real number of steps and real position in the flow at every screen, never compressed or re-ordered to make the remaining effort look smaller than it is. The total step count should be visible from the very first screen, not revealed one step at a time, so no one commits to step one without knowing there are three, or five, more behind it.
Re-chunking has a real/fake line. Splitting a step into finer sub-steps is only legitimate when each resulting step is still genuinely distinct, required content, contact details and identity verification are two real things, not one field split in half to manufacture an extra step. No step shown, at either chunking, may be fabricated, reordered to hide what's left, or padded purely to inflate the count.
What's being tested. Whether an honest completion cue helps people follow through on a booking they already want to make, not whether the count can be tuned to disguise how long the form actually takes. See Not Testing Is Still a Bet.
Control vs. two treatments
All three arms are visitors starting a home loan appointment booking flow, randomly assigned at the first screen. Every field and every appointment option offered is identical across all three; only how, or whether, progress is shown differs. Control and Treatment A are shown at the flow's natural step 2 of 4; Treatment B shows the same real work re-chunked into 6 genuine steps, generic and unbranded.
Control · no progress shown
Loan & property details
Estimated property value, deposit amount, and the purpose of the loan.
Treatment A · honest step count
Step 2 of 4About 3 min left
Loan & property details
Estimated property value, deposit amount, and the purpose of the loan.
Treatment B · re-chunked, 6 steps
Step 4 of 6About 2 min left
Loan amount & purpose
Deposit amount and the purpose of the loan. Property details are already confirmed.
All three arms require the exact same real information: contact details, identity verification, property value, deposit amount, loan purpose, meeting preference, and a chosen time. Only whether progress is shown, and how finely the same steps are chunked, differs.
Why Treatment B looks different. It doesn't change how much work is left, real or perceived honestly, it changes how many genuine pieces that work is split into. “Loan & property details” becomes its own property step and its own loan-amount step; no field is added, removed, or hidden. A customer who has done the identical real work now sits at step 4 of 6 instead of step 2 of 4. That's the exact lever the original study used: the same 10 purchases left to go read as 0% done on a blank 10-stamp card, or already under way on a 12-stamp card with 2 stamps pre-filled, identical real distance, different perceived proximity.
Sample size to actually trust this
Baseline34%completed booking, control
Target42%+8pp, the smaller effect (Treatment A)
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm574visitors starting the flow
Total sample1,7223 arms
Assumed volume900/wkvisitors entering the flow
Est. run time~3 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². Sized against Treatment A's smaller expected effect, so Treatment B, assumed to reach 48% (+14pp) if the re-chunking mechanism holds, is over-powered by comparison.
Outcome variables
✓Primary: % of visitors who start the flow and confirm an appointment, each treatment versus control
✓Secondary: the difference between Treatment A and Treatment B specifically, isolating whether re-chunking the same real work into more granular steps adds anything beyond simply showing an honest count at the flow's natural chunking
✓Secondary: step-by-step drop-off in each arm, whether the lift is concentrated near the end (the goal-gradient signature) or spread evenly across the flow
✓Secondary: median time to complete the flow, whether visible progress also speeds people up, not just raises who finishes
✓Guardrail: show-up rate for the booked appointment itself must not fall in either treatment arm, a rushed booking made just to finish a bar is worse than no booking at all
✓Guardrail: rate of visitors who abandon after seeing the total step count on screen one must not spike in either treatment, confirming upfront honesty about the length, 4 steps or 6, doesn't itself deter people
The mechanism behind this:Goal Gradient Effect, full study and citation on the Principles page. See also Pseudo-Set Framing, a related completion pull driven by how a group of steps is framed rather than by shrinking distance to a goal.
Experiment blueprintDigital & Product Experiment
Retail Banking · Savings & DepositsAlso applies to: super/retirement contribution apps, budgeting apps with an auto-save or round-up featureNot yet run
Does timing a savings increase to a raise beat asking for it right now?
Save More Tomorrow tied the increase to an employer's known pay schedule. Does the same commitment hold when the raise is only spotted in a bank deposit?
Builds on Smart Defaults · Thaler & Benartzi (2004), Journal of Political Economy
Theory
The study. Save More Tomorrow let employees commit today to a contribution-rate increase that wouldn't take effect until their next scheduled pay raise, so take-home pay never actually shrank on the day it changed. Of 162 employees who'd just declined an immediate increase, 78% accepted the deferred version instead, and average savings rates rose from 3.5% to 13.6% of pay over four raises, about 40 months, with 78% still enrolled at the end.
The gap. Every tested version of this relies on an employer who knows the exact date and size of the next raise in advance. A savings or banking app has no such schedule to work from: the best it can do is notice, after the fact, that a customer's regular salary deposit just got bigger.
The prediction. If the mechanism is really about timing the ask against a gain instead of a felt cut, the source of the raise shouldn't matter, scheduled or detected, only that the increase is committed to in advance and only takes effect once the extra money has already arrived.
Hypothesis
Customers offered a one-tap commitment to increase their auto-save rate at their next detected pay rise, instead of being asked to increase it immediately, accept some increase to their savings rate at a higher rate than customers asked to increase it right now, even though the amount of the increase on offer is identical in both arms.
Assumed baseline
Illustrative, not real product data. Assume 9% of customers shown an in-app prompt to increase their auto-save rate by 2% right now accept it.
Ethical guardrail
The risk. A jump in a bank deposit isn't always a raise. A one-off bonus, a tax refund, or a new but temporary higher-paying gig would all look identical to a real, ongoing pay rise the moment it lands, and treating any of them as permanent would auto-escalate someone's savings against income that's about to disappear.
The guardrail. A detected increase should hold across at least two consecutive pay cycles before it triggers the commitment, and any triggered increase must be reversible with one tap and refundable in full for a short window afterwards, not just cancellable going forward. Getting the timing wrong should cost the customer nothing.
What's being tested. Whether committing in advance, against a gain instead of a felt cut, is what drives acceptance, not whether raise-detection can be tuned to trigger more often than it should. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms are customers shown a prompt to increase their auto-save rate by the same 2%. The only thing that changes is when the increase takes effect and what it's measured against.
Control · increase now
Increase your auto-save
Move your auto-save rate from 5% to 7% of each pay, starting with your very next pay.
Treatment · increase at next raise
Increase it at your next raise
We'll watch for a rise in your regular pay and move your auto-save from 5% to 7% only once it lands. Your take-home stays the same until then.
The 2% increase on offer, and the account it applies to, is identical in both arms. Only the timing, and what the increase is measured against, a felt cut today or a gain that hasn't arrived yet, differs.
Sample size to actually trust this
Baseline9%accepted, increase now
Target15%+6pp, ~67% relative
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm456customers prompted
Total sample9122 arms
Assumed volume550/wkeligible customers prompted
Est. run time~2 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number. Swap in your own weekly eligible-customer count and the run time follows.
Outcome variables
✓Primary: % of prompted customers who accept some increase to their auto-save rate, control's immediate accept versus treatment's deferred commitment
✓Secondary: of treatment customers who commit, % whose increase actually triggers within 90 days, whether the commitment survives long enough for a real raise to be detected
✓Guardrail: % of triggered increases reversed or refunded within 14 days, catching a raise the detector got wrong
✓Guardrail: overall auto-save opt-out rate in the 30 days after the prompt, the ask itself must not drive people to switch saving off altogether
The mechanism behind this:Smart Defaults, full study and citation on the Principles page.
Three fronts, one mechanism
Reactance shows up at three different moments
When a choice gets taken away, when someone's asked to make one, and when access to one gets restricted. Three separate blueprints test each moment on its own, all built on the same enriched Not Enough Choice principle. Pick one below, or scroll through all three in order.
Removing a choice
Does explaining why a feature was removed reduce the backlash?
Not Enough Choice already claims transparency defuses reactance. This tests that claim directly, for the first time, against a real product change.
Does framing a tier as restricted increase upgrade interest?
The founding reactance study proved restriction increases desire. The sharpest of the three to get right, since it only stays honest if the restriction named is real.
The mechanism. Not Enough Choice found reactance is triggered specifically by an active removal, not by an option simply never being available, and its own “where it can help” section claims plainly that explaining the reason should defuse it.
The gap. That claim has never actually been tested on a real removal notice. Nobody has isolated the explanation itself as the only variable.
This experiment. Both arms get the identical removal, on the identical timeline, with the identical balance migration. Only whether a reason is given differs.
Hypothesis
Customers told why a feature is being removed contact support or lodge a complaint about it at a lower rate than customers told only that it's being removed, even though the outcome for their money is identical in both arms.
Assumed baseline
Illustrative, not real product data. Assume 22% of customers who receive a plain “this feature is going away” notice contact support or complain about it within 30 days.
Ethical guardrail
The guardrail. Balance migration must be identical, immediate, and lossless in both arms. Nothing about a customer's money may differ based on which notice they happened to receive.
The line. The reason given in the treatment notice has to be the real, actual reason the feature is being retired, not a plausible-sounding cover story invented to test messaging. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms are customers affected by the same real feature retirement. The only thing that changes is whether the notice explains why.
Control · no reason given
This feature is going away
Sub-accounts will no longer be available from 1 October. Your balance will move to your main savings account automatically.
Treatment · reason given
This feature is going away
Sub-accounts will no longer be available from 1 October. We're retiring them to keep pricing and features consistent across every savings account. Your balance will move automatically.
Sample size to actually trust this
Baseline22%complained, no reason given
Target14%−8pp with a reason
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm358affected customers
Total sample7162 arms
Assumed volume500/wkcustomers affected by the removal
Est. run time~2 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: % of affected customers who contact support or lodge a complaint within 30 days
✓Secondary: account closure rate in the 30 days after the notice
✓Guardrail: support contact-handling time and volume, the explanation itself must not generate confused follow-up questions that offset the drop in complaints
The study. Guéguen and Pascual's real street experiment found adding “but you are free to accept or to refuse” to an identical request for bus fare raised compliance from about 10% to about 47.5%.
The gap. That test was a single, low-stakes, in-person cash request. Nobody has tested whether the same freedom-naming line moves a digital, repeatable financial decision.
This experiment. It tests the identical mechanism on a real in-app ask: turning on round-up savings, a genuinely optional feature that's easy to feel nagged into.
Hypothesis
Customers shown an in-app prompt to turn on round-up savings that explicitly names their freedom to decline accept it at a higher rate than customers shown the identical prompt without that line.
Assumed baseline
Illustrative, not real product data. Assume 11% of customers shown a plain round-up savings prompt turn it on.
Control vs. treatment
Both arms see the identical round-up savings offer. Only whether the prompt names the freedom to decline differs.
Control · plain ask
Turn on round-up savings
Every purchase gets rounded up, the extra goes straight into savings.
Treatment · freedom named
Turn on round-up savings
Every purchase gets rounded up, the extra goes straight into savings. You're free to turn this on or leave it off, whichever suits you.
Sample size to actually trust this
Baseline11%opted in, plain ask
Target17%+6pp with freedom named
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm521customers prompted
Total sample1,0422 arms
Assumed volume600/wkeligible customers prompted
Est. run time~2 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: % of prompted customers who turn on round-up savings
✓Secondary: 30-day retention of the setting, whether freedom-framing attracts people who then quietly turn it back off
✓Guardrail: opt-out rate within 7 days of turning it on
Retail Banking · Fees & PricingAlso applies to: any tiered product upsell with a real, checkable eligibility gateNot yet run
Does framing a tier as restricted increase upgrade interest?
The founding reactance study proved restriction increases desire. Does that hold when the restriction is a real, honest eligibility gate, not a threat?
The study. Hammock and Brehm's founding reactance experiment found that actively restricting an option makes it more wanted, not just less available, the same lever plenty of marketing already leans on informally.
The gap. Nobody on this site has tested that lever as an honest upsell mechanism, one where the restriction named is real, not manufactured urgency.
This experiment. It tests whether naming a real, existing eligibility gate increases interest in upgrading, without inventing a scarcity that isn't there.
Hypothesis
Standard-tier customers shown a premium-tier message that names a genuine eligibility restriction click through to check their eligibility at a higher rate than customers shown the identical benefits with no restriction framing.
Assumed baseline
Illustrative, not real product data. Assume 6% of standard-tier customers shown a plain premium-benefits message click through to “see if you qualify.”
Ethical guardrail
The risk. A restriction that reads as real but isn't invites exactly the manufactured urgency this site's own Scarcity and Signpost Effect articles warn against.
The guardrail. The message may only state a restriction that's true at the moment it's shown: a real, current eligibility gate or a genuinely limited rollout, never a fabricated headcount or a countdown with no real deadline behind it.
What's being tested. Whether naming a real restriction changes interest, not whether a restriction can be invented to manufacture demand. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms see the identical premium benefits and face the identical real eligibility gate. Only whether the message names that restriction differs.
Control · plain benefits
Premium benefits
Lower fees, priority support, and higher transfer limits.
Treatment · restriction named
Premium benefits
Premium is currently offered to a limited group of customers. Lower fees, priority support, and higher transfer limits.
Sample size to actually trust this
Baseline6%clicked through, plain benefits
Target9%+3pp, restriction named
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm1,205standard-tier customers shown the message
Total sample2,4102 arms
Assumed volume900/wkeligible customers shown the message
Est. run time~3 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: click-through rate to “see if you qualify”
✓Secondary: completed upgrades among those who click through
✓Guardrail: complaint rate mentioning the exclusivity claim specifically, catching the message reading as manipulative even when it's technically true
Retail Banking · Credit & LendingAlso applies to: any rewards or loyalty programme with both a practical and an indulgent redemption optionNot yet run
Does completing a responsible credit action change how someone redeems reward points right after?
Same points, same redemption screen. Does having just done the responsible thing make someone more likely to pick the indulgent option over the practical one?
Builds on Licensing Effect · Khan & Dhar (2006), Journal of Marketing Research
Theory
The study. Khan and Dhar had people first imagine committing to volunteer work, then choose between a pair of jeans and a vacuum cleaner of equal price. Imagining the virtuous act pushed choice towards the jeans, the indulgent option, compared with people who imagined no such act first. A second study found the same shift towards luxury sunglasses over practical ones.
The gap. Both studies used an imagined, hypothetical virtuous act in a lab setting. Nothing on this site has tested whether a real, completed responsible action inside a live product changes a real choice made minutes later.
The prediction. If the mechanism is a genuine temporary boost to feeling like a responsible person, it shouldn't matter that the virtuous act here is financial (setting up an extra loan repayment) rather than social (volunteering): a real completed action should shift a real redemption choice the same way an imagined one shifted a hypothetical one.
Hypothesis
Customers who have just set up an automatic extra repayment on their credit card, then reach a reward-points redemption screen offering a practical option (a statement credit) and an indulgent option (a dining or travel voucher) of equal point value, choose the indulgent option more often than customers who reach the identical redemption screen without having just taken that action.
Assumed baseline
Illustrative, not real product data. Assume 35% of customers reaching the redemption screen without a prior prompt choose the indulgent voucher over the statement credit.
Ethical guardrail
The risk. This experiment measures a real effect on real spending choices. If treated carelessly, the same finding could be misread as licence to engineer a virtuous moment purely to sell more discretionary redemptions, which would work against the customer's own financial interest.
The guardrail. The responsible action being measured, the extra repayment opt-in, has to already exist as a genuinely useful feature offered on its own merits, not one invented or timed to manufacture the licensing effect. Both redemption options stay equally visible and equally easy to pick in every arm, and a guardrail metric tracks whether the extra repayment itself is cancelled again soon after, which would signal it was ticked as a box rather than adopted.
What's being tested. Whether completing a responsible action shifts a later, unrelated choice, not whether the responsible-action prompt itself can be tuned to appear more often than it should. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms see the identical redemption screen with the same two options at the same point value. The only difference is what happened in the screen immediately before it.
Control · straight to redemption
Redeem your points
12,000 points ready. Choose: $120 statement credit, or a $120 dining voucher.
Treatment · after the responsible action
Extra repayment set up
You'll pay an extra $50 off your card balance each month from now on. Nice work.
Tapping "Continue to rewards" in the treatment arm leads to the exact same redemption screen shown in the control arm. Only what the customer just did, and saw confirmed, differs.
Sample size to actually trust this
Baseline35%chose indulgent, control
Target45%+10pp, treatment
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm375customers reaching redemption
Total sample7502 arms
Assumed volume100/dayeligible redemptions
Est. run time~8 daysincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number. Swap in your own daily eligible-redemption count and the run time follows.
Outcome variables
✓Primary: % of redeeming customers who choose the indulgent voucher over the statement credit, control versus treatment
✓Secondary: overall redemption completion rate, checking the extra step in the treatment arm doesn't itself suppress redemption
✓Guardrail: % of treatment customers who cancel the extra repayment within 30 days, catching a responsible action taken only to unlock the redemption screen
✓Guardrail: average outstanding card balance 60 days after redemption, checking the shift towards indulgent redemption isn't paired with a rise in carried balance
The mechanism behind this:Licensing Effect, full study and citation on the Principles page.
Experiment blueprintDigital & Product Experiment
Retail Banking · Savings & DepositsAlso applies to: any product asking someone to commit a windfall, bonus, or refund to a future-focused optionNot yet run
Does a savings commitment made a week in advance survive contact with the actual moment?
Same customer, same decision, one week apart. Does deciding in advance land on a different number than deciding once the money's actually there?
Builds on Present Bias · Read & van Leeuwen (1998), Organizational Behavior and Human Decision Processes
Theory
The study. Read and van Leeuwen had office workers choose a snack for the following week, either fruit or junk food, at two different points: a week ahead of time, or on the day itself. Choosing a week ahead, hunger made almost no difference to the choice. Choosing on the day, hungry workers picked junk food far more often than workers who'd just eaten, even though everyone was choosing the exact same two snacks.
The gap. The study's reward is food. Nothing on this site has tested whether the same advance-versus-in-the-moment reversal shows up when the reward is money someone could spend right now instead of saving.
The prediction. If the mechanism is really about how steeply a reward's value drops once "now" is on the table, it shouldn't matter that the choice here is a savings percentage instead of a snack: a commitment made a week before a bonus or refund lands should favour saving more than the identical choice made once the money is actually sitting in the account.
Hypothesis
Customers who choose what share of an upcoming bonus or tax refund to transfer to savings a week before it lands commit a higher average percentage than customers who make the identical choice on the day the money actually arrives.
Assumed baseline
Illustrative, not real product data. Assume customers choosing on the day the money lands commit an average of 15% of it to savings.
Ethical guardrail
The risk. Locking someone into a savings commitment made a week before the money arrives could work against them if their circumstances change in that week, an unexpected bill, a job loss, an emergency, that makes the earlier number wrong by the time it would apply.
The guardrail. This only runs on customers who've already opted into a windfall-savings feature on its own merits, and the advance commitment stays editable or fully cancellable right up to the moment the money lands, with one tap and no penalty. Getting the timing wrong should cost the customer nothing.
What's being tested. Whether the timing of the ask changes the commitment, not whether advance commitments can be made harder to reverse. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms make the same choice about the same upcoming payment. The only difference is when they're asked.
Control · asked on the day
Your refund just landed
$800 arrived today. How much would you like to move to savings?
Treatment · asked a week ahead
Your refund lands in 7 days
Expecting around $800 on the 14th. How much would you like to move to savings when it arrives?
The estimated amount and the destination account are identical in both arms. Only the date on which the customer sets the percentage, and whether the money has already landed, differs.
Sample size to actually trust this
Baseline15%avg. committed, on the day
Target25%avg. committed, week ahead
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm100customers making a choice
Total sample2002 arms
Assumed volume25/dayeligible bonus/refund events
Est. run time~8 daysincl. buffer
Two-sample t-test for a continuous outcome, equal allocation: n = 2 × (zα/2 + zβ)² × σ² / (μ₁−μ₂)², assuming a standard deviation of 25 percentage points around each arm's average. The volume figure is an assumption to make the timeline concrete, not a real product number. Swap in your own daily eligible-event count and the run time follows.
Outcome variables
✓Primary: average % of the windfall committed to savings, control's on-the-day choice versus treatment's week-ahead choice
✓Secondary: of treatment customers, % who reduce or cancel their commitment before the money lands, the reversal Present Bias predicts
✓Guardrail: overdraft or missed-payment incidents in the 30 days after, checking the committed amount isn't leaving customers short
✓Guardrail: windfall-feature opt-out rate, treatment versus control, catching reactance to being asked to commit early rather than a genuine timing effect
Subscriptions & Telecom · Fees & PricingAlso applies to: any recurring plan with a real, checkable cheaper alternativeNot yet run
Does showing what you've already paid reduce comparison shopping at renewal?
Arkes and Blumer's field study proved a real sunk cost changes real behaviour. Does simply showing someone their own sunk cost suppress comparison shopping, even when a cheaper plan is one tap away?
The study. Arkes and Blumer's real field experiment found that theatergoers who paid full price for a season of plays attended more of it than theatergoers given the identical season at a random discount. Continued spending tracked what had already been spent, not what the season was actually worth going forward.
The gap. That study measured attendance after the sunk cost had already been paid. Nobody on this site has tested whether actively reminding someone of their own sunk cost, at the exact moment they'd otherwise consider switching, changes what they do next.
This experiment. It tests whether showing a customer their real, year-to-date spend on a plan, right on the renewal screen, reduces how often they check a cheaper alternative before renewing.
Hypothesis
Customers shown their own cumulative spend on the current plan at renewal click "compare other plans" less often than customers shown an identical renewal screen with no spend figure, even though the same cheaper plan is available to both.
Assumed baseline
Illustrative, not real product data. Assume 22% of customers shown a plain renewal screen click "compare other plans" before renewing.
Ethical guardrail
The risk. This tests a lever that could just as easily be used to trap customers on a worse deal by exploiting the same fallacy the Sunk Cost Fallacy article warns about.
The guardrail. The figure shown must be the customer's own real, verifiable spend, never rounded up or estimated in the company's favour, and the "compare other plans" option must stay exactly as visible and functional in both arms. Nothing about switching may get harder; only what's shown alongside the choice changes.
What's being tested. Whether showing a true sunk-cost figure changes behaviour on its own, not whether comparison shopping should be made harder. A result confirming the effect is a reason to avoid shipping this pattern, not a reason to ship it. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms see the identical renewal price and the identical "compare other plans" link. Only whether the screen states the customer's own year-to-date spend differs.
Control · plain renewal
Your plan renews
$69.99/month, unchanged.
Treatment · sunk cost shown
Your plan renews
$69.99/month, unchanged. You've paid $840 into this plan this year.
Sample size to actually trust this
Baseline22%clicked “compare,” plain screen
Target16%−6pp, sunk cost shown
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm667renewing customers shown the screen
Total sample1,3342 arms
Assumed volume800/wkrenewals reaching this screen
Est. run time~3 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: click-through rate to “compare other plans”
✓Secondary: overall renewal completion rate
✓Guardrail: complaint and negative-sentiment rate mentioning feeling pressured, plus churn rate 90 days later, catching a suppressed decision that just resurfaces as regret
The mechanism behind this:Sunk Cost Fallacy, full study and citation on the Principles page.
Experiment blueprintDigital & Product Experiment
Subscriptions · Fees & PricingAlso applies to: any checkout offering a monthly vs. annual choiceNot yet run
Does a truthful daily-equivalent price increase annual-plan uptake, even with the total shown right next to it?
Gourville's studies reframed a cost as pennies a day instead of stating the total. Does the same lift survive when the real annual total sits right next to it, not hidden?
The study. Gourville's lab studies found that restating a cost as a small daily amount, instead of one aggregate total, changed which other expenses it got compared against, and increased agreement to pay it.
The gap. Gourville's studies pitted the daily framing against the total, not alongside it. Most real checkouts that use a daily-equivalent price disclose the real total too, for legal and trust reasons, so whether the effect survives full disclosure is a live, untested question.
This experiment. It tests whether adding a truthful daily-equivalent figure next to the real annual total, changing nothing else about what's disclosed, still shifts more customers towards the annual plan.
Hypothesis
Customers shown "33¢/day, billed annually at $120" select the annual plan over the identical monthly alternative more often than customers shown "$120/year" alone, with the true total visible in both arms.
Assumed baseline
Illustrative, not real product data. Assume 18% of customers choose the annual plan over the monthly plan when shown the plain annual total.
Ethical guardrail
The risk. A daily-equivalent figure that isn't a true, correctly rounded division of the real total would mislead, exactly the failure mode the Temporal Reframing article names as the dishonest version of this lever.
The guardrail. The daily figure must be the exact annual total divided by 365, rounded normally, and the annual total must appear in the same type size, not fine print. Nothing about the monthly plan's price or visibility changes between arms.
What's being tested. Whether a fully disclosed daily reframe still moves choice, not whether disclosure can be minimised to make the reframe work harder. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms show the identical monthly plan as the alternative and the identical real annual total. Only whether a daily-equivalent figure appears alongside that total differs.
Control · total only
Annual plan
$120/year, billed annually.
Monthly: $12.99/mo
Treatment · daily equivalent added
Annual plan
33¢/day, billed annually at $120/year.
Monthly: $12.99/mo
Sample size to actually trust this
Baseline18%chose annual, total only
Target24%+6pp, daily equivalent added
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm719checkouts reaching the plan choice
Total sample1,4382 arms
Assumed volume500/wkcheckouts reaching the plan choice
Est. run time~4 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: share choosing the annual plan over monthly
✓Secondary: overall checkout completion rate, either plan
✓Guardrail: refund and support-contact rate within 30 days mentioning the price, catching a choice that felt smaller at checkout than it does on the first bill
The mechanism behind this:Temporal Reframing, full study and citation on the Principles page.
Experiment blueprintDigital & Product Experiment
Add-on Products · Fees & PricingAlso applies to: any cross-sell offering two or more add-ons with real, equal-sized discountsNot yet run
Does stating one discount on a bundle beat splitting the identical discount across its parts?
Yadav and Monroe found a saving stated on the bundle carries more weight than the same saving split across items. Does that hold when the two items are financial add-ons, not retail goods?
The study. Yadav and Monroe found that an additional discount stated directly on a combined bundle price carried more weight in perceived transaction value than the identical dollar amount split across the bundle's individual items.
The gap. The original study tested this with described purchase scenarios for physical goods. Nobody on this site has tested whether the same framing shift changes real take-up of a financial add-on cross-sell, where the "items" are services rather than products.
This experiment. It tests whether the identical total discount, stated once on a two-add-on bundle instead of once per add-on, increases how often customers take both.
Hypothesis
Customers offered "save $2 when you bundle ID theft protection and device insurance together" take both add-ons more often than customers offered "save $1 off ID theft protection" and "save $1 off device insurance" as two separate line items, for the identical $2 total saving.
Assumed baseline
Illustrative, not real product data. Assume 11% of customers take both add-ons when the two discounts are shown separately, one per item.
Ethical guardrail
The risk. A bundle framing that makes a saving feel bigger than it is could sell add-ons a customer wouldn't otherwise want, then get cancelled once the bill arrives and the framing has worn off.
The guardrail. The bundle discount must equal, not exceed, the sum of what the two item-level discounts would have been; no inflated "bundle-only" saving may be invented to make the framing test unfairly favourable.
What's being tested. Whether the framing itself changes take-up for an identical real saving, not whether a bundle can be priced to look better than it is. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms offer the identical two add-ons at the identical final prices. Only whether the $2 saving is stated once, on the bundle, or twice, one per item, differs.
Control · per-item discounts
Two add-ons
ID theft protection: $4/mo, save $1. Device insurance: $4/mo, save $1.
Treatment · bundle discount
Bundle both, save $2
ID theft protection + device insurance, together, save $2/mo when bundled.
Sample size to actually trust this
Baseline11%took both, per-item discounts
Target16%+5pp, bundle discount
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm729customers shown the cross-sell
Total sample1,4582 arms
Assumed volume350/wkcustomers reaching the cross-sell
Est. run time~6 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: share of customers taking both add-ons
✓Secondary: share taking exactly one add-on, checking whether the bundle framing cannibalises single-item take-up
✓Guardrail: cancellation rate of the bundled add-ons within 60 days, catching a deal that felt bigger at checkout than the add-ons turned out to be worth
The mechanism behind this:Bundling, full study and citation on the Principles page.
Experiment blueprintDigital & Product Experiment
BNPL Provider · Credit & LendingAlso applies to: payday lenders, overdraft products, any short-term credit where fees compound on a missed paymentNot yet run
Does naming the real dollar cost of repeat late fees cut how much someone borrows next?
A real field test on payday loans found stating dollar fees over time cut future borrowing by 11%. Does the identical format work inside a BNPL app, at the moment of the next purchase?
Builds on Bertrand & Morse (2011)'s real field experiment on payday loan disclosure formats, covered in full in The Fine Print Nobody Reads · also builds on Disclosure Backfire for what a badly-designed version of this could do instead
Theory
The study. Bertrand and Morse ran a real field experiment inside 77 payday loan stores, randomising 1,441 real borrowers to a control group or one of several disclosure formats shown at the counter. An APR figure, and a comparison to credit-card APRs, barely moved anything. A format that instead named the real dollar fees a borrower would accumulate if the loan renewed repeatedly cut their borrowing over the following four months by about 11%.
The gap. That disclosure was delivered by a person, on paper, at a physical counter, for a payday loan. Nobody has tested whether the identical format, a real dollar-fees-over-time figure, delivered inside a digital checkout screen instead of by a person, moves behaviour the same way for Buy Now, Pay Later, a product built around exactly the repeat, compounding late fee the original study measured.
This experiment. It tests the same decision Bertrand and Morse tested: an existing high-cost-credit borrower's next borrowing decision, not whether to complete the purchase in front of them right now. The dollar-cost figure is tied to the current order's own real terms, a projection, not a personal history lookup, mirroring exactly what the original disclosure format showed.
Hypothesis
Existing BNPL users with at least one prior late payment, shown the real dollar late-fee cost of missing payments on their current order at the moment they confirm it, go on to miss fewer payments and take out fewer new BNPL purchases over the following months than users shown only the standard terms link.
Assumed baseline
Illustrative, not real BNPL provider data. Assume 24% of users with a prior late payment miss a payment again on their next order when shown only the standard terms link.
Ethical guardrail
The line. Showing a real dollar-cost projection to someone who has already missed a payment is not a neutral change, so the restraint is worth stating before the design. This isn't built to talk anyone out of BNPL when it fits how they actually intend to pay. Plenty of BNPL purchases are paid on time and cost nothing extra.
The restraint. The figure shown is a real projection tied to the current order's own instalment terms, never an invented worst case, and it appears alongside the purchase, not in place of it: nothing about the checkout flow itself is slowed down or blocked.
The risk this specifically guards against. The whole reason to test this rather than assume it works is Disclosure Backfire: disclosure doesn't reliably help just because it's honest. A late-fee warning that reads as an accusation, or that pushes a borrower towards a less transparent product instead of paying more carefully, would be the intervention backfiring, not succeeding.
Control vs. treatment
Both arms see the identical order and the identical instalment schedule. Only whether the late-fee cost is named in real dollars differs.
If you miss 2 of these 4 payments, the way an average repeat late payer does, this order costs $30 more in late fees: $180 total.
The $30 figure is this order's own real late-fee schedule applied twice, not an invented or worst-case number, and confirming the order takes the identical number of taps in both arms.
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². This covers the primary missed-payment outcome only. The volume figure is an assumption to make the timeline concrete, not a real provider number.
Outcome variables
✓Primary: share of the current order's instalments paid late or missed
✓Secondary: total BNPL amount borrowed over the following 4 months, the same window and outcome type Bertrand & Morse's own result was measured on
✓Guardrail: checkout abandonment rate on the current order, catching a warning that just kills otherwise-affordable purchases instead of changing repayment behaviour
✓Guardrail: rate of switching to a less transparent, higher-cost credit product for a purchase the user would otherwise have made with BNPL, checking the disclosure didn't just push the same borrowing somewhere worse
The mechanism behind this:The Fine Print Nobody Reads, the full Bertrand & Morse teardown and the wider case for why most disclosure goes unread or backfires. See also Disclosure Backfire on the Principles page.
Experiment blueprintDigital & Product Experiment
Financial Services · Credit CardsAlso applies to: membership organisations, insurers, and any subscription business with an annual renewal decisionNot yet run
Does a printed annual member guide increase renewal more than a digital-only version?
Touch is proven to raise how much people value a physical object. Nobody has tested whether that shows up in an actual renewal decision a year later.
The studies. Peck and Shu found that physically touching an object, even briefly, raises how much a person values it and how much they'll pay for it. Atasoy and Morewedge found the same underlying mechanism running the other direction: identical content is valued and paid for less once it becomes a purely digital good, because a digital copy generates a weaker sense that it's “mine.”
The gap. Both studies tested everyday consumer objects and generic digital content: mugs, sunglasses, souvenir photos, books, films. Neither tested a recurring benefit mailer tied to an actual, dated renewal decision for a paid product. Whether the same physical-ownership effect shows up in real renewal behaviour, months after the object arrived, is untested.
This experiment. It tests one level up from a lab object: a premium credit card's own annual member guide, sent either as a printed booklet or as a digital-only equivalent inside the app, with renewal at the card's next annual fee date as the real outcome.
Hypothesis
Cardholders who receive a printed annual member guide, in addition to the digital version every cardholder already gets in the app, will report higher perceived value of the card at 90 days and renew at a higher rate at their next annual fee date, compared with cardholders who receive the digital-only version.
Assumed baseline
Illustrative, not real card-programme data. Assume the premium card's current annual renewal rate is 85%, with the digital-only guide as the existing standard.
Ethical guardrail
The line. A physical mailer costs materially more per cardholder than a push notification. Testing it against a cheaper default is a fair comparison. But only if the digital-only arm still gets the full, real content, not a deliberately thinner version built to make print look better by comparison.
The restraint. Both arms receive identical content. The only difference is the medium the guide arrives in, and every cardholder still gets the in-app digital version regardless of which arm they're in.
The risk this specifically guards against. A cardholder who notices peers received a printed booklet and they didn't could read the difference as being treated as a lesser member, an unfairness read this site covers under Perceived Fairness. The complaint-rate guardrail below exists to catch that directly.
Control vs. treatment
Every cardholder in both arms sees the identical digital guide in the app. Only the treatment arm also receives a printed booklet in the mail.
Control · digital-only
Your Member Guide
This year's benefits, travel perks, and partner offers, all in the app.
Treatment · printed booklet + digital
Your Member Guide, printed
The same content, bound as a booklet and mailed to the cardholder's address on file.
Arrives by mail
The printed booklet supplements the in-app guide; it never replaces it. Randomisation happens at the cardholder level, not the household level, to avoid one printed copy being shared across two cardholders at the same address.
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The long run time is structural, not padding. Enrolling the full sample takes about 7–8 months at the assumed volume. But the outcome itself only fires on each cardholder's own annual renewal date, so the last cardholders enrolled don't produce a result until roughly a year after they join.
Outcome variables
✓Primary: renewed or cancelled at the cardholder's own next annual fee date
✓Secondary: self-reported perceived value of the card, 1–7 scale, surveyed at 90 days
✓Guardrail: complaint or support-contact rate about not receiving a printed guide, among digital-only cardholders
✓Guardrail: print and postage cost per incremental renewal, weighed against the annual fee retained, so a real effect can still be judged not worth the cost
Retail Banking · Savings & DepositsAlso applies to: any investing app, super fund, or robo-advisor letting users hand-pick from a flat list of portfolios or fundsNot yet run
Does trimming a 10-option investing menu to 3 curated ones change how evenly people spread their money?
Benartzi and Thaler found real retirement savers roughly split their contributions evenly across however many funds a plan happened to offer. Does a real micro-investing app's own menu do the same thing to a new user's first portfolio?
The study. Benartzi and Thaler found retirement savers roughly split their contributions evenly across however many funds their employer's plan happened to list, and the resulting stock/bond mix tracked the mix of funds on offer far more closely than any stated risk preference.
The gap. That's field data from employer-chosen 401(k) menus, not a menu a product team can freely redesign and re-test inside its own app, and it predates today's single-screen, pick-as-many-as-you-like investing menus like CommSec Pocket's own. No test yet checks whether trimming a real investing app's own option list changes how evenly a real user's first contribution ends up spread across whatever they pick.
This experiment. It tests whether replacing a flat, 10-option "Invest in" menu with a curated set of 3 pre-grouped options, each already diversified inside itself, changes how many portfolios a new user selects and whether the resulting split across them still tracks pure equal-weighting by count, inside the same signup flow, for the same app.
Hypothesis
New users shown the trimmed, 3-option menu are more likely to end up with a genuinely diversified first portfolio than users shown the original 10-option menu, even though both groups still tend to split their money close to evenly across however many options they pick. The 10-option group is more likely to unintentionally overweight several overlapping single-country or single-theme funds by picking, say, 4 of the 10 and splitting evenly across them.
Assumed baseline
Illustrative, not real/disclosed data. Assume 45% of new signups shown the existing 10-option menu end up with a first-session portfolio classified as genuinely diversified, spanning at least two distinct, non-overlapping asset classes rather than several near-duplicate single-country or single-theme picks split evenly.
Ethical guardrail
The guardrail. The test can only change how options are grouped and how many are shown by default, never nudge a user toward investing more than they intended or hide a legitimate option's existence entirely. Every option in the original 10-item menu must stay reachable through the curated screen, just organised differently, and any "already diversified" curated option must actually be diversified, not marketed as such while concentrated.
What's being tested. Whether restructuring the menu reduces the accidental overlap naive equal-weighting produces, not whether a shorter, more persuasive menu can funnel new users into house-preferred, higher-fee products. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms are new sign-ups reaching the "Invest in" screen for the first time, randomly assigned at account creation. Every option available in the control's flat list is still reachable in the treatment, just organised differently.
Control · flat 10-option menu
Invest in
Aussie Top 200, Aussie Dividends, Aussie Sustainability, Aussie Corporate Bonds, Global 100, Diversified Equities, Emerging Markets, Health Wise, Sustainability Leaders, Tech Savvy.
Treatment · 3 curated options
Invest in
Balanced Starter: already spread across shares, property and bonds. Aussie Shares: broad exposure to the ASX. Go further: see all 10 individual options.
Nothing removed: every original option is still one tap away behind “Go further.”
Same account type, same underlying ten options, same signup flow. Only how many choices are shown by default, and how they're grouped, differs between arms.
Sample size to actually trust this
Baseline45%diversified first portfolio, control
Target60%+15pp assumed lift, treatment
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm171new signups reaching the menu
Total sample3422 arms
Assumed volume1,500/monew app signups
Est. run time<1 moto reach 342 at 1,500/mo split evenly
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number. Swap in an actual monthly signup count and the timeline follows.
Outcome variables
✓Primary: proportion of new accounts whose first-session portfolio is classified as genuinely diversified, spanning at least two distinct, non-overlapping asset classes, treatment versus control
✓Secondary: average number of options selected per new account, confirming the trimmed menu changes allocation quality rather than just suppressing engagement
✓Secondary: total dollar amount invested in the first session, confirming a decluttered menu doesn't reduce how much people are willing to commit
✓Guardrail: click-through rate on "Go further" must not collapse to near zero, confirming the other 7 options stay genuinely reachable, not just nominally present
✓Guardrail: support contacts about missing or hidden options must stay within the normal range for new signups
The research behind this:Diversification Heuristic, full citation and the real CommSec Pocket menu behind this test. See also Choice Overload, the companion finding that more options can make choosing itself harder even before the 1/n split kicks in.
Experiment blueprintDigital & Product Experiment
Retail Banking · Credit & LendingAlso applies to: any windfall, bonus, or lump sum an app can detect and apply toward an existing debtNot yet run
Does directing a tax refund to pay down a credit card increase spending right after?
Same refund, same balance reduction. Does confirming the paydown happened make someone spend more on everything else in the weeks after?
Builds on Licensing Effect · Khan & Dhar (2006), Journal of Marketing Research
Theory
The study. Khan and Dhar had people first imagine committing to volunteer work, then choose between two equally-priced products, one practical, one indulgent. Imagining the virtuous act shifted choice toward the indulgent option compared with people given no such prompt first, evidence that a completed (or even merely imagined) responsible act can license a later indulgence.
The gap. The reward-redemption experiment already on this site tests this mechanism on a small, already-earned resource, points, choosing between two redemptions of equal value. Nothing on this site tests it on new money entering someone's life: a real lump sum with no assignment yet, and a real choice about whether it goes toward debt.
The prediction. If confirming a virtuous act is what does the licensing, not just the act itself, then two customers whose refund is applied to their card balance identically should still end up spending differently, depending only on whether the app told them so.
Hypothesis
Among customers already enrolled in an automatic Refund Auto-Payoff feature, customers who receive an explicit confirmation that their tax refund was applied to their credit card balance show a bigger rise in discretionary card spending over the following 30 days than customers whose identical refund is applied to their balance without any distinct confirmation, just a line in their ordinary transaction history.
Assumed baseline
Illustrative, not real product data. Assume 40% of customers whose refund is applied without a distinct confirmation show a rise in discretionary card spend above their own prior 30-day average in the month that follows, ordinary post-windfall spending, refund or no confirmation.
Ethical guardrail
The risk. A bank has a real financial interest in this specific finding: a confirmation message that increases discretionary spending after a debt paydown is a way to win back, in interest and fees, some of what the customer just saved. That incentive makes it tempting to design the confirmation to maximise the effect rather than to honestly inform the customer.
The guardrail. The underlying paydown is identical and unaffected in both arms; every enrolled customer's refund is applied to their balance the same way regardless of which arm they land in. Only whether they're told about it, not whether it happens, its size, or its timing, is being tested. The confirmation message itself states only true, specific facts, the dollar amount applied and the new balance, with no added encouragement to spend, celebrate, or reward the moment.
What's being tested. Whether making a real debt reduction visible changes later spending, not whether the confirmation should be redesigned to push that spending further. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms are already enrolled in Refund Auto-Payoff, and both refunds are applied to the card balance identically, in full, the moment the refund lands. The only difference is whether that fact is confirmed to the customer as a distinct moment.
Control · no confirmation moment
Transaction history
ATO refund applied · $1,200. Card ending 4821. Balance: $2,340.
Treatment · explicit confirmation
Refund applied
Your $1,200 tax refund paid down your card. Nice work, your new balance is $2,340.
Both customers' $1,200 refund reduces their balance by exactly $1,200. Only whether the app pauses to tell them so, as its own screen rather than a line in a list, differs.
Sample size to actually trust this
Baseline40%rose above own baseline, control
Target52%+12pp, treatment
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm270enrolled customers with a detected refund
Total sample5402 arms
Assumed volume45/dayduring tax-refund season
Est. run time~14 daysincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number, and tax-refund season is itself seasonal, so a real run would need to fit inside that window.
Outcome variables
✓Primary: % of customers whose discretionary card spend in the 30 days after the refund exceeds their own prior 30-day average, control versus treatment
✓Secondary: average card balance at day 30, checking whether a confirmed paydown gets quietly re-spent back up
✓Guardrail: % of treatment customers who opt out of refund confirmation messaging, a signal the moment itself is unwanted
✓Guardrail: average interest paid over the following 90 days, checking whether any licensed spending costs the customer more than the paydown just saved them
The mechanism behind this:Licensing Effect, full study and citation on the Principles page.
Cards & Rewards · Fees & PricingAlso applies to: interest-rate offers, loyalty tier upgrades, or any two-option choice where a peer benchmark can be shown truthfullyNot yet run
Does showing a customer where a cashback rate ranks against similar spenders pull them toward the lower-paying card?
Same two real rates, same real dollars. Does knowing one card would leave you ahead of similar customers make people choose it over the card that simply pays more?
Builds on Positional Concern · Solnick & Hemenway (1998), Journal of Economic Behavior & Organization
Theory
The study. Solnick and Hemenway asked people to choose between two hypothetical worlds: earning $50,000 while others earned $25,000, or earning $100,000 while others earned $200,000. A majority picked the world with less money outright, as long as it left them ahead of everyone around them.
The gap. That study was a hypothetical survey with no real money and no real product on the other end of the choice. Nothing on this site tests whether the same pull shows up when the two options are real, dollar-denominated financial products a customer could actually pick.
The prediction. If relative standing carries real weight in an actual purchase decision, adding a truthful peer benchmark to two unchanged cashback rates should move some customers off the higher-paying card and onto the one that merely ranks better.
Hypothesis. Among customers choosing between two real cashback rates on their own card spend, customers who are also shown a truthful benchmark of what similar-spend customers typically earn choose the lower-rate card at a higher rate than customers shown the same two rates with no benchmark.
Assumed baseline
Illustrative, not real product data. Assume 6% of customers pick the lower-rate card when shown only the two rates, most likely people choosing on a card feature other than cashback, rather than misreading the numbers.
Ethical guardrail
The risk. A card issuer earns more margin on the lower-cashback card. A rank benchmark that pulls customers toward it, dressed up as helpful context, would cost real customers real money for the issuer's benefit.
The guardrail. Both rates are real, unchanged card products either arm could choose without this test running at all. The benchmark shown is a real, independently computed figure, not a number tuned to produce the result, and it's calculated the same way in both arms; only whether it's displayed differs.
What's being tested. Whether visible relative standing shifts a real choice away from the higher-paying option, not whether that shift should be built into how the cards are sold. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms see the same two card offers side by side, Card A paying a higher cashback rate and Card B paying a lower one. The treatment arm adds one extra line to each card: how that rate compares with what similar-spend customers typically earn.
Control · rates only
Choose your card
Card A: 2.2% cashback on everyday spend. Card B: 1.8% cashback on everyday spend.
Treatment · rates plus benchmark
Choose your card
Card A: 2.2% cashback. Similar spenders typically earn 2.4% on this card. Card B: 1.8% cashback. Similar spenders typically earn 1.3% on this card.
Card A still pays more than Card B in both arms, in real dollars, on identical spend. Only whether each rate's standing against similar customers is shown changes between them.
Sample size to actually trust this
Baseline6%chose Card B, control
Target15%+9pp, treatment
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm180customers mid card-choice flow
Total sample3602 arms
Assumed volume25/daythrough the card-upgrade flow
Est. run time~15 daysincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: % of customers who choose Card B, the lower-cashback-rate card, control versus treatment
✓Secondary: average annual cashback forgone per customer who switches to Card B under the benchmark, the real dollar cost of the effect if it appears
✓Guardrail: comprehension check score on a short follow-up question about which card pays more, ruling out simple confusion as the cause of any shift
✓Guardrail: complaint or support-contact rate referencing the card comparison screen
The mechanism behind this:Positional Concern, full study and citation on the Principles page.
Distinct from:Social Norm, which tests what most people do; this tests whether a customer ends up ahead of or behind others, regardless of what the group's own behaviour is.
Experiment blueprintDigital & Product Experiment
Retail Banking · Credit & LendingAlso applies to: any automated decline, denial, or rejection notice a customer receives with no person actually delivering itNot yet run
Does a benevolence-framed decline reduce how harshly customers judge the bank after a loan rejection?
Same decline, same reason codes. Does opening with one sentence of genuine care change how a customer feels about the bank that sent it?
Builds on Shooting the Messenger · John, Blunden & Liu (2019), Journal of Experimental Psychology: General
Theory
The study. John, Blunden and Liu found that people who deliver bad news get rated less likeable, even when a random process, not the messenger, produced the outcome. The penalty shrank when the messenger explicitly signalled a benevolent motive while delivering the news.
The gap. That effect was shown for a person delivering news to another person. A loan decline is usually the opposite: a templated, automated message with no human messenger at all. Nothing on this site tests whether benevolence-framing still works when the “messenger” is a system-generated letter.
The prediction. If genuine care in the delivery softens the penalty even for a live human messenger, adding an explicit, honest statement of care to an otherwise identical automated decline should soften how a declined applicant feels about the bank, without changing the decision itself at all.
Hypothesis. Among applicants who receive an identical loan decline, same underlying decision, same reason codes, applicants whose decline opens with an explicit, genuine statement of care report less negative sentiment toward the bank on a short post-decision survey than applicants who receive the current generic decline message.
Assumed baseline
Illustrative, not real product data. Assume 34% of applicants who receive the current generic decline rate their experience 1 or 2 stars out of 5 on a short post-decision survey.
Ethical guardrail
The risk. A bank has an interest in fewer complaints and better sentiment scores, regardless of whether the underlying decision was actually fair. Warmer language could be used to make a bad or biased decision feel better instead of making the decision itself better, or to quietly discourage a customer with a real grievance from escalating it.
The guardrail. The underlying credit decision, the reason codes shown, and the applicant's appeal rights are identical in both arms; only the opening framing sentence changes. That sentence states only true things, real next steps and real resources, and never implies the decision could change just by asking.
What's being tested. Whether framing changes how a customer feels about an unchanged decision, not whether framing should be used to reduce complaints about decisions that are actually wrong. See Not Testing Is Still a Bet.
Control vs. treatment
Both arms receive the identical decision, reason code, and reapplication timeline. The only difference is the opening of the decline message itself.
Control · current generic decline
Application update
Your personal loan application has been declined based on our lending criteria. Reason: insufficient serviceable income. You may reapply after 90 days.
Treatment · benevolence-framed opening
Application update
We know this isn't the outcome you were hoping for, and we want to be upfront about why. Your application was declined due to insufficient serviceable income. You may reapply after 90 days, and we've included what could change next time.
The decision, the reason code, and the reapplication window are identical in both arms. Only the opening sentence and its framing differ.
Sample size to actually trust this
Baseline34%1–2 star rating, control
Target24%−10pp, treatment
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm320declined applicants surveyed
Total sample6402 arms
Assumed volume60/daydeclined applications
Est. run time~12 daysincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: % of declined applicants rating their experience 1 or 2 stars out of 5 on a short post-decision survey, control versus treatment
✓Secondary: complaint or escalation rate within 14 days of the decline
✓Guardrail: rate of customers who use the formal appeals process, checking that warmer framing doesn't discourage a legitimate appeal
✓Guardrail: reapplication rate within 90 days, checking that warmer framing doesn't quietly discourage applicants who could genuinely qualify next time
The mechanism behind this:Shooting the Messenger, full study and citation on the Principles page.
Experiment blueprintDigital & Product Experiment
Retail Banking · Credit & LendingAlso applies to: any booked call or appointment where the customer's own follow-through, not the reminder itself, determines whether they show upNot yet run
Does asking a customer to state their own plan reduce no-shows on a home loan appointment?
A phone appointment only needs the customer to answer. A branch appointment needs them to actually get there. Does asking them to say what they'll do, or how they'll get there change whether they show up, either way?
Builds on Implementation Intentions · Gollwitzer & Brandstätter (1997), Journal of Personality and Social Psychology
Theory
The study. Gollwitzer and Brandstätter found that specifying exactly when and where you'll act on a goal makes you far more likely to actually follow through than holding the same goal without a specific plan.
The gap. CommBank's real appointment flow (commbank.com.au/appointment) lets a customer book a home loan appointment by phone or in-branch. Both already tell the customer everything an implementation intention would need: the confirmed time, and exactly what to have ready. But it's all stated to the customer, never asked back. An in-branch appointment also carries a barrier a phone call never has: actually getting to a specific address on time.
This experiment. It tests four different ways of asking for a plan on the confirmation screen. One names a vivid, specific moment. One has the customer state the whole plan back in their own words. One names what might get in the way and plans around it. One is a light, optional suggestion with nothing to fill in. Running the identical four against both a phone appointment and an in-branch appointment separates two real questions. Which approach works at all, and whether the same one works as well for a remembering problem as it does for a travelling one.
Hypothesis. Among customers who book a home loan appointment, by phone or in-branch, customers shown any of the four confirmation-screen prompts show a lower no-show rate than customers who are only ever told the details. Whether the four prompts rank the same way for both appointment types is a separate question. A travelling problem might call for a different one of them than a remembering problem does, and that's itself part of what this experiment would show.
Assumed baseline
Illustrative, not real CommBank data. Assume 22% of booked home loan phone appointments end in a no-show under the current reminder-call script.
Illustrative, not real CommBank data. Assume 28% of booked in-branch home loan appointments end in a no-show, higher than the phone rate, since travelling somewhere by a set time is a real added barrier a phone call never has.
Ethical guardrail
The risk. A prompt that presumes a specific time, travel mode or personal obstacle can feel intrusive if it asks for more than a customer wants to give before they've even arrived. A prompt with no real payoff to the customer can also read as the bank telling them what to think, rather than helping them plan something they'd genuinely benefit from planning.
The guardrail. The appointment, the specialist or branch, and the actual information required are identical across all five arms, whichever appointment type is shown. Every confirmation-screen prompt is optional and skippable, any options it offers are neutral and non-exhaustive, and none of them presume which obstacle, travel mode or moment matters most to a given customer.
What's being tested. Whether asking the customer to state their own plan changes follow-through on an identical appointment, not whether the confirmation screen or reminder call itself should exist. See Not Testing Is Still a Bet.
The reminder-fatigue risk. The reminder call itself carries a real, separate cost this experiment doesn't measure: repeated contact can lift today's response while quietly raising how many customers opt out of future contact altogether. See Reminder Fatigue.
The real flow today
Spotted in the wild
Real screenshots from Commonwealth Bank's live home loan appointment booking flow (commbank.com.au/digital/homeloanappointment), captured booking a real phone appointment slot. The same flow also offers an in-branch option; no real screenshots of that specific path were captured for this page.
The confirmation screen: a date, time and specialist, plus “Add to Calendar.” The specialist's name is blurred here for their privacy.
Further down the same confirmation: a reminder call is already promised, “to confirm the time and location and to ask you a couple of quick questions,” followed by a real document checklist.
This is a real, current product flow, not a demonstration that any change to it works. The reminder call and checklist already exist; nothing here shows the proposed prompts actually being tested.
Control vs. four treatments
All five arms book the identical appointment, with the identical specialist or branch and identical information required, whichever appointment type is shown below. Use the toggle to see how the same four prompts adapt to a phone appointment versus an in-branch one.
Control · today's screen
Appointment confirmed
Thursday 10:30am with your Home Lending Specialist. We'll call a few days before to confirm the time and go through what to bring.
Treatment A · vivid cue
Appointment confirmed
Thursday 10:30am with your Home Lending Specialist.
Where will you be when the call comes in?
Customers tell us this makes the call feel less rushed, one less thing to think about on the day.
Thursday 10:30am with your Home Lending Specialist.
Entirely optional: some customers find it helps to have a rough mental picture of 10:30am ahead of time, where they'll be when the call comes in, and what they'd do if it catches them at a bad time.
Nothing to fill in here, just something worth having in the back of your mind.
Control · today's screen
Appointment confirmed
Thursday 10:30am with your Home Lending Specialist, CommBank George Street. Bring photo ID and your last two payslips.
Treatment A · vivid cue
Appointment confirmed
Thursday 10:30am, CommBank George Street.
What will you be doing right before you leave?
Customers tell us this makes the trip feel less rushed, one less thing to think about on the day.
Finishing breakfastAt work alreadyDropping the kids off
Entirely optional: some customers find it helps to have a rough mental picture of 10:30am ahead of time, when they'll leave and how they'll get there.
Nothing to fill in here, just something worth having in the back of your mind.
All five arms get the same appointment, the same specialist or branch and the same information requirements. Only whether the customer is asked to state a plan back differs. These are illustrative recreations of an app screen, not real screenshots.
Why four treatments, not one. Each one tests a different psychological lever from the same underlying theory. Treatment A makes the moment itself more vivid and easier to notice when it arrives. Treatment B has the customer build the whole if-then plan themselves, the mechanism the site's own principle article centres. Treatment C names a specific obstacle in advance and plans a response to it, the goal-shielding effect from the same research tradition. Treatment D tests whether any interaction is actually necessary at all, offering the identical idea as something to simply read. If Treatment D performs close to Control, active response is doing the real work. If it performs close to A, B or C, a lighter, optional suggestion might be doing most of that work on its own.
Sample size to actually trust this
Baseline22%no-show rate, control
Target, Treatment A18%−4pp, vivid cue
Target, Treatment B18%−4pp, self-stated plan
Target, Treatment C19%−3pp, name the obstacle
Target, Treatment D19.5%−2.5pp, text only
Significanceα = 0.05two-sided
Power80%1−β
Sample per arm4,127booked appointments
Total sample20,6355 arms
Assumed volume40/daybookings through this flow
Est. run time~74 wksincl. buffer
Two-proportion test, equal allocation: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁−p₂)². Sized against Treatment D's smaller expected effect, the weakest of the four since it asks nothing of the customer beyond reading; Treatments A, B and C are correspondingly over-powered at this sample. If Treatment D's real effect turns out even smaller than assumed here, running it as a smaller, separate exploratory test rather than a full arm of this trial would be the more honest call. The volume figure is an assumption to make the timeline concrete, not a real product number.
Same test, same logic, sized against Treatment D again. A lower real-world volume and a higher baseline no-show rate both push this well past a year. At this real volume, a full-powered test of the weakest arm alone would take close to three years. That's a strong signal Treatment D belongs in a higher-volume channel first, or as a much smaller confirmatory follow-up once A, B and C have already run. The volume figure is an assumption to make the timeline concrete, not a real product number.
Outcome variables
✓Primary: % of booked appointments ending in a no-show, each treatment versus control, for whichever appointment type is shown
✓Secondary: pairwise comparisons among Treatments A through D, isolating which mechanism, a vivid cue, a self-generated plan, a named obstacle, or a passive suggestion, produces the larger effect
✓Secondary: whether the ranking among the four treatments differs between the phone and branch versions, testing whether a remembering problem and a travelling problem call for different tools
✓Guardrail: confirmation-screen abandonment rate must not rise in any treatment, every prompt has to stay skippable in practice, not just on paper
✓Guardrail: customer complaints or negative feedback specifically referencing any of the four prompts
The mechanism behind this:Implementation Intentions, full study and citation on the Principles page, including the goal-shielding mechanism Treatment C is built on.
Also worth citing: Nickerson, D. W., & Rogers, T. (2010). “Do You Have a Voting Plan? Implementation Intentions, Voter Turnout, and Organic Plan Making.” Psychological Science, 21(2), 194–199. A real field experiment, 287,228 voters ahead of the 2008 US presidential election: a call that helped people form a specific plan (what time, coming from where, doing what beforehand) raised turnout by 4.1 percentage points overall, and by 9.1 points in single-voter households. The branch version of Treatment B adapts that exact three-part structure directly.
Also worth citing: Martin, S. J., Bassi, S., & Dunbar-Rees, R. (2012). “Commitments, Norms and Custard Creams: A Social Influence Approach to Reducing Did Not Attends (DNAs).” Journal of the Royal Society of Medicine, 105(3), 101–104. A real NHS trial found asking patients to state their own appointment time back to staff, rather than simply being told it, measurably reduced missed appointments, the same self-stated-plan mechanism behind Treatment B here.