Home / Pricing pages / Research brief
Download raw .md

Research Brief — Pricing-Page Psychology

Fact-check reference for the YouTube script

Confidence key SOLID — primary source located, figure verified CONTESTED — real source exists, but methodology or interpretation is disputed SHAKY — widely repeated, source is weak, absent, or circular BLACKLISTED — never repeat this claim

Cross-reference: choice overload (Iyengar jam study, Scheibehenne meta, Chernev moderators) is fully covered in decision-fatigue-research-brief.md — do not rehash here. The relevant bridge: Chernev's "a dominant option makes overload disappear" is the empirical case for a highlighted "Most Popular" tier.


️ PART 1 — CORRECTIONS

The standard tellings of six pricing-page staples needed changing. In each case the corrected version is either more accurate and more powerful, or is your honesty beat.

1. The Economist decoy — a classroom demo, not a field result THE HONESTY BEAT

What actually happened — Ariely, Predictably Irrational (2008), ch. 1. The pricing on economist.com was real (Ariely reproduces the actual ad), but the experiment was a survey of 100 MIT Sloan MBA students, never a field test. Web-only $59 / print-only $125 / print+web $125 → 16 / 0 / 84. Decoy removed, different 100 students → 68 / 32. SOLID on the numbers, but The Economist never measured anything. Nobody knows what the decoy did to actual subscriptions.

Three more layers, all verified: - Frederick, Lee & Baskin (2014, JMR, footnote 8) note the demo "is identical to one discussed in Kivetz, Netzer, and Srinivasan (2004)," whose version produced a milder 43% → 72%. Ariely's 32→84 is the flashy outlier retelling. - Ariely credibility caveat: the 2012 PNAS honesty paper was retracted (Sept 2021) after Data Colada showed the insurance data was fabricated; Duke's investigation (disclosed 2024) concluded the data was falsified but found no evidence Ariely knowingly falsified it. The Economist demo is unrelated, but if you lean on Ariely, expect comments. Coherent arbitrariness (§2) is co-authored with Loewenstein & Prelec — safer. - And the kicker below: the effect the demo illustrates largely fails with real products (Correction 2).

Safe wording: "Ariely showed the ad to a hundred MBA students... and when he deleted the middle option for another hundred students, preferences flipped. It's a great demo. It is not a field experiment — and what happened next in the research is the part nobody tells you."

2. The decoy effect largely fails outside numeric tables — get this exactly right

Frederick, Lee & Baskin (2014), JMR 51(4): 487–507 — "The Limits of Attraction." SOLID — full primary PDF read. 38 studies total. The scoreboard, from their own General Discussion: - Fully abstract numeric-matrix stimuli: significant attraction effect in 4 of 5 studies - Numeric + perceptual/verbal representation: 2 of 5 - At least one attribute directly experienced or shown (photos, tasted drinks, real mints, popcorn, fruit): 0 of 27 — and one significant reversal ("repulsion effect"; e.g., bottled water target 70%→52% with decoy added, p<.01)

The smoking-gun manipulation: the same gambles produced a decoy effect when probability was a number (Study 2b, n=791: target 21%→37%, p<.001) and nothing when it was a pie chart (34% vs 35%). Study 3b: n=4,033 via Google Surveys, same split. Verbatim: attraction occurs "when stimuli [are] represented numerically, but not otherwise"; boundary conditions "seem to be so restrictive that its practical validity should be questioned."

Huber, Payne & Puto's own 2014 response ("Let's Be Honest About the Attraction Effect," JMR 51(4): 520–525) SOLID — concedes a lot: - The 1982 paper was "designed as a demonstration study"; "We did not set out to suggest a tool for marketing practice"; stimuli were iteratively tuned until choice sets "generated significant and replicable demonstrations." - Five inhibiting conditions (verbatim list): (1) strong prior trade-offs, (2) inability to identify the dominance relationship quickly and easily, (3) cross-respondent value heterogeneity, (4) strong dislike of the decoy, or (5) strong liking for the decoy. - "We suspect that the asymmetric dominance effect occurs rarely in the marketplace today." - Huber's own unpublished conjoint test: 586 respondents × 20 choices, ~4,000 dominated choice sets — "he could not detect any consistent increase of the target's share." - The pivot that saves your Act Three: "we more often observe successful use of compromise in the marketplace" — see §4.

Also Yang & Lynn (2014, JMR 51(4): 508–513): 91 attempts, 23 product classes → only 11 reliable attraction effects; pictures and meaningful verbal descriptions reduced it to chance. SOLID. 2020s consensus: real but fragile, flips with presentation format (Cataldo & Cohen 2019 — by-attribute layouts yes, by-alternative layouts no/reverse), killed by time pressure (Pettibone 2012). Note: a SaaS pricing page with numeric feature counts in a comparison table is closer to the format where it works; three photographed products is the format where it doesn't. That nuance is the useful takeaway.

Original for the record: Huber, Payne & Puto (1982, JCR 9(1)): six categories, all text/numeric two-attribute stimuli; average target share increase +9.2 percentage points (n=153, secondary-verified) — not the "30% sales boost" of blog legend.

3. SSN anchoring ("coherent arbitrariness") — real paper, contested replication

Ariely, Loewenstein & Prelec (2003, QJE 118(1)). 55 MIT Sloan MBA students, six products (avg retail ≈$70), incentive-compatible BDM auctions. Top vs bottom SSN quintile WTP: cordless keyboard $55.64 vs $16.09 (3.46×); correlations r = .32–.52 across products. SOLID — Table verified against the authors' own reprint. - Misquote trap: the famous "57–107% more" is above-median vs below-median, per product. The quintile comparison is the ~2.2–3.5× claim. Don't mix them. - Replication asterisk (say it): Fudenberg, Levine & Maniadis (2012, AEJ:Micro): anchor–valuation correlations −0.11 to 0.21, only one significant; quintile ratios 1.1–1.4 vs Ariely's 2.2–3.45. Maniadis, Tufano & List (2014, AER 104(1)): direct replication of the annoying-sounds anchoring, n=116 — effect ~half the original size and p = 0.253. Bergman et al. (2010, Econ Letters): found anchoring (~+45%) but smaller, decreasing with cognitive ability. Counter-critique: Data Colada #7 argues the "failure" is overstated. - The defensible line: anchoring on estimates is bulletproof (§1 below); anchoring on willingness to pay is real but probably one-half to one-third the size of the famous demo. CONTESTED as a WTP effect.

4. "Remove the dollar sign" — one lunch menu, and the overall test was not significant

Yang, Kimes & Sessarego (2009), Cornell Hospitality Report 9(8). Full report read. St. Andrew's Café (Culinary Institute of America, Hyde Park NY), lunch only, Aug–Nov 2007, final n = 201 checks, randomized by table. Three formats: "$20.00", "20.", and scripted "twenty dollars". - The overall ANCOVA effect of price format was NOT statistically significant. One linear contrast was: numeral-only parties spent $3.70 more than the restaurant average (p<.05). The widely quoted "≈8% more per person" ($23.00→$24.87) is a derived figure the authors explicitly say "was not the subject of our statistical tests." - Confound: the no-$ menu also dropped cents and rounded ("20."), so sign-removal is entangled with rounding. - The scripted "twenty dollars" format — which the authors predicted would win — tied for worst, contradicting the folk "spell it out" advice AND their own hypothesis. - CONTESTED as a finding; BLACKLISTED as a law. Safe wording: "one Cornell field study, one restaurant, 201 checks, and the headline number wasn't even the tested one."

5. Charm pricing — the real field numbers are stranger and better than the blog version

Anderson & Simester (2003, QME 1: 93–110), primary PDF read. Mail-order women's-clothing catalogs, real randomized field experiments: - Pilot: the same dress at $34 → 16 units, $39 → 21 units, $44 → 17 units. The $39 price outsold the $5-cheaper price. Across 4 dresses: 66 units at $9-endings vs 46 ($5 lower) and 45 ($5 higher) — ≈+40%, while a $10 price spread did nothing. (Tiny unit counts — say "in a small pilot.") - Full experiments: Study 1 (60,000 catalogs) ≈+35%; Study 2 (62,500) ≈+15% overall, +22% for new items, ~+10% n.s. for established items; Study 3 (270,000) ≈+7%, and a "Sale" cue alone (+0.215 coeff) beat the $9-ending effect. $9-endings added little when a Sale cue was already present. $9.50 endings did nothing or hurt. - Their conclusion: 9-endings work as an information cue ("this is a deal"), strongest when customers lack other price information — not as arithmetic hypnosis. SOLID, but report it as conditional. - When charm pricing backfires: Wadhwa & Zhang (2015, JCR 41(5)): rounded prices evaluate better for hedonic/feeling purchases (champagne at $40.00 beat $39.72/$40.28; camera-for-vacation vs camera-for-class flipped the effect; interactions p≤.004). Stiving (2000, Mgmt Sci): 0-endings signal quality — why luxury prices are round. Preregistered null: Escher et al. (2026, Frontiers, N=729): $9.99 vs $10.00 — no reliable effect on purchase intention. SOLID collectively.

6. "Good-better-best / 3 tiers is optimal" — no primary source exists

The legitimate neighbors are the compromise effect (§4 — middles gain share) and choice-model papers (Kivetz, Netzer & Srinivasan 2004 — formalizations, not tier-count prescriptions). Nothing establishes three tiers as revenue-optimal. HBR's "The Good-Better-Best Approach to Pricing" (Rafi Mohammed, 2018) is a practitioner playbook — its headline case (Allstate Your Choice Auto, 3.9M policies by 2008) is company-reported, no controlled test. BLACKLISTED as "research shows"; fine as "the compromise effect is why the pattern exists, and here's the practitioner rule of thumb, labeled as such."


PART 2 — THE VERIFIED SPINE

§1 — Anchoring: the robust foundation

Tversky & Kahneman (1974, Science 185: 1124–1131). Wheel of fortune, then estimate % of African countries in the UN: median estimates 25 vs 45 for anchors 10 vs 65. Verbatim detail worth using: "Payoffs for accuracy did not reduce the anchoring effect." SOLID — primary verified. Traps: the paper reports no n for this demo and never says the wheel was rigged (that's Kahneman's 2011 retelling).

Northcraft & Neale (1987, OBHDP 39: 84–97) — experts anchor too. Real Tucson house (appraised & listed at $74,900), 10-page MLS packet, real tours. Fake listing prices $65,900 / $71,900 / $77,900 / $83,900. - Amateurs (n=48): mean appraisals $63,571 → $72,196 across anchors (p<.001). - Professional agents (n=21, ~7 yrs experience; ±12% conditions): $67,811 vs $75,190 — a $7,379 swing from the same house (p<.01). Experiment 2 (54 amateurs, 47 experts, $134,900 house): all effects p<.001. - The denial data (the good part): 56.2% of amateurs but only 24.0% of experts admitted considering the listing price (Exp 1); verbatim, experts' checklists "flatly denied their use of listing price." Agents claimed a 5% credibility window; ±12% anchors moved them anyway. SOLID — full primary PDF verified. - ️ The circulating "92% vs 56%" denial figures are wrong. Delicious meta-detail: Northcraft & Neale's own intro misquotes the 1974 wheel numbers (swapping anchor 65 and estimate 45). Even the anchoring paper anchored wrong on the anchoring paper.

Replication status — Many Labs 1 (Klein et al. 2014, Social Psychology 45): anchoring is arguably the best-replicated effect in psychology. 36 samples, N=6,344. Anchoring took 4 of the top 5 effect sizes in the whole project: babies born/day d = 2.42, Mt. Everest d = 2.23, Chicago d = 1.79, SF→NYC d = 1.17; 94–100% of individual samples significant; replications came out larger than the originals. SOLID. Honest caveat: these are anchors on factual estimates, not prices — pair with §2's WTP asterisk.

Incidental anchors — Critcher & Gilovich (2008, JBDM 21): restaurant named "Studio 97" vs "Studio 17" → diners would spend $32.84 vs $24.58 (p=.03); jersey #94 vs #54 shifted judged probabilities. CONTESTED — real paper, but n≈200 per study and p-values .03–.05. Say "suggestive."

The purchase-limit soup study — usable only with a flag. Wansink, Kent & Hoch (1998, JMR 35(1)): 3 Iowa supermarkets, "Campbell's 79¢" end-cap, 906 shoppers analyzed. Cans bought: no limit 3.32 · "limit 4" 3.58 · "limit 12" 7.0 (p<.01) — the arbitrary "12" more than doubled purchases. Verified: this paper is NOT retracted (Crossref clean), predates the Food and Brand Lab era, and has heavyweight co-authors (Hoch). But lead author Wansink had ~13–18 retractions and resigned from Cornell after a misconduct finding. CONTESTED — if used, disclose the Wansink situation on camera; never replicated at scale.

Reference prices ("was $X, now $Y"): Urbany, Bearden & Weilbaker (1988, JCR 15(1)): plausible advertised reference prices raised perceived value — and exaggerated, implausible ones had generally the same positive effects, even among skeptics. SOLID direction (keep qualitative). This is the peer-reviewed basis for strikethrough pricing.

§2 — Coherent arbitrariness (see Correction 3 for the full treatment)

Two bonus details, both verified: (a) coherence is the real finding — >95% of subjects valued the keyboard > trackball and rare wine > average wine; the level was arbitrary, the structure was orderly. Demand curves look rational even when built on sand. (b) A companion study with 77 MIT executives: SSN–bid correlations ~.29 (p<.01), and 72% claimed the anchor didn't influence them — pair with Northcraft's denying realtors for a "you won't feel it working" beat. SOLID.

§3 — Decoy / asymmetric dominance (see Correction 2 — that IS this section)

§4 — The compromise effect: the one that actually holds up

Simonson (1989, JCR 16(2)) — the original. Brands gain share when they become the middle option. Verified table data (via S&T 1992 reprint): calculator batteries — brand share 43% → 60% when a more extreme option was added, even when the extreme option was marked unavailable (ns 119–126). SOLID. ️ The freestanding "+17.5% average" figure circulates from secondary sources only — I could not verify it in the primary. Don't use it.

Simonson & Tversky (1992, JMR 29(3): 281–295) — full primary PDF read. 22 experiments, n=100–220 each, real color pictures from a merchandise catalog (note: pictures — and the effects still appeared, unlike the decoy). The Minolta camera set, exact: - Two options (n=106): X-370 $169.99 50% / Maxxum 3000i $239.99 50% - Add Maxxum 7000i $469.99 (n=115): 22% / 57% / 21% — the mid camera's share vs the cheap one went from 50% to 72% (p<.01) - Adding a premium option didn't sell the premium option — it sold the middle. SOLID - Extremeness aversion, verbatim: "the attractiveness of an option is enhanced if it is an intermediate option in the choice set and is diminished if it is an extreme option."

Replication status — this is the robust one. Neumann, Böckenholt & Sinha (2016, JCP 26(2)): meta-analysis, 142 observations — extremeness aversion is robust; the same product is 29% to 88% more likely to be chosen when intermediate than when extreme. SOLID. Plus Huber/Payne/Puto's 2014 line: compromise, not attraction, is what marketers successfully use. Boundaries: weakens with larger sets and by-alternative formats — it's a tilt (share shifts), not magic.

Script architecture note: this is the clean arc — the famous effect (decoy) is fragile; the boring cousin (compromise) is the one that survives meta-analysis and is why good-better-best exists.

§5 — Center-stage effect: real, small, thin

Valenzuela & Raghubir (2009, JCP 19(2)): center gum chosen 50% vs 29.2% (left) vs 20.8% (right) — but n=48. Mechanism: people infer the middle option is the popular one. Raghubir & Valenzuela (2006, OBHDP 99): The Weakest Link contestants in the two center positions reached the final 42.5% and won 45% of the time vs 17.5% / 10% for the extremes. Rodway et al. (2012): center preference replicates in horizontal and vertical arrays (incl. real socks). CONTESTED overall — a handful of labs, modest ns, moderated by array size and goal. ️ There is no Wimbledon study — that's a false memory in circulation; it's a TV game show. Safe use: "middle placement gets a measurable tilt in small studies" — supports putting the plan you want sold in the center, weakly.

§6 — Left-digit / charm pricing (with Correction 5)

Thomas & Morwitz (2005, JCR 32(1)) — the mechanism. $2.99 vs $3.00 changes perceived magnitude; $3.59 vs $3.60 does not (left digit unchanged: F<1). Study 1b: magnitude ratings 35.6 vs 55.8 (p<.01) for 2.99/3.00; nothing for 2.79/2.80 or 3.19/3.20. Effect dies when the comparison price is far away ($2 distance: F<1). Latency data: "4.99" judged faster than "5.00" against a $5.50 standard — it genuinely feels farther below. SOLID — primary PDF read. Small lab samples (27–154); perception, not sales — which is why A&S (Correction 5) and the next entry matter.

Strulov-Shlain (2023, Review of Economic Studies 90(5): 2612–2645) — the modern estimate. NielsenIQ scanner data: 25 US chains, ~3,500 products, 78M observations, $4.3B in sales, 2006–2019. In-sample: 41% of prices end .99; 87% end in 9. - Bias parameter θ ≈ 0.2: consumers respond to a 1-cent increase that crosses a dollar ($4.99→$5.00) "as if it were more than a twenty-cent increase." A within-dollar cent is perceived as ~0.8¢. - The twist: firms are the irrational ones. Given θ=0.2, essentially all prices should end in 99 — firms use it on only 30–40%, pricing as if θ were 0.01–0.03. All 25 chains underestimate the bias. - Retailers forgo an estimated 1–4% of potential gross profits. SOLID — published figures verified. - Trap: "consumers see $4.99 as $4" is too strong — the model is a weighted blend (~$4.60–4.80 perceived).

§7 — Price presentation micro-effects

Effect Finding Source Confidence
Precise prices seem smaller $395,425 judged smaller than $395,000; in >27,000 real-estate transactions (S. Florida + Long Island), round-listed homes sold ≈0.73% lower than precise-listed comparables (~$3,600 on $500k) Thomas, Simon & Kadiyali (2010), Marketing Science 29(1): 175–190 SOLID (0.73% via reputable secondary)
Precise anchors constrain adjustment Wholesale-cost guesses adjust much farther from $5,000 than from $4,988 or $5,012 — precise anchors evoke a finer mental scale Janiszewski & Uy (2008), Psych Science 19(2) SOLID concept; per-study means paywalled — don't quote exact deltas
Precise first offers anchor negotiations harder $4,983 offers drew more conciliatory counteroffers than $5,000 Mason et al. (2013), JESP 49 SOLID, lab; can backfire vs experts (Loschelder)
Font size Sale price in smaller font than regular price → lower perceived magnitude, higher purchase likelihood Coulter & Coulter (2005), JCP 15(1), 3 experiments SHAKY — one team, lab-only, later boundary/reversal papers (J. Retailing 2022). Not a law.
Syllables / commas "$1,499.00" vs "$1499" — longer verbal encoding → larger perceived magnitude, even reading silently. Strip commas and ".00" Coulter, Choi & Monroe (2012), JCP 22(3) SOLID citation, lab-only
Currency sign See Correction 4 — overall test n.s. Yang, Kimes & Sessarego (2009) CONTESTED/BLACKLISTED as law

§8 — Partitioned & drip pricing: the strongest field evidence in the whole space

** Blake, Moshary, Sweeney & Tadelis (2021), Marketing Science 40(4) — the StubHub experiment. Aug 19–31, 2015; several million desktop users, ~50/50 randomized. Control saw all-in prices while browsing; treatment saw base prices with ~15% buyer fees revealed only at checkout. All numbers verified against the NBER WP (25186) tables: - Drip-fee users spent +20.64% (SE 1.38) — "about 21% more" - +14.1% more likely to purchase at least once; +13.24% more transactions - Conditional on buying, paid +5.42% more per order — they bought better tickets, not just more ("quality upgrade effect"; ≥28% of the revenue gain) - The funnel tells the story: drip users clicked listings ~19% more, then were ~45% more likely to bail at the final checkout stage when fees appeared — and still spent 21% more overall - Experienced users (10+ prior visits) still spent ~15% more. Experience doesn't immunize. - StubHub switched the whole platform to back-end fees on Sept 1, 2015. Sellers responded by listing higher-quality inventory (row-A listings +15%). SOLID - Framing trap: this is a revenue result, not a welfare-positive one — the paper is about how hiding fees distorts choice. It's simultaneously the best evidence that drip pricing "works" and the reason it's now illegal in key categories (below).

Morwitz, Greenleaf & Johnson (1998, JMR 35(4)) — the original partitioned-pricing paper. Study 1: incentivized auction for a jar of pennies; a 15% buyer's premium group paid more in total — bidders anchored on the bid and under-adjusted. Study 2: telephone survey, catalog offer. Verified via Morwitz's own 2016 JCP review: 12.2–35.6% of consumers ignored the surcharge completely when recalling total price; in one condition only 21.9% actually calculated the total (54.8% used a heuristic, 23% ignored). SOLID design/direction; exact cell means paywalled. Trap: it was a lab auction + phone survey, not a "phone auction."

Hossain & Morgan (2006, BE J. Econ Analysis & Policy) — eBay shipping. 80 real auctions (CDs, Xbox games). Low opening bid + $3.99 shipping beat $4 opening + free shipping: CDs +~21% revenue, 9/10 matched pairs; pooled 16/20 pairs (p=.005). High-reserve Xbox: hidden-fee variant won 10/10 pairs (+11%). Boundary: the effect vanished when the total reserve hit ~53% of retail — shrouding fails when the fee is too large relative to value. Follow-up: Brown, Hossain & Morgan (2010, QJE 125(2)): raising shipping boosts revenue especially when hidden; disappears when shipping is displayed on the search page. SOLID - Nuance for fairness: Bertini & Wathieu (Mktg Sci 2008) — partitioning can raise or suppress demand depending on which component gets attention. Not uniformly demand-increasing.

Regulation (dates verified): FTC "Junk Fees Rule" (16 CFR 464): announced Dec 17, 2024 (4–1 vote), effective May 12, 2025 — live-event tickets and short-term lodging must display total price including mandatory fees upfront. California SB 478 economy-wide since July 1, 2024. ️ The DOT airline ancillary-fee rule was vacated by the Fifth Circuit — do NOT cite it as in force. SOLID

§9 — Temporal reframing: per-day, monthly, annual

(Coordinate with the paywall brief — covered here from the display angle only.)

Gourville (1998, JCR 24(4)) — pennies-a-day. Charity payroll deduction: "85 cents per day" → 52% agreed; "$300 per year" → 30% — and note the sleight: $0.85/day ≈ $310/yr, more money, and it still won. Mechanism: per-day framing retrieves trivial comparisons (coffee); aggregate framing retrieves big ones. SOLID (headline verified via multiple academic citations; original paywalled). Boundary (Gourville 2003, Marketing Letters): backfires when the daily number stops being trivial — people prefer monthly framing for rent at "$25/day" or taxes at "$58/day." - ️ Trap: the "$1/day vs $350/year" version is the Atlas & Bartels (2018) stimulus, not Gourville's.

Atlas & Bartels (2018, JCR 45(2)) — 8 experiments + field test: periodic pricing raises intentions partly by increasing perceived benefits (not just shrinking perceived cost), and extends beyond trivial amounts in some contexts. SOLID.

** Hershfield, Shu & Benartzi (2020, Marketing Science 39(6)) — the modern field result. Fintech savings app (Acorns): framing the identical recurring deposit as "$5/day" vs "$150/month" roughly quadrupled enrollment (~30% vs ~7%) — and the daily frame eliminated the participation gap between lowest- and highest-income users. SOLID. This is your best per-day-framing number: recent, field, peer-reviewed.

Flat-rate bias — Lambrecht & Skiera (2006, JMR 43(2)). 10,882 DSL customers: 48.1% exhibited flat-rate bias vs 8.5% pay-per-use bias; over half of the flat-rate-biased paid ≥100% more than their cheapest tariff — knowingly, happily (insurance + taxi-meter effects, confirmed), and it did NOT increase churn. Pay-per-use bias did. SOLID. The honest case for why unlimited plans print money and customers don't resent them.

"$X/mo billed annually" specifically: no direct peer-reviewed test exists. It's an untested extrapolation from PAD/periodic-pricing research. Label it as such.

§10 — Price ordering: high-to-low vs low-to-high

Suk, Lee & Lichtenstein (2012, JMR 49(5)) — the actual evidence for "anchor high first." 8-week field experiment in a bar, 13 beers: descending menu (expensive first) → average paid $6.02 vs ascending → $5.78. +$0.24/beer, ~4%, significant. Mechanism: high-to-low sets a high reference; trading down feels like a quality loss. SOLID (figures via reputable secondaries of the JMR paper; venue/direction confirmed in abstract). That's the entire real basis for "show your expensive plan first" — a 4% nudge in one bar plus labs. Not nothing; not a law.

§11 — SaaS practitioner data: vendor, not research

Price Intelligently → ProfitWell (Patrick Campbell) → acquired by Paddle, 2022 (~$200M). Major citability problem: old profitwell.com data posts now redirect to Paddle pages where the underlying data is often gone.

Claim in circulation Reality Confidence
"Companies spend only 6–8 hours on pricing" Campbell's stated claim is ~8 hours over the company's lifetime — self-reported survey of PI's audience, methodology never published SHAKY — attribute to Campbell, never "research shows"
"Value-based pricing lifts revenue 30–40%" Traces to PI's own service pitch laundered into a finding BLACKLISTED
Monetization beats acquisition 2–4× per 1% improvement ProfitWell panel analysis (~23.4k companies); original posts largely dead links Vendor-grade CONTESTED
Annual plans reduce churn ProfitWell panel (~2,500 companies), direction consistent; content now redirected Vendor-grade CONTESTED
Freemium conversion benchmarks Cite OpenView (2–5%) and ChartMogul (3–5% good, 8–12% great) instead of "ProfitWell says" CONTESTED (vendor, but published)
"Revisit pricing quarterly → grow faster" No primary publication located BLACKLISTED — conference lore

PART 3 — HOOK CANDIDATES (verified, striking)

  1. The StubHub number: "StubHub randomly hid its ~15% fees until checkout for half of several million visitors. Those people spent 21% more — and bought better seats. Two weeks later StubHub switched the whole site. Ten years later, that pricing style is illegal for ticket sites." (§8, SOLID)
  2. The $39 dress: "A catalog sold the same dress at $34, $39, and $44. The $39 version outsold the $34 version. Charging five dollars MORE sold more units." (Correction 5, SOLID — flag small pilot; back with the 60k–270k-catalog studies)
  3. The expert anchor: "Real-estate agents toured a real house. A fake listing price moved their professional appraisals by $7,379 — and their written reports 'flatly denied' using the listing price at all." (§1, SOLID)
  4. The penny that costs twenty cents: "Across 78 million supermarket price observations, crossing from $4.99 to $5.00 hits demand like a 20-cent increase — and the punchline is that retailers underprice this, leaving 1–4% of profit on the table." (§6, SOLID)
  5. The decoy scoreboard: "In 27 studies where people could actually see or taste the products, the famous decoy effect worked zero times." (Correction 2, SOLID)
  6. $5/day: "Reframing $150/month as $5/day quadrupled savings-plan enrollment — 7% to 30% — in a real fintech field experiment." (§9, SOLID)

PART 4 — NOVEL-ANGLE CANDIDATES

  • "The famous effect is the fragile one." Decoy (famous) barely survives contact with real products; compromise (obscure) survives a 142-observation meta-analysis and is what the market actually uses. Almost no creator gets this right — it's the video's spine.
  • "You won't feel it working." Three-way sourced denial pattern: Northcraft's experts flatly denied using the anchor; 72% of Ariely's executives claimed no influence; T&K's subjects weren't fixed by accuracy payoffs. Introspection is not a defense.
  • "The irrational actor is the firm." Strulov-Shlain's inversion — consumers' left-digit bias is real, and retailers systematically misprice around it. Fresh take vs. the tired "brains are broken" framing.
  • Drip pricing's quality-upgrade effect: hiding fees doesn't just extract more money — it changes what people buy (better seats, +5.4%/order). Price display shapes product choice, not just conversion.
  • The regulation arc: 1998 lab paper → 2015 platform-scale experiment → 2024 FTC rule. A rare complete story of psychology research becoming law.
  • Flat-rate bias as the honest ending: half of subscribers overpay ≥100% on flat plans, don't churn, and are happier — the uncomfortable case that some "exploitation" is what customers actively want (insurance against the taxi-meter feeling).
  • The repulsion effect: badly executed decoys backfire — a moldy orange next to your target makes people flee the category. Practical stakes for copying "add a decoy tier" advice blindly.

PART 5 — THE BLACKLIST

Never say these on camera:

Claim Why
"The Economist tripled subscriptions with a decoy" Nobody measured Economist conversions. Classroom demo, n=100 MBA students.
"The decoy effect boosts sales ~30%" 1982 original: +9.2 share points, text stimuli tuned to produce the effect.
"The decoy effect was debunked" Also wrong — it replicates with numeric attribute tables (4/5); it fails with perceptual stimuli (0/27). Format-dependent.
"Removing the dollar sign makes people spend 8% more" Overall test n.s.; one contrast p<.05; one restaurant, 201 checks; confounded with rounding.
"Prices ending in 9 always win" Conditional: new items yes, Sale-cue present barely, hedonic/luxury can reverse, $9.50 hurts, preregistered null exists.
"Research shows 3 tiers is optimal" No primary source. Compromise effect ≠ tier-count prescription.
"Consumers read $4.99 as $4" Model says perceived ≈$4.60–4.80; the 20¢ figure applies only at dollar crossings.
"People bid their SSN" without the replication asterisk AER/AEJ replications: half the size to nothing on WTP.
"57–107% more" as the SSN quintile gap That's above/below median. Quintiles = 2.2–3.5×.
"92% of amateurs vs 56% of experts denied the anchor" Real figures: 56.2% vs 24.0% mentioned it (Exp 1).
The Wimbledon center-stage study Doesn't exist. It's The Weakest Link (game show), 42.5%/17.5%.
Soup study without the Wansink disclosure Paper not retracted, but lead author had ~13–18 retractions; disclose or drop.
"Gourville: $1/day vs $350/year" That stimulus is Atlas & Bartels 2018. Gourville: $0.85/day vs $300/yr, 52% vs 30%.
"Value-based pricing lifts revenue 30–40% (research)" It's Price Intelligently's service pitch.
"Companies spend 6–8 hours/year on pricing" Campbell's claim is 8 hours over the company's lifetime, unpublished survey.
DOT airline fee rule as current law Vacated by the Fifth Circuit. FTC rule (tickets/lodging) and CA SB 478 are the live ones.
"The middle option gets chosen 66% of the time" (any universal center %) Center-stage flagship result is n=48 gum choice (50%).
"+17.5% compromise-effect average" Secondary-source figure; not verified in Simonson 1989. Use the 43→60% or camera 50→57% instead.

APPENDIX — AT-A-GLANCE VERDICTS

Tactic on the pricing page Verdict Best evidence
Show an expensive plan/anchor first Real, small (~4% in the one field test) Suk et al. 2012 bar study
Strikethrough "was/now" reference prices Works, even when implausible Urbany et al. 1988
Decoy tier (asymmetric dominance) Only in numeric comparison-table formats; fails/reverses with rich stimuli FLB 2014; HPP 2014
Compromise / middle tier Robust — the workhorse effect Neumann et al. 2016 meta (142 obs)
Center placement of featured plan Weak tilt, thin literature Valenzuela & Raghubir 2009 (n=48)
Charm pricing (.99 / $X9) Field-verified but conditional; can backfire on premium/hedonic A&S 2003; Strulov-Shlain 2023; Wadhwa & Zhang 2015
Precise (non-round) prices Seem smaller; 27k-transaction field support Thomas, Simon & Kadiyali 2010
Small price font One lab team; not a law Coulter & Coulter 2005
Drop "$", commas, ".00" "$"-removal overclaimed (n.s. overall); comma/cents-stripping has lab support Yang et al. 2009; Coulter et al. 2012
Drip/partitioned fees The strongest effect in the space (+21% spend) — and now regulated Blake et al. 2021; FTC 2025
Per-day reframing ("$5/day") Strong, field-replicated; backfires for large daily amounts Hershfield et al. 2020; Gourville 2003
"$X/mo billed annually" No direct peer-reviewed test — practitioner extrapolation
Exactly 3 tiers Myth as a research claim

PART 6 — PRIMARY SOURCES

Anchoring & coherent arbitrariness Tversky & Kahneman (1974) — https://www.cs.tufts.edu/comp/150AIH/pdf/TverskyKa74.pdf Northcraft & Neale (1987) — https://www.smallprojectsbureau.com/wp-content/uploads/2020/01/northcraft_neale.pdf Klein et al., Many Labs 1 (2014) — https://stanford.edu/~knutson/jdm/klein14.pdf Critcher & Gilovich (2008) — http://static1.1.sqspcdn.com/static/f/409296/3798405/1249679424930/Critcher_Anchoring.pdf Wansink, Kent & Hoch (1998) — https://journals.sagepub.com/doi/10.1177/002224379803500108 Ariely, Loewenstein & Prelec (2003) — https://web.mit.edu/ariely/www/MIT/Chapters/CA.pdf (QJE: https://academic.oup.com/qje/article/118/1/73/1917051) Fudenberg, Levine & Maniadis (2012) — https://www.aeaweb.org/articles?id=10.1257/mic.4.2.131 Maniadis, Tufano & List (2014) — https://www.aeaweb.org/articles?id=10.1257/aer.104.1.277 Data Colada #7 (counter-critique) — https://datacolada.org/7 Urbany, Bearden & Weilbaker (1988) — https://academic.oup.com/jcr/article-abstract/15/1/95/1840979

Decoy, compromise, center-stage Huber, Payne & Puto (1982) — DOI 10.1086/208899 Frederick, Lee & Baskin (2014) — https://web2-bschool.nus.edu.sg/wp-content/uploads/media_rp/publications/tF83O1430805722.pdf Huber, Payne & Puto (2014) — https://people.duke.edu/~jch8/bio/Papers/HuberPaynePutoJMR%202014.pdf Yang & Lynn (2014) — DOI 10.1509/jmr.14.0020 Simonson (1989) — DOI 10.1086/209205 Simonson & Tversky (1992) — https://cognition.aau.at/bg/BA/Simon%20&%20tversky,%201992.pdf Neumann, Böckenholt & Sinha (2016) — DOI 10.1016/j.jcps.2015.05.005 Valenzuela & Raghubir (2009) — DOI 10.1016/j.jcps.2009.02.011 Raghubir & Valenzuela (2006) — OBHDP 99(1), 66–80 Rodway, Schepman & Lambert (2012) — DOI 10.1002/acp.1812 Mohammed, HBR good-better-best (2018) — https://hbr.org/2018/09/the-good-better-best-approach-to-pricing

Left-digit, precision, presentation Thomas & Morwitz (2005) — archived: web.archive.org copy of forum.johnson.cornell.edu/faculty/mthomas/LeftDigitEffect.pdf Anderson & Simester (2003) — https://www.kellogg.northwestern.edu/faculty/anderson_e/htm/personalpage_files/Papers/Effects_of_9_Price_Endings_on_Retail_Sales.pdf Strulov-Shlain (2023) — https://academic.oup.com/restud/article-abstract/90/5/2612/6931812 (PDF: https://gwern.net/doc/economics/2023-strulovshlain.pdf) Wadhwa & Zhang (2015) — https://www.smallprojectsbureau.com/wp-content/uploads/2020/01/wadhwa-zhang-2015.pdf Stiving (2000) — Management Science 46(12), 1617–1629 Escher et al. (2026, preregistered null) — https://www.frontiersin.org/journals/behavioral-economics/articles/10.3389/frbhe.2026.1828446/full Janiszewski & Uy (2008) — https://journals.sagepub.com/doi/10.1111/j.1467-9280.2008.02057.x Thomas, Simon & Kadiyali (2010) — https://pubsonline.informs.org/doi/10.1287/mksc.1090.0512 Mason et al. (2013) — https://www.sciencedirect.com/science/article/abs/pii/S0022103113000401 Coulter & Coulter (2005) — https://myscp.onlinelibrary.wiley.com/doi/abs/10.1207/s15327663jcp1501_9 Coulter, Choi & Monroe (2012) — https://myscp.onlinelibrary.wiley.com/doi/abs/10.1016/j.jcps.2011.11.005 Yang, Kimes & Sessarego (2009) — https://ecommons.cornell.edu/bitstreams/d9504484-4912-4291-a65c-f5b44461302b/download

Partitioned/drip pricing, framing, ordering Morwitz, Greenleaf & Johnson (1998) — DOI 10.1177/002224379803500404 (2016 review w/ figures: Morwitz et al., JCP 26(1), Wharton-hosted PDF) Blake, Moshary, Sweeney & Tadelis (2021) — DOI 10.1287/mksc.2020.1261 (WP: https://www.nber.org/papers/w25186) Hossain & Morgan (2006) — https://faculty.haas.berkeley.edu/rjmorgan/ebay.pdf Brown, Hossain & Morgan (2010) — QJE 125(2), 859–876 FTC Junk Fees Rule — https://www.federalregister.gov/documents/2025/01/10/2024-30293/trade-regulation-rule-on-unfair-or-deceptive-fees California SB 478 — https://oag.ca.gov/hiddenfees Gourville (1998) — JCR 24(4), 395–408; Gourville (2003), Marketing Letters 14(2) Atlas & Bartels (2018) — JCR 45(2), 350–367 Hershfield, Shu & Benartzi (2020) — Marketing Science 39(6) Lambrecht & Skiera (2006) — JMR 43(2), 212–223 (authors' PDF via uni-frankfurt.de) Suk, Lee & Lichtenstein (2012) — JMR 49(5), 708–717

SaaS practitioner (vendor-grade) OpenView Product Benchmarks; ChartMogul SaaS Conversion Report; Paddle/ProfitWell (note: pre-2022 data posts largely dead-linked)