Research Brief — Decision Fatigue & Choice in App Forms

Fact-check reference for the YouTube script

Confidence key SOLID — primary source located, figure verified CONTESTED — real source exists, but methodology or interpretation is disputed SHAKY — widely repeated, source is weak, absent, or circular BLACKLISTED — never repeat this claim


️ PART 1 — CORRECTIONS TO YOUR ORIGINAL BRIEF

Six of your talking points needed changing. In every case the corrected version is either more accurate and more powerful, or has been replaced by a stronger stat.

1. The jam study — three errors

You said Reality
"30 jams vs 6 jams" 24 vs 6. The "30" is from the other two experiments in the same paper (30 essay topics, 30 chocolates). You merged them.
"30% higher sales" The 30% is a purchase rate, not a lift. ~30% of the 6-jam group bought vs ~3% of the 24-jam group — roughly a 10x ratio. Your version understates the paper's own finding by ~30x.
(implied) settled science The effect has largely failed to replicate. See §2 below.

Exact figures — Iyengar & Lepper (2000), JPSP 79(6): 995–1006. Draeger's Supermarket, Menlo Park, CA. ~754 shoppers observed, 502 passed a display.

6 jams 24 jams
Stopped at booth 104/260 = 40.0% 145/242 = 59.9%
Purchased (of stoppers) 31/104 = 29.8% 4/145 = 2.8%
Purchased (of all passersby) 31/260 = 11.9% 4/242 = 1.7%
Jams actually sampled 1.38 1.50

SOLID — https://faculty.washington.edu/jdb/345/345%20Articles/Iyengar%20&%20Lepper%20(2000).pdf

The sampling detail is the single best insight in the whole study and almost nobody uses it. Both groups tasted the same number of jams. Nobody explored the big assortment and got tired — the overload was anticipatory. People priced the job by looking at it and left. That maps onto app UI far better than the standard telling.

Also note the inconvenient half: the 24-jam display attracted 50% more foot traffic. Big assortments are better at acquisition, worse at conversion. Worth acknowledging — it makes you sound like you actually read the paper.


2. "70–90% of people never change defaults" — no source exists

I could not find any study producing that range. It circulates in design blogs unattributed. The real numbers are considerably stronger:

Figure Source Confidence
1.2% of 4,540 Facebook profiles changed the default searchability setting; 0.06% restricted profile visibility Gross & Acquisti, ACM WPES '05 — https://www.heinz.cmu.edu/~acquisti/papers/privacy-facebook-gross-acquisti.pdf SOLID
"Less than 5%" of surveyed MS Word users had changed any setting Jared Spool, UIE, 2011 — https://archive.uie.com/brainsparks/2011/09/14/do-users-change-their-settings/ CONTESTED — informal convenience sample, never peer-reviewed. Cite as industry lore, not science.
61% of auto-enrolled 401k employees stayed on both the default rate and default fund — a combination only ~1% of self-selecting employees ever chose Madrian & Shea, QJE 116(4), 2001 — https://www.nber.org/papers/w7682 SOLID

Use Madrian & Shea. It's peer-reviewed, the contrast is dramatic, and it demonstrates something stronger than "people don't change defaults" — it shows the default becomes the choice.


3. "Some Netflix users spend more time browsing than watching" — no source

I searched hard. Nothing in Nielsen, Netflix's publications, or any credible outlet supports browse-time exceeding watch-time. BLACKLISTED.

Replacements, all sourced:

Figure Source Confidence
10.5 minutes per session deciding what to watch; 1 in 5 abandoned the session entirely; catalogue 2.7M titles Nielsen State of Play, Aug 2023 — https://www.nielsen.com/news-center/2023/nielsens-state-of-play-report-delivers-new-insights-as-streamings-next-evolution-brings-content-discovery-challenges-for-viewers/ SOLID
7 min 24 sec average (your recollection — confirmed, but it's the 2019 figure) Nielsen Total Audience Report Q1 2019 SOLID
~80% of hours streamed come from recommendations, 20% from search Gomez-Uribe & Hunt, ACM TMIS 6(4), 2015 — https://dl.acm.org/doi/10.1145/2843948 SOLID — Netflix's own VP of Product and CPO, peer-reviewed

Do not use McKinsey's "75% of Netflix viewing" or "35% of Amazon purchases" — both trace to a single uncited 2013 McKinsey sentence. BLACKLISTED.

️ Netflix's "60–90 seconds" line is real but hedged with "perhaps," has no methodology, and says users lose interest — not that they leave. Attribute it if used; don't assert it.


4. Trader Joe's — soften the claims

You said Reality
"40,000+ items at other stores" FMI's measured figure is 33,248 (2025). The 40–50k numbers are journalistic shorthand from Fortune/CNN, not FMI.
"Highest revenue per square foot" Probably true among grocers — but every TJ's input is an estimate (privately held, no filings) and Costco is in the same band at $1,803–$2,043, from an audited 10-K.
Metric Figure Confidence
Trader Joe's SKUs ~4,000 (range in circulation: 2,000–4,500; never company-confirmed) CONTESTED — say "roughly"
Average US supermarket items 33,248 (2025); 31,795 (2024) SOLID — FMI, https://www.fmi.org/our-research/food-industry-facts
Walmart supercenter 100,000+ CONTESTED (CNN 2019)
Aldi ~1,400–2,000, >90% private label CONTESTED (2019 figures, likely higher now)
Average supermarket sales/sq ft $822 (computed: FMI $668,377/wk × 52 ÷ 42,272 sq ft) SOLID
Costco sales/sq ft $2,043 total revenue · $1,803 ex-gasoline SOLID — FY2025 10-K
Trader Joe's sales/sq ft ~$2,200–2,700 estimated CONTESTED — all inputs are estimates
TJ's private label share ~80% CONTESTED (Fortune 2010, confirmed by execs on TJ's own podcast)

Safe wording: "Trader Joe's does somewhere north of two thousand dollars per square foot — roughly three times the average American supermarket. Costco is the only grocer in the same league. And because Trader Joe's is private, every one of those numbers is an outside estimate."

Avoid: "Highest-revenue grocer per square foot in America, full stop."

Also worth knowing: curation is the enabling condition, not the cause. The real drivers are ~80% private label (captures the manufacturer's margin), ~$5M of annual sales per SKU vs ~$600K at a 33k-SKU chain (enormous supplier leverage), no coupons/loyalty/slotting fees, no e-commerce, and a store one-third the average size. Attributing TJ's success purely to choice psychology would be the intellectual sloppiness you said you wanted to avoid.


5. Head & Shoulders "26 → 15 varieties, +10% sales" — BLACKLISTED

Not in your brief, but it's the stat everyone reaches for here and you'd likely have hit it. No primary source exists. It circulates via Inc. and TED talks with zero citation, and is absent from the Wharton, Kellogg, and Boatwright literature reviews — which is telling. The Inc. headline says profits while the body says sales. Unstable in its own retelling.

You don't need it. Boatwright & Nunes is peer-reviewed and stronger — see §8.


6. "Decision fatigue" as a mechanism — the framing problem

Your video is titled around it, so here's the honest state of the science:

Claim Confidence
Willpower is a depletable resource (ego depletion) SHAKY — Hagger et al. 2016 (23 labs, N=2,141): d = 0.04, CI includes zero. Vohs et al. 2021 (36 labs, N=3,531): d = 0.06, non-significant — and Vohs is one of the original proponents.
Danziger parole-judges study ("hungry judges") SHAKY — Weinshall-Margel & Shapard showed case ordering isn't random (unrepresented prisoners go last and win 15% vs 35%). Glöckner showed a purely rational scheduling judge reproduces the same curve with zero depletion.
Users abandon flows with many hard decisions in a row SOLID as a product observation — supported by funnel data and the Chernev moderators, and it does not require ego depletion to be true

Recommendation: keep "decision fatigue" as the title and the descriptive label — it's the term your audience knows and it accurately names the phenomenon. But build the mechanism on extraneous cognitive load and the Chernev moderators, both of which are solid and give identical practical advice. The script does exactly this, and the Act Two admission converts the weakness into your strongest credibility moment.


PART 2 — THE VERIFIED SPINE

Everything below is stated in the script and checks out.

§1 — Choice overload: the honest picture

The replication problem — Scheibehenne, Greifeneder & Todd (2010), JCR 37(3): 409–425. 63 data points, 50 experiments, 5,036 participants. Mean effect D = 0.02, 95% CI [−0.09, 0.12]. Publication bias detected. Jam study specifically failed to replicate in a German supermarket (Scheibehenne 2008). SOLID — https://scheibehenne.com/ScheibehenneGreifenederTodd2010.pdf

Schwartz's own concession — PBS NewsHour, 29 Jan 2014: "Does choice overload always occur? Of course not. Does it affect all people, in all domains of decision making? Of course not." And: "we don't yet know what factors determine when choice overload will occur and when it won't." SOLID — https://www.pbs.org/newshour/economy/is-the-famous-paradox-of-choic

The rescue — Chernev, Böckenholt & Goodman (2015), JCP 25(2): 333–358. 99 observations, 53 studies, 7,202 participants. Without moderators the effect is non-significant; with moderators it's significant (b = .17, p < .001). SOLID — https://chernev.com/wp-content/uploads/2017/02/ChoiceOverload_JCP_2015.pdf

Moderator Estimate Maps to signup as
Effort-minimizing decision goal .56 They came to use the app, not configure it
Choice set complexity .55 They can't tell what your options mean
Decision task difficulty / time pressure .37 They're doing this in a spare 90 seconds
Preference uncertainty .32 They've never used your product

Plus: a dominant option makes overload disappear. That's the empirical case for a preselected default, and it's the cleanest bridge from Act Two to Act Three.

Barry SchwartzThe Paradox of Choice, 2004 (rev. 2016 w/ Cheek). Maximizer vs satisficer from Schwartz et al. (2002), JPSP 83(5): 1178–1197. His four mechanisms (paralysis, opportunity cost, escalated expectations, self-blame) are CONTESTED — say "Schwartz argues," not "research shows."

Hick's Law — handle with care. SHAKY as a general UI rule. Applies to simple reaction-time tasks with equiprobable options; breaks down above ~8 options (Seibel 1963: 1,023 alternatives were only 20–30ms slower than 31); erased by practice and high stimulus-response compatibility; doesn't apply at all to scanning an unsorted list (that's linear). Ironic footnote: the best-validated HCI application of Hick's law (Landauer & Nachbar 1985) argues for broad, shallow menus — i.e. more options per screen. UX writers citing it to justify hiding options are citing a result that points the other way. The script deliberately omits Hick's Law. — https://web.ics.purdue.edu/~dws/pubs/ProctorSchneider_2018_QJEP.pdf

Miller's 7±2 — do not use as a UI limit. SHAKY. Miller himself called the correspondence "only a coincidence" and said there's "nothing magical about the number seven." NN/g: "often misunderstood" — and decisively, users don't need to hold menu items in short-term memory because the options are visible on screen. That's recognition, not recall. Cowan (2001) puts the real working-memory figure at ~4 chunks (3–5 population average), and only when rehearsal is blocked. — https://www.nngroup.com/articles/chunking/

Cognitive Load Theory — the safest frame in the video. Sweller. Intrinsic load = complexity actually in the task. Extraneous load = created by how it's presented. ️ Germane load was removed as a separate category in Sweller, van Merriënboer & Paas (2019) — many UX articles still teach the outdated three-load model. Use two. SOLID — https://leadinglearner.me/wp-content/uploads/2019/02/sweller2019_article_cognitivearchitectureandinstru.pdf


§2 — Least effort & how users actually behave

Zipf (1949), Human Behavior and the Principle of Least Effort. The correction that makes this beat work: it is not "people are lazy." Verbatim — a person "will strive to solve his problems in such a way as to minimize the total work that he must expend in solving both his immediate problems and his probable future problems." Technically: least average rate of probable work. People are optimizing over an expected future workload. SOLID

Steve Krug, Don't Make Me Think, Ch. 2 — all three verbatim: "We don't read pages. We scan them." / "We don't make optimal choices. We satisfice." (crediting Herbert Simon, 1956) / "We don't figure out how things work. We muddle through." Plus the bridge to defaults: "once they find something that works, they rarely look for a better way." SOLID

NN/g eye-tracking: "Users' eyes are drawn to empty fields" — filled fields become harder to locate. Directly supports your blank-field thesis. SOLID (observational) — https://www.nngroup.com/articles/form-design-placeholders/


§3 — Form abandonment: the numbers

Zuko Analytics — 739 forms, 93,022,997 form views, ~60M starts, ~38M completions, 73.8% mobile. SOLID (vendor telemetry, disclosed sample; notably they publish findings that undercut their own sales pitch).

Metric Figure
Viewers who never touch a single field ~32–35% (60M starts ÷ 93M views = 64.5%)
Starter → completion overall 51.71% (desktop 54.5%, mobile 47.5%)
Registration forms 60.7%
Password field abandonment 10.5% — the highest of any field type (email 6.4%, phone 6.3%)
Zuko's own conclusion "The number of inputs is not the primary factor"

Honest caveat: the ~1-in-3 figure is view-to-start drop-off. It's the best available proxy for "bailed at first sight," but it bundles in irrelevance, price shock, and slow loads. No study isolates "screen of empty fields" as a causal variable. State the funnel number; frame the causation as your interpretation. The script does this.

Baymard Institute — 4,400+ moderated sessions, 335 top-grossing US/EU sites, 200,000+ research hours. Cart abandonment reasons (n≈2,219 US adults, excluding the 42% "just browsing"):

Reason %
Extra costs too high 40%
Delivery too slow 20%
Didn't trust site with card 19%
Site required account creation 18%
Too long / complicated checkout 17%

Version drift trap: Baymard's "too long/complicated" figure has been published at 26% (2019), 22%, and 17% (current). Forced-account-creation is 18% on the current list but 24% in their Jan 2023 article. Always state the year. SOLID — https://baymard.com/lists/cart-abandonment-rate

Other Baymard figures: average checkout is 5.1 steps, 11.3 form fields (2024), down from 12.7 in 2019; most sites need only 8. Up to 19% of checkout abandonment among existing account holders is caused by password-reset issues; 82% of e-commerce sites impose overly complex password requirements. SOLID


§4 — Field count: it is NOT a simple law

This is where most CRO content is wrong, and where you're most likely to get called out. Report it as conditional.

The famous "11 fields → 4 fields = +120%" stat is real but weak: it's Imaginary Landscape, a Chicago web agency, studying its own site, comparing Oct–Nov 2007 to Apr–May 2008. Not an A/B test. Raw counts: 10 conversions vs 26. 95% CI on the lift runs +8% to +341%. CONTESTED — cite only with caveats, or skip.

Counter-evidence that reducing fields can hurt: - Michael Aagaard cut 9 fields → 6 and lost 14%. His diagnosis: "I removed all the fields that people actually want to interact with and only left the crappy ones." Re-test keeping 9 fields but rewriting the labels: +19.2%. CONTESTED (case study) - MarketingExperiments has cases where adding fields lifted conversion +109%. CONTESTED - Zuko's 93M-session dataset: field count "is not the primary factor." SOLID

The best-sourced dose-response is Marketo/MarketingExperiments: 5 fields = 13.4% conversion / $31.24 CPL · 7 fields = 12.0% · 9 fields = 10.0% / $41.90 CPL. CONTESTED (named company, clean monotonic pattern, but no sample size disclosed).

The lesson your script uses: it's not fewer fields, it's fewer decisions. Prefilling a field removes the decision without removing the field. That's why the "after" screenshot keeps four of the five inputs and still converts better — only Occasion is deleted, and it's deleted because it changes nothing, not because it's a field.

Expedia "$12M from one field": BLACKLISTED as a fact. Traceable to Joe Megibow (VP Analytics, Expedia) at a 2010 conference, reported by the now-defunct silicon.com. No published data, no methodology — and CXL retells the same story as "$1 million." A 12x discrepancy between two reputable sources. If you use it, label it a war story.

"$300 Million Button" (Jared Spool, 2009) is better sourced but the company is anonymous. The genuinely useful part: 160,000 password reset requests per day, of which 75% never completed the purchase. CONTESTED


§5 — Defaults

Finding Figure Source Confidence
401k auto-enrollment participation 37% → 86% Madrian & Shea, QJE 2001 SOLID
Stayed on BOTH defaults 61% of auto-enrolled (71% of participants) same SOLID
Would have chosen that combo freely ~1% same SOLID
Effect by demographic +67pp for lowest earners vs +26pp for highest same SOLID
️ The catch Auto-enrolled cohort's average contribution was 2pp LOWER than opt-in cohort's — the 3% default was too low same SOLID
Organ donation, controlled experiment Opt-in 42% · Opt-out 82% · Forced choice 79% (n=161) Johnson & Goldstein, Science 302, 2003 SOLID
Country level Germany ~12% · Austria >99% same SOLID
️ But 5 countries that switched opt-in→opt-out saw no increase in actual donation rates Dallacker et al., Public Health 236, 2024 SOLID
Save More Tomorrow 78% take-up; savings 3.5% → 13.6% over 40 months; 98% still enrolled through 2nd raise Thaler & Benartzi, JPE 2004 SOLID
Google's own internal doc "Of the tiny fraction of end users who try to change the default, many will become frustrated and simply leave the default as originally set" US v. Google, 2024 SOLID

Per-country organ donation decimals (Denmark 4.25%, UK 17.17%, etc.) circulate widely but are chart readings I could not verify against the primary PDF. Stick to Germany ~12% / Austria >99%. SHAKY

Why defaults work — users read them as advice. Three independent sources say the same thing: - Spool's Word users: "Microsoft must know what they are doing." - Madrian & Shea: "many employees taking the default as investment advice." - Nielsen: users assume systems present optimal defaults.

This inferred endorsement is exactly what makes exploiting a default a betrayal rather than a nudge — and it's the strongest ethical argument available to you, sourced three ways.

The honest counterweight (worth having in your pocket): when Mozilla switched Firefox's default from Google to Yahoo in 2014, Google's share fell from 80–90% to 60–70%. Most users stayed with the new default; a large minority actively switched back. Defaults are powerful, not deterministic.

"Smart defaults" provenance: Jenifer Tidwell, "Good Defaults" pattern (1999) → Jakob Nielsen, "The Power of Defaults" (NN/g, 2005) → Luke Wroblewski, Web Form Design, Rosenfeld Media, 2008 — where the exact phrase enters the vocabulary. Current NN/g guidance (Dec 2024): "Assume that users will not change defaults and plan for this in your designs."

The ethical line: Harry Brignull coined "dark pattern" (2010) as "a manipulative or deceptive trick in software that gets users to complete an action that they would not otherwise have done." Nielsen's criterion for a good default is "the most common value" — a user-benefit test that is the exact inverse. And the law has already drawn the line: CJEU, Planet49, Case C-673/17 (1 Oct 2019) — a pre-ticked consent checkbox is not valid consent under GDPR. SOLID


§6 — Recommendation surfaces

Claim Figure Confidence
Netflix: recommendations vs search ~80% / ~20% of hours streamed SOLID — Gomez-Uribe & Hunt, ACM TMIS 2015
Netflix homepage share "2 of every 3 hours streamed on Netflix are discovered" there SOLID (same)
Netflix business value "save us more than $1B per year"; "reduced churn by several percentage points" SOLID (same)
Spotify "One third of all new artist discoveries happen on personalized, algorithmic playlists" SOLID — 2022 Investor Day
️ Spotify share of total listening from algorithmic surfaces Not published. Don't let this get conflated with the discovery figure.
YouTube ">70%" of watch time from recommendations — Neal Mohan, CPO, at CES 2018 CONTESTED — conference remark, never published, 8 years old. Attribute by name and year.
Amazon "35% of purchases" Uncited McKinsey 2013 line BLACKLISTED

§7 — Authentication

Claim Figure Confidence
Social login share of actual logins 14% of logins; 25% of MAU used it at least once; 62% of logins still username/password SOLID-ish — Auth0/Okta 2022, "thousands of websites and apps"
Google's share of social logins >73% (75% on Auth0 specifically). Apple #2. same
️ Stated preference vs revealed behavior Gigya's survey: 88% of consumers "have used social login." Auth0's server logs: 14% of logins. The best contrast in this section. Use it.
Independent research — cuts against the vendor narrative Social-login registrants "exhibit shorter tenures and lower levels of activity"; "the marginal privacy cost associated with social login is larger than its marginal convenience benefit" SOLID — Newberry, Kim & Wagman, Economics Letters, 2025
Lab research 15% of participants refused Facebook Connect on privacy grounds even in a controlled setting SOLID — Egelman, CHI 2013
Passkeys Sign-in conversion 92% vs 54%; registration 63% vs ~25% CONTESTED — Dashlane/Google, 7 months, but on the most passkey-friendly audience imaginable. Don't generalize.

There is no credible independent measurement of signup-conversion lift from adding social login. Every lift number in circulation is vendor-published or unsourced. The defensible case rests on the password-abandonment data (§3), not on a lift stat.

Apple App Store Guideline 4.8 — note it is no longer "you must offer Sign in with Apple." Since 2022 it's technology-neutral: if you use any third-party/social login for the primary account, you must also offer an equivalent service that (a) limits collection to name + email, (b) lets users keep their email private, and (c) doesn't collect in-app interactions for advertising without consent. If you ship only email/password, 4.8 doesn't apply at all. SOLID — https://developer.apple.com/app-store/review/guidelines/

Apple private relay operational trap (worth a future video): relay addresses are forwarding addresses — mail must come from a domain registered and SPF-verified with Apple or it bounces silently. And Apple is consolidating privaterelay.appleid.com and icloud.com onto a new private.icloud.com domain, no launch date announced as of June 2026. Unprepared services will break password resets and verification codes.


§8 — Assortment reduction: the real evidence

** Boatwright & Nunes (2001), Journal of Marketing 65(3): 50–63. An online grocer cut SKUs dramatically in 94% of categories. Sales rose an average of 11% across 42 categories. Sales rose in >2/3 of categories; 75% of households increased overall spend. ️ Their own caveat: "category sales depend on the total number of SKUs offered" — there's a floor. SOLID

** Borle, Boatwright, Kadane, Nunes & Shmueli (2005), Marketing Science 24(4): 616–622. The same authors, the same retailer, measured at the store level instead of the category level — and the sign flipped. "The reduction in assortment reduces overall store sales." Purchase frequency fell; expected time between deliveries rose ~25%**. Purchase amounts fell ~4.8%.

This is the most important nuance in the entire brief. Cutting made every category perform better and made the store perform worse. App translation: removing things can improve every screen you measure while quietly reducing how often people open the app at all. Per-session engagement is the category metric; retention and session frequency are the store metric. SOLID — it's Appendix A2 in the script.

** Broniarczyk, Hoyer & McAlister (1998), JMR 35(2): 166–176. Consumers judge variety by heuristic cues, not by counting. Three conditions: 1. Category space must stay constant — if the physical footprint shrinks, people notice 2. Cut only low-preference items — reductions up to 54% often go unnoticed 3. ️ The exception: "if someone's favorite is deleted, even if it is generally a low-preferred item, then that consumer will see a diminished assortment"

This maps to app UI more rigorously than anything else in the brief. You can remove half your low-usage features without users perceiving less capability — provided the visual footprint doesn't shrink and you don't delete the one obscure thing a given user depends on. "Low usage overall" is the wrong deletion criterion. "Low usage AND not anyone's favorite" is the right one. A feature used by 2% of users constantly ≠ a feature used by 30% of users once.

This is the direct justification for "View All Times" in the after-screenshot: the popular options are surfaced, but the footprint is preserved and nobody's favorite is hidden.

Walmart Project Impact (2008–2011) — the cautionary tale:

Fact Figure Confidence
SKUs cut ~15% of items CONTESTED (analyst estimate)
Consecutive quarters of US comp declines 9 SOLID (Reuters + Walmart FY12 Q3 filing)
Items restored 8,500 (~11% more products per store), tagged "It's Back" SOLID
Category bleed Salty snacks and granola bars declined at Walmart while growing at competitors SOLID — devastating detail
️ "$1.85 billion lost" Traces to marketing blogs citing each other. No filing, no analyst note. BLACKLISTED

Bill Simon (Walmart US CEO): "Our customers can't buy it if we don't sell it. And if we don't sell it they will go somewhere else to buy it." SOLID

Why it failed — it broke Broniarczyk's rules directly: they cut brands customers were loyal to, not redundancy, and they shrank the space (Project Impact removed "Action Alley," the main promotional corridor).

Contemporary counterweight: Dollar General cut 1,500+ SKUs and the CEO called it "the cornerstone of our stabilization of retail" — in-stocks up ~250bps, inventory down 5.7%. SOLID (earnings call, March 2026).


§9 — Progressive disclosure & multi-step forms: weaker than you'd think

"Multi-step forms have 86% higher conversion rates" — BLACKLISTED. Attributed to "a HubSpot report" across dozens of blogs. The report doesn't exist. The trail dead-ends in mutual citation between form vendors.

The most useful quote in this area — GOV.UK's "One Thing Per Page" is the most-cited pattern in the space. Its own author, Tim Paul of GDS, commented on his own post: "I wish we had some easy-to-share quant data on this as well, but I'm not aware of any." SOLID and an excellent honesty beat if you ever cover this.

Nielsen's progressive disclosure article (NN/g, 2006) cites no studies, no sample sizes, no measured deltas. Its strongest evidentiary sentence is "Research says that these are groundless worries" — with no citation. The principle is sound; the numbers do not exist.

Foot-in-the-door is real — Freedman & Fraser (1966), JPSP 4(2): 195–202. 52.8% vs 22.2% control (p < .02); Experiment II: 76.0% vs 16.7%. SOLID — but Burger (1999) shows it's driven by six competing processes and can backfire. Small-n 1960s door-to-door experiments generalize to signup flows by analogy only.

️ The Zeigarnik effect is dead — BLACKLISTED it as a justification for progress bars. Ghibellini & Meier (2025), Humanities and Social Sciences Communications (Nature Portfolio) meta-analysis: d_z = 0.15; weighted interrupted-to-completed recall ratio = 0.99 (i.e. no advantage); interrupted tasks are 49.16% of recalled tasks — indistinguishable from chance. Authors: "The current findings do not support a memory advantage for interrupted tasks."

What DOES replicate is the Ovsiankina effect: interrupted tasks are resumed 66.79% of the time. So progress indicators may work — just not for the reason every CRO blog gives. This is a strong, checkable credibility moment (script Appendix A4). SOLID


§10 — Placeholder text as labels

NN/g, "Placeholders in Form Fields Are Harmful" (2014, updated 2018). Your "before" screenshot uses placeholder-as-label on every field. Seven documented failure modes:

  1. Disappearing text strains short-term memory
  2. Users can't verify entries before submitting
  3. Error correction requires deleting your answer to re-read the question
  4. Keyboard/tab navigation disrupted
  5. "Users' eyes are drawn to empty fields" — filled fields become harder to find
  6. Placeholders get mistaken for pre-filled values, so users skip the field
  7. Some implementations don't auto-clear

Accessibility: grey placeholder text typically fails contrast requirements; screen readers handle it inconsistently. WCAG: implicates 1.3.1 (Info and Relationships), 3.3.2 (Labels or Instructions), 1.4.3 (Contrast Minimum).

Recommendation: persistent labels above the field. SOLID (observational — frame as "NN/g's usability research," not "a study found") — https://www.nngroup.com/articles/form-design-placeholders/

Quantitative backing for form guidelines generally: Seckler, Heinz, Bargas-Avila, Opwis & Tuch (2014), CHI 2014, pp. 1275–1284. N = 65, controlled eye-tracking. Guideline-compliant forms produced faster completion, fewer submission attempts, fewer eye movements, higher satisfaction. SOLID — peer-reviewed, disclosed n. ️ NN/g's widely-quoted "78% vs 42% one-try error-free submissions" is attributed to this paper but I could not confirm which study it comes from. Say "Nielsen Norman Group reports" rather than "a study found."


§11 — Retention benchmarks

Claim Figure Confidence
"Average app loses 77% of DAU within 3 days" Day 1: 29.17% · Day 3: 23.42% · Day 7: 17.28% · Day 30: 9.55%. Top 10 apps: Day 1 74.67%, Day 30 59.80% CONTESTED — Andrew Chen / Quettra, 125M phones, 2015 data. If you say 77%, say the year. It's 11 years old and Quettra no longer publishes.
B2B tech 3-month retention Median 2.5%, top performers 15.6% CONTESTED — Amplitude, 2,600+ companies, Sept 2023–Sept 2024
Activation ↔ retention link 69% of products with strong early activation were also strong 3-month retention performers CONTESTED — same. This is the figure that supports the onboarding thesis.
Amplitude null finding No relationship between top-quartile user acquisition and top-quartile retention CONTESTED — same

There is no credible primary source for "X% of users abandon during onboarding." Every version I traced ends at an SEO aggregator. BLACKLISTED all of them.


PART 3 — THE BLACKLIST

Never say these on camera:

Claim Why
"30 jams vs 6 jams" It was 24. "30" is a different experiment.
"Reducing to 6 jams increased sales 30%" It's a ~10x purchase-rate ratio, not a 30% lift.
"70–90% never change defaults" No study produces this range.
"Netflix users spend more time browsing than watching" No source exists.
"75% of Netflix viewing from recommendations" (McKinsey) Uncited 2013 sentence. Use Netflix's own 80%.
"35% of Amazon purchases from recommendations" Same uncited sentence.
"P&G cut Head & Shoulders 26→15, sales up 10%" No primary source anywhere.
"Walmart lost $1.85 billion on Project Impact" Marketing blogs only. Use the 9 quarters / 8,500 items.
"Multi-step forms convert 86% better (HubSpot)" The HubSpot report doesn't exist.
"The Zeigarnik effect means progress bars work" 2025 meta-analysis found the effect essentially null.
"88% of consumers use social login" Stated preference. Server logs say 14%.
"Expedia made $12M from one field" (as fact) No published data; also retold as $1M.
"81% of people abandon online forms" n=502, 56% aged 55+, self-reported.
"X% abandon during onboarding" Every version traces to SEO aggregators.
"7±2 means max 7 menu items" Miller called it a coincidence; NN/g calls the UI usage a misunderstanding.
"Better checkout UX = 35% more conversions" (Baymard) That's Baymard's modeled upside, not a measured result.

PART 4 — PRIMARY SOURCES

Choice & cognition Iyengar & Lepper (2000) — https://faculty.washington.edu/jdb/345/345%20Articles/Iyengar%20&%20Lepper%20(2000).pdf Scheibehenne, Greifeneder & Todd (2010) — https://scheibehenne.com/ScheibehenneGreifenederTodd2010.pdf Chernev, Böckenholt & Goodman (2015) — https://chernev.com/wp-content/uploads/2017/02/ChoiceOverload_JCP_2015.pdf Schwartz, PBS NewsHour (2014) — https://www.pbs.org/newshour/economy/is-the-famous-paradox-of-choic Schwartz et al. (2002), maximizer/satisficer — https://works.swarthmore.edu/fac-psychology/101/ Sweller, van Merriënboer & Paas (2019) — https://leadinglearner.me/wp-content/uploads/2019/02/sweller2019_article_cognitivearchitectureandinstru.pdf Cowan (2001) — https://www.cambridge.org/core/services/aop-cambridge-core/content/view/44023F1147D4A1D44BDC0AD226838496/S0140525X01003922a.pdf Miller (1956) — https://labs.la.utexas.edu/gilden/files/2016/04/MagicNumberSeven-Miller1956.pdf NN/g on chunking / 7±2 — https://www.nngroup.com/articles/chunking/ Proctor & Schneider (2018), Hick's law review — https://web.ics.purdue.edu/~dws/pubs/ProctorSchneider_2018_QJEP.pdf

Decision fatigue / ego depletion Danziger et al. (2011) — https://www.pnas.org/doi/10.1073/pnas.1018033108 Weinshall-Margel & Shapard critique — https://pmc.ncbi.nlm.nih.gov/articles/PMC3198355 Hagger et al. (2016), 23-lab replication — https://journals.sagepub.com/doi/10.1177/1745691616652873 Vohs et al. (2021), 36-lab replication — https://www.asc.upenn.edu/sites/default/files/2022-06/Vohs%20et%20al%202021.pdf

Defaults Madrian & Shea (2001) — https://www.nber.org/papers/w7682 Johnson & Goldstein (2003) — https://www.science.org/doi/10.1126/science.1091721 Thaler & Benartzi (2004) — https://www.journals.uchicago.edu/doi/10.1086/380085 Gross & Acquisti (2005) — https://www.heinz.cmu.edu/~acquisti/papers/privacy-facebook-gross-acquisti.pdf Spool, "Do Users Change Their Settings?" — https://archive.uie.com/brainsparks/2011/09/14/do-users-change-their-settings/ Nielsen, "The Power of Defaults" — https://www.nngroup.com/articles/the-power-of-defaults/ NN/g, "The Danger of Defaults" (2024) — https://www.nngroup.com/videos/the-danger-of-defaults/ Tidwell, "Good Defaults" (1999) — https://www.mit.edu/~jtidwell/language/good_defaults.html Dallacker et al. (2024), organ donation — https://www.mpib-berlin.mpg.de/press-releases/organ-donation

Forms & conversion Baymard methodology — https://baymard.com/research/methodology Baymard cart abandonment — https://baymard.com/lists/cart-abandonment-rate Baymard password requirements — https://baymard.com/blog/password-requirements-and-password-reset Zuko benchmarking data — https://www.zuko.io/benchmarking/about-the-data NN/g placeholders — https://www.nngroup.com/articles/form-design-placeholders/ Seckler et al., CHI 2014 — https://dl.acm.org/doi/10.1145/2556288.2557265 Ghibellini & Meier (2025), Zeigarnik meta-analysis — https://www.nature.com/articles/s41599-025-05000-w Freedman & Fraser (1966) — https://www.bulidomics.com/w/images/6/6c/Freedman_fraser_footinthedoor_jpsp1966.pdf GDS, "One Thing Per Page" — https://designnotes.blog.gov.uk/2015/07/03/one-thing-per-page/

Auth Gomez-Uribe & Hunt (2015), Netflix — https://dl.acm.org/doi/10.1145/2843948 Auth0/Okta social login report — https://assets.ctfassets.net/2ntc334xpx65/77U9sLFO7rD7t9zdI6Q1SV/a8e2054b5affc0280769516eee70b0ea/Social-Login-Report.pdf Newberry, Kim & Wagman (2025) — https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5100863 Egelman, CHI 2013 — https://blues.cs.berkeley.edu/?p=93 Apple App Store Guidelines — https://developer.apple.com/app-store/review/guidelines/

Retail assortment Boatwright & Nunes (2001) — https://journals.sagepub.com/doi/10.1509/jmkg.65.3.50.18330 Borle et al. (2005) — https://ideas.repec.org/a/inm/ormksc/v24y2005i4p616-622.html Broniarczyk, Hoyer & McAlister (1998) — https://www.semanticscholar.org/paper/472e568670075463ba10c92bcb29818ea9dd86d5 FMI Food Industry Facts — https://www.fmi.org/our-research/food-industry-facts Costco FY2025 10-K — https://www.sec.gov/Archives/edgar/data/909832/000090983225000101/cost-20250831.htm Fortune, "Inside the secret world of Trader Joe's" — https://fortune.com/2010/08/23/inside-the-secret-world-of-trader-joes/ RetailWire, Walmart reverses SKU rationalization — https://retailwire.com/discussion/walmart-reverses-course-on-sku-rationalization/ Nielsen State of Play 2023 — https://www.nielsen.com/news-center/2023/nielsens-state-of-play-report-delivers-new-insights-as-streamings-next-evolution-brings-content-discovery-challenges-for-viewers/