# Research Brief — Onboarding & Personalization: Why the Best Onboardings Ask You Questions
### Fact-check reference for the YouTube script

**Confidence key**
`SOLID` — primary source located, figure verified
`CONTESTED` — real source exists, but methodology or interpretation is disputed
`SHAKY` — widely repeated, source is weak, absent, or circular
`BLACKLISTED` — never repeat this claim

**Scope guard — what the prior two videos already own.** The forms/decision-fatigue video covered choice overload, smart defaults, social login, field counts, and cognitive load. The progress-bars video covered endowed progress, goal gradient, Zeigarnik/Ovsiankina, labor illusion (Buell & Norton), and Duolingo's streak tests. This brief is the *question-asking, commitment, effort, and personalization* layer. Where a beat touches prior ground, it's cross-referenced, not re-derived.

---

# ⚠️ PART 1 — CORRECTIONS TO POPULAR CLAIMS

### 1. Duolingo's "delayed signup +20% DAU" — real, but almost everyone misquotes it

The primary source is Gina Gotthilf (then VP of Growth, Duolingo), interviewed by First Round Review. Exact words: *"Simply moving the sign-up screen back a few steps led to about a 20% increase in DAUs."* `SOLID` (company-published, single named source, no methodology or CI disclosed).

Three corrections to the folklore version:
- It's a **DAU** increase, not "+20% signups" or "+20% conversion." Secondary blogs routinely swap the metric.
- The +20% was the **initial move**. The subsequent soft-wall/hard-wall optimization (replacing a red "Discard my progress" button with a subtle "Later") added a further **8.2% DAU** — a separate number that gets blended into the first.
- It is one company's A/B result on one product circa 2016–17. Duolingo's own launch threshold was a 1% lift; Gotthilf explicitly warns that new apps should be hunting 20–30% effects because small ones are noise at low traffic.

Same interview, same status (company-published): winning notification copy alone **+5% DAU**; streak "weekend amulet" **+2.1% D7 / +4% D14**; badges **+2.4% DAU** and **+116% friends added**; in-app coach (Duo mascot) **+7.2% D14**; notifications sent at **23.5 hours** after last use.

### 2. "Asking questions in onboarding works because personalization" — the causal story is mostly untested

The honest hierarchy of evidence:
- That *asking a question changes the asker's subsequent behavior* is peer-reviewed and meta-analyzed (Part 2, §A). Small effects, real.
- That *personalized content outperforms generic content* is peer-reviewed (Tam & Ho; Matz et al., with fights — §C).
- That *a personalization quiz at signup causally improves activation/retention* has **no public controlled study**. Duolingo's company-published A/B tests are the closest thing that exists. Every "quizzes lift retention X%" number in circulation is vendor marketing.

This mirrors the confidence-gap finding from the progress-bars brief: industry certainty exceeds the literature.

### 3. Quiz-vendor conversion stats are selection-biased marketing — don't launder them as research

Interact's 2026 report: **40.1%** of quiz starters become leads; **65%** quiz completion. Unbounce-derived comparison: quiz pages 30–40% vs standard forms 6.6%. `SHAKY` as generalizable claims — quiz platforms measuring quizzes built by customers who chose quizzes, versus all forms everywhere. Different products, different audiences, different intent. Cite only as "quiz-platform vendors report," never "studies show."

### 4. "Users judge your app in 50 milliseconds" — real study, wrong conclusion

Lindgaard et al. (2006), *Behaviour & Information Technology* 25(2): visual-appeal ratings after **50ms** exposure correlate highly with ratings after 500ms. `SOLID` for what it says: first aesthetic impressions form fast and are consistent. It says **nothing** about abandonment, decisions, trust, or onboarding. The "you have 50ms before users leave" version is an extrapolation the paper never makes. Related: "users decide in the first 3/7/10 seconds" has no traceable source at all — `BLACKLISTED`.

### 5. Facebook's "7 friends in 10 days" — a heuristic from a talk, not an experiment

Chamath Palihapitiya, Growth Hackers Conference 2012: *"The single biggest thing we realized was to get any individual to 7 friends in 10 days. That was it."* `SOLID` as a quote; `CONTESTED` as evidence. It was a **correlational threshold** read off retention curves — engaged users happened to have ≥7 friends by day 10 — not a randomized finding, and Facebook never published the analysis. Mode Analytics and others have shown similar-looking "magic number" thresholds appear in almost any dataset if you go looking. Use it as the origin story of activation metrics, framed as pattern-mining, not proof.

### 6. "Apps lose 77% of users in 3 days" — already flagged; stays flagged

Quettra/Andrew Chen, 2015, 125M Android phones. `CONTESTED` and 11 years old — covered in the forms video brief (§11). If used at all, say the year. The adjacent stat — **~1 in 4 users abandon an app after a single use** (Localytics: 25% in 2015, 23% in 2016, 21% in 2017–18) — is vendor telemetry from 37,000 apps, directionally credible, also dated. `CONTESTED`; say "Localytics measured," with the year.

### 7. The big personalization percentages are stated-preference surveys, not behavior

- Epsilon (2018): **80%** more likely to buy from brands offering personalized experiences — vendor survey of stated attitudes.
- Accenture (2018): **91%** more likely to shop with brands that recognize/remember them — same genre.
- McKinsey "Next in Personalization" (2021): **71%** expect personalization, **76%** frustrated without it; "personalization leaders drive 40% more revenue" is a McKinsey **estimate**, not a measured experiment.

All `SHAKY` as causal claims. Server logs vs. surveys was the sharpest contrast in the forms video (88% "use social login" vs 14% of actual logins) — same trap here. People *say* they want personalization; the measured evidence for when it actually pays is Part 2 §C, and it's far more conditional.

### 8. Duolingo "Birdbrain AI" retention numbers circulating on SEO blogs are fabricated

"17% higher success rates," "35% better vocabulary retention," "67% MAU growth from AI personalization" — none of this appears in anything Duolingo published. Duolingo's engineering blog confirms Birdbrain exists (difficulty prediction model) but publishes no retention percentages for it. `BLACKLISTED` all Birdbrain numbers.

---

# PART 2 — THE VERIFIED SPINE

## §A — Question-asking effects: asking is an intervention

**⭐ Morwitz, Johnson & Schmittlein (1993)**, *JCR* 20(1): 46–61. Large consumer panels (~100,000 households; 40,000+ analyzed). Households asked a single purchase-intent question about automobiles: **3.3%** bought within 6 months vs **2.4%** of those not asked — roughly a **35–40% relative lift from one question**. Similar effect for PCs. Repeated asking polarized: low-intent respondents asked repeatedly became *less* likely to buy. `SOLID` — this is the cleanest big-n foundation for the whole video. One question, no persuasion, measurable behavior change.

**⭐ Wood, Conner, Miles, Sandberg, Taylor, Godin & Sheeran (2016)**, *Personality and Social Psychology Review* 20(3): 245–268. Meta-analysis, **116 published tests**: overall **d+ = 0.24** — small, positive, real. By question type: self-prediction questions strongest (**d+ = 0.29**), intentions-only weakest (**d+ = 0.12**). Effects larger for **socially desirable, easy** behaviors and in **student samples**; little support for any single mechanism (attitude accessibility, dissonance, fluency all under-supported). `SOLID` — and state the size honestly: this is a nudge, not a lever.
⚠️ Caveats to carry: the literature skews to health/civic behaviors; effects shrink for difficult behaviors — and "keep using this app for months" is a difficult behavior. Durability exists in spots (Godin et al. found blood-donation effects at 6 and 12 months post-questionnaire) but decay is expected as accessibility fades.

**Sherman (1980)**, *JPSP* 39(2): 211–221, "self-erasing errors of prediction." Residents asked to *predict* whether they'd volunteer 3 hours for the American Cancer Society; called back 3 days later: **31% of predictors agreed vs 4% of directly-asked controls**. People over-predict their virtue, then behave to make the prediction true. `SOLID` (small-n era, but the canonical self-prophecy demo). This is the direct ancestor of "What's your goal?" screens: the user who taps "Lose 10 lbs" or "15 min/day" has just made a self-prediction.

**Commitment & consistency — the strongest primaries behind Cialdini's chapter:**
- **Moriarty (1975)**, *JPSP* 31(2). Staged beach thefts: bystanders who'd been asked "watch my things?" intervened **95% (19/20)** vs **20% (4/20)** of controls. `SOLID` result / tiny cells — say the raw numbers, not just the percentages. ⚠️ The French "replications" of this paradigm are by Nicolas Guéguen, whose corpus is under serious integrity investigation — do not cite Guéguen for anything.
- **Cialdini, Cacioppo, Bassett & Miller (1978)** — low-ball: agreeing first, then learning the 7am cost → **~56% vs ~31%** compliance. A 2016 meta-analysis (23 studies, n=4,733) finds the low-ball holds: **r = .16, OR = 2.47**. `SOLID` — one of the better-surviving classic compliance effects.
- **Cioffi & Garner (1996)**, *PSPB* 22(2): 133–147. **Active choice** (ticking a box to volunteer) produced stronger commitment and follow-through than **passive choice** (not opting out), persisting 6 weeks. `SOLID` qualitatively. ⚠️ The "49% vs 17% showed up" figures often attached to it — I could not verify against the paper; don't state them. Direct product translation: making the user *tap* their goal beats preselecting it for them — a genuine tension with the defaults gospel from the forms video, and worth playing as such.
- Foot-in-the-door (Freedman & Fraser 1966) — covered in the forms video brief (§9); reference, don't rebuild.

**The reflexive kicker (novel, rigorous):** mere-measurement means your onboarding survey **contaminates your own analytics**. "Users who set a goal retain 30% better" confounds (a) self-selection, (b) the question-behavior effect itself. The only clean test is randomizing whether the question is asked — which almost nobody does publicly. Duolingo's A/B culture is the exception, which is exactly why their numbers are the only ones worth quoting.

## §B — Effort effects: the IKEA effect and its dangerous boundary condition

**⭐ Norton, Mochon & Ariely (2012)**, *Journal of Consumer Psychology* 22(3): 453–460 (HBS WP 11-091). Four studies — IKEA boxes, origami, Lego:
- Builders bid **63% more** for boxes they assembled themselves than non-builders bid for identical pre-assembled ones.
- Origami folders valued their (amateur) creations at ~**5×** non-builders' valuations — and near what non-builders would pay for *expert*-made origami. Builders also wrongly expected others to share their inflated valuation.
- **The boundary condition that is the whole point for onboarding:** *"labor leads to love only when labor results in successful completion."* Participants who built then **destroyed** their creations, or who **failed to complete** them, showed **no IKEA effect**. Partial assembly < complete assembly.
`SOLID` — and conceptually replicated by Sarstedt, Neubert & Barth (2017) with loom bands, including the destruction condition.

**App translation (the sharpest transfer in this brief):** effort invested in a personalization quiz only converts to attachment if the user *finishes and sees the product of their labor*. An abandoned 20-question quiz is pure cost — the effect literally requires completion. This is the effort-side twin of the endowed-progress logic from the progress-bars video, and it argues for (a) short quizzes, (b) an explicit "here's your plan, built from your answers" payoff screen as the completion artifact.

**Effort justification — the classic is shakier than the IKEA effect.** Aronson & Mills (1959): severe initiation (reading embarrassing material aloud) → higher liking for a dull group. Gerard & Mathewson (1966) replicated with electric shocks, supporting it. But: n=63 female undergrads in the original, demand-characteristics critiques, and a modern replication attempt (~70 participants) found **no** severity effect on group identity — and mild initiation *plus reward* beat severe initiation. `CONTESTED` — say "the classic dissonance studies," attribute, don't state as settled. Lean on IKEA (better-replicated, quantified) and treat Aronson & Mills as historical color.

**Labor illusion (Buell & Norton 2011)** — "we're building your plan…" fake-work screens after quizzes. Covered in the progress-bars brief; reference only. Note the connection: quiz answers give the labor illusion its *script* ("analyzing your answers…"), which is part of why quizzes and fake-processing screens travel together.

## §C — Personalization payoffs: what's actually measured

**Tam & Ho (2005)**, *Information Systems Research* 16(3): 271–291, and **Tam & Ho (2006)**, *MIS Quarterly* 30(4): 865–890. The canonical peer-reviewed web-personalization experiments (lab + field). Mechanism findings: personalization works through **content relevance** and **self-reference** (content framed around *you*), which increase attention, elaboration, and acceptance of recommended options — and **goal specificity** moderates: personalization persuades most when the user's goal is vague. `SOLID` for direction and mechanism; no single dramatic percentage to quote, and that's fine — this is the "why," not the hook.

**Matz, Kosinski, Nave & Stillwell (2017)**, *PNAS* 114(48): 12714–12719. Three Facebook field experiments, **3.5M+ people reached**: ads matched to users' psychological traits (extraversion, openness — inferred from likes) produced **up to 40% more clicks and up to 50% more purchases** than mismatched/generic ads. `CONTESTED` — and you must carry the fight if you use it:
- **Eckles, Gordon & Johnson (PNAS letter)**: Facebook doesn't randomly assign users to campaigns; the delivery algorithm optimizes each campaign differently → internal validity threat.
- **Sharp, Danenberg & Bellman (PNAS letter)**: creative-quality differences, not targeting, could explain results.
- **Matz et al. replies**: algorithmic confounds are unlikely to produce this pattern; effects are robust.
The defensible on-camera framing: "the biggest field test of psychological targeting found up-to-40%-higher click rates — with a real methodological fight about how much was targeting versus Facebook's own delivery machinery."

**Aguirre, Mahr, Grewal, de Ruyter & Wetzels (2015)**, *Journal of Retailing* 91(1): 34–49 — the **personalization paradox**. Personalization built on **covertly** collected data backfires: field data show *"sharp drops in click-through rates when customers realize their personal information has been collected without their consent."* **Overt** collection (asking!) plus trust cues flips the effect positive. `SOLID` — and it's the single best scientific justification for the quiz pattern: **asking is the disclosure mechanism that makes personalization feel legitimate instead of creepy.** The quiz isn't just data collection; it's consent theater that actually matters.

**Locke & Latham (2002)**, *American Psychologist* 57(9) — 35-year synthesis, ~400 studies: **specific, difficult goals beat "do your best" with d ≈ 0.42–0.80**, one of the most robust effects in organizational psychology (goal-commitment and feedback required as moderators). `SOLID` for the science; **analogy only** for goal-selection screens — the studies are work/lab tasks, not app onboarding. Duolingo's daily-goal picker is a plausible application (specific goal + public commitment + streak feedback loops all three moderators), but say "applies by analogy."

**Superhuman's concierge onboarding — well documented, company-published** (First Round Review, written by growth lead Gaurav Vohra; plus Rahul Vohra on Lenny's/20VC):
- Every user got a mandatory 1:1 video call — 90 minutes in the early days, later compressed to **30 minutes**; Rahul Vohra personally onboarded the first ~100 users.
- **65%+** of human-onboarded customers fully switched their email vs **~30%** self-serve.
- Each ramped specialist ≈ **$650K ARR/year** (~40 calls/week); manually onboarded users showed **~2× referral rates**; one specialist logged **3,900+** onboardings.
- Later productization moved self-serve activation **40% → 50%**, and the company wound down universal concierge onboarding when costs beat benefits.
`SOLID` as "Superhuman reports" — self-reported, no control group, survivor-told. The onboarding call is question-asking at maximum fidelity: a human interviews you, then configures the product to the answers. Related: Rahul Vohra's PMF survey ("how disappointed if you could no longer use…", **40% 'very disappointed'** threshold, adapted from Sean Ellis) — itself a question-based instrument, company-published methodology.

**Activation benchmarks — vendor telemetry, label as such:**
- OpenView/Pendo 2023 Product Benchmarks (~1,000 companies): **median activation ~30%** of signups reach first value. `CONTESTED` (self-reported definitions vary by product).
- Amplitude's **69%** activation→retention association — **already used in the prior video**; cross-reference, don't respend it.
- "Time to value" has no standardized definition across these reports; never quote a TTV benchmark as if comparable across products.

## §D — When asking backfires

**Length kills honestly and predictably.** SurveyMonkey's own completion telemetry: abandonment climbs once surveys pass **7–8 minutes** (completion drops 5–20%); per-question drop-off is steepest **up to ~15 questions**, past which fewer than half of starters finish. `CONTESTED` (vendor data, surveys ≠ onboarding quizzes, but it's the best dose-response available). The first-question cost finding (Liu & Wronski 2018 — first-question wording is ~3.2× costlier than later text) is in the progress-bars brief; cross-reference for the "your first question is the expensive one" beat.

**Noom — the documented cautionary tale, two layers:**
1. *Legal:* Noom paid **$56M (+$6M in credits)** to settle a ~2M-member class action (S.D.N.Y., 2022) over its trial-to-subscription flow: quiz-driven trial funnel into auto-renewed lump-sum charges with cancellation friction (cancel-via-your-coach). FTC's Oct 2021 enforcement policy statement on dark patterns is the regulatory backdrop. `SOLID` — the settlement is about the *subscription mechanics*, not the quiz itself; don't overclaim "the quiz was ruled a dark pattern."
2. *Design critique (attributed, not research):* Jason Hreha (ex-Stanford Persuasive Tech Lab, built Walmart's behavioral science team) documented his Noom onboarding at **~45 minutes** of sequential questionnaires before meaningful value, and argues the length functions as a **motivation filter**: only the highly committed survive it, which then inflates Noom's published outcome stats (e.g., "78% lost weight") via survivorship. `CONTESTED` (one expert's teardown) — but the screening logic is sound and it's the best articulation of "the quiz as a filter, not a service."
⚠️ Do not state a specific Noom question count ("59 questions" etc.) — teardowns disagree and the flow changes constantly. Say "tens of minutes of questions, by multiple documented teardowns."

**Covert personalization backfires** — Aguirre et al. (§C): same personalization, resented when the data collection was invisible. The quiz pattern's dark mirror: personalizing from *behavioral surveillance* without asking produces the creepiness penalty that asking avoids.

**Mere-measurement cuts both ways** — Morwitz et al. found repeated intent questions **decreased** purchase among low-intent respondents. Asking an uncommitted user to declare a goal can crystallize "actually, no." Polarization, not uniform lift.

**Question order and trust** — Hreha's core critique of Noom generalizes: intensely personal questions (weight, mental health, medical history) *before* any value or trust is established invert the self-disclosure norm. No controlled onboarding study exists on this; frame via Aguirre's vulnerability mechanism.

---

# PART 3 — HOOK CANDIDATES (verified, ranked)

1. **"One survey question made people 35% more likely to buy a car."** 3.3% vs 2.4% across 40,000+ households, and nobody sold them anything — they were just *asked*. (Morwitz 1993, `SOLID`) — cleanest cold open for the whole thesis.
2. **"95% vs 20%."** Ask a stranger to watch your stuff and they'll chase down a thief; don't ask and they'wll watch him walk away. 19/20 vs 4/20, staged thefts, Jones Beach 1972. (Moriarty 1975, `SOLID`, tiny n — show the raw counts.)
3. **"Duolingo deleted the signup screen and DAUs rose 20%."** The best-documented onboarding A/B result in existence — and it's about *removing* a question, which sets up the real thesis: ask the right questions, not the most questions. (Gotthilf/First Round, company-published.)
4. **"People paid 63% more for a box because they screwed it together themselves"** — and the effect vanishes if they don't finish. (Norton, Mochon & Ariely 2012, `SOLID`.)
5. **"31% vs 4%."** Predict-your-own-behavior question vs direct ask, Cancer Society volunteering. (Sherman 1980, `SOLID`.)
6. **"A $56 million quiz."** Noom's quiz-driven funnel ended in the largest dark-patterns class settlement of its era. (`SOLID` on the settlement; careful framing per §D.)

# PART 4 — NOVEL-ANGLE CANDIDATES

1. **The quiz is the commitment, not the data.** Strongest synthesis available: mere-measurement (§A) + active choice (Cioffi & Garner) + IKEA effect (§B) all predict onboarding questions work *even if the answers are never used*. The personalization payoff (§C) is real but conditional; the commitment payoff fires regardless. No one else covering onboarding makes this separation.
2. **The completion cliff.** IKEA effect requires *finished* labor — so a quiz abandoned at question 14 of 20 is strictly worse than no quiz. Marries §B's boundary condition to §D's length data into one actionable rule: every question you add raises both drop-off *and* the stakes of drop-off.
3. **Your onboarding survey is contaminating your data.** Question-behavior effect means the act of asking changes retention — so "goal-setters retain better" dashboards are doubly confounded. Only randomized ask/don't-ask designs settle it; almost nobody runs them.
4. **Asking as anti-creepiness technology.** Aguirre's paradox: identical personalization, opposite reaction depending on whether data was volunteered or harvested. The quiz's underrated function is making personalization *legible*. (Ties to GDPR/consent line from the forms video.)
5. **The quiz as survivorship filter.** Hreha's Noom point: long onboarding manufactures impressive outcome stats by shedding everyone else. Any "our program works" claim downstream of a brutal funnel is selection, not treatment.
6. **The confidence gap, part 3.** As with progress bars: zero public controlled experiments isolate "personalization quiz → retention." A running honesty franchise for the channel.

# PART 5 — BLACKLIST

Never say on camera:

| Claim | Why |
|---|---|
| "Quizzes convert 30–40% vs 6% for forms" (as research) | Vendor marketing with selection bias; attribute to Interact/Unbounce or skip |
| "Duolingo increased signups 20% by delaying signup" | Metric is DAU, not signups; +8.2% is a separate later result |
| Any Birdbrain/AI retention percentage ("35% better retention," "17% higher success") | SEO blogspam; Duolingo published no such numbers |
| "Users decide in the first 50 milliseconds" (as abandonment) | Lindgaard measured visual-appeal rating consistency only |
| "Users decide in the first 3/7 seconds (or 3 days)" | No traceable source |
| "Facebook proved 7 friends in 10 days causes retention" | Correlational heuristic from a 2012 talk; never published |
| "Onboarding should be ≤3 screens" | No research source exists; it's a design-template convention |
| "80%/91% of consumers demand personalization" (as behavior) | 2017–18 stated-preference vendor surveys |
| "Personalization drives 40% more revenue" (as measured) | McKinsey estimate, not an experiment |
| "The Noom quiz was ruled an illegal dark pattern" | Settlement was about auto-renewal/cancellation, not the quiz |
| "Severe initiation makes people love groups" (as settled) | Aronson & Mills is contested; modern replication failed |
| Guéguen's commitment/bystander replications | Author under research-integrity investigation |
| Cioffi & Garner "49% vs 17% showed up" | Figures unverifiable against the paper; use the qualitative finding |
| "Superhuman proved concierge onboarding works" | Self-reported, no control; say "Superhuman reports" |
| "X% of users abandon during onboarding" | Still no primary source (carried over from forms brief) |
| "Asking about intentions always increases behavior" | Repeated asking *decreased* purchases among low-intent users (Morwitz) |

# PART 6 — PRIMARY SOURCES

**Question-behavior / commitment**
- Morwitz, Johnson & Schmittlein (1993), *JCR* 20(1) — https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1424067
- Wood et al. (2016), *PSPR* 20(3) — https://pubmed.ncbi.nlm.nih.gov/26162771/ ; OA: https://eprints.whiterose.ac.uk/id/eprint/90306/
- Sherman (1980), *JPSP* 39(2) — https://psycnet.apa.org/record/1981-23660-001
- Moriarty (1975), *JPSP* 31(2) — https://www.scribd.com/document/622928083/moriarty1975
- Cialdini, Cacioppo, Bassett & Miller (1978), *JPSP* 36(5); low-ball meta-analysis (2016) — https://www.sciencedirect.com/science/article/abs/pii/S1162908816300470
- Cioffi & Garner (1996), *PSPB* 22(2) — https://journals.sagepub.com/doi/abs/10.1177/0146167296222003
- Wilding et al. (2016) QBE review, *ERSP* — https://www.tandfonline.com/doi/full/10.1080/10463283.2016.1245940

**Effort effects**
- Norton, Mochon & Ariely (2012), *JCP* 22(3) — https://www.hbs.edu/ris/Publication%20Files/11-091.pdf
- Sarstedt, Neubert & Barth (2017), conceptual replication — via https://en.wikipedia.org/wiki/IKEA_effect
- Aronson & Mills (1959), *J. Abnormal & Social Psych* 59(2); Gerard & Mathewson (1966), *JESP* 2(3) — https://www.sciencedirect.com/science/article/abs/pii/0022103166900849
- Buell & Norton (2011), labor illusion — see progress-bars brief

**Personalization**
- Tam & Ho (2005), *ISR* 16(3); Tam & Ho (2006), *MISQ* 30(4) — https://aisel.aisnet.org/misq/vol30/iss4/6/
- Matz, Kosinski, Nave & Stillwell (2017), *PNAS* 114(48) — https://www.pnas.org/doi/10.1073/pnas.1710966114
- Eckles et al. critique — https://www.pnas.org/content/115/23/E5254 ; Sharp et al. reply thread — https://www.pnas.org/doi/10.1073/pnas.1811106115
- Aguirre et al. (2015), *Journal of Retailing* 91(1) — https://openaccess.city.ac.uk/id/eprint/15747/1/AGUIRRE%20et%20al%20%202015%20(2).pdf
- Locke & Latham (2002), *American Psychologist* 57(9) — https://www-2.rotman.utoronto.ca/facbios/file/09%20-%20Locke%20&%20Latham%202002%20AP.pdf
- Lindgaard et al. (2006), *BIT* 25(2) — https://www.researchgate.net/publication/220208334

**Company-published / practitioner**
- First Round Review — Gotthilf/Duolingo A/B tests — https://review.firstround.com/the-tenets-of-a-b-testing-from-duolingos-master-growth-hacker/
- First Round Review — Superhuman onboarding playbook — https://review.firstround.com/superhuman-onboarding-playbook/
- Rahul Vohra PMF engine — https://review.firstround.com/how-superhuman-built-an-engine-to-find-product-market-fit/
- Hreha, Noom onboarding critique — https://www.thebehavioralscientist.com/articles/noom-product-critique-onboarding
- Noom $56M settlement — https://journals.law.unc.edu/ncjolt/blogs/diet-app-noom-agrees-to-pay-56-million-to-settle-class-suit/ ; https://www.huntonprivacyblog.com/2022/02/22/fitness-app-agrees-to-pay-56-million-to-settle-class-action-alleging-dark-pattern-practices/
- FTC dark-patterns enforcement policy (Oct 2021) — https://www.ftc.gov/news-events/news/press-releases/2021/10/ftc-ramp-enforcement-against-illegal-dark-patterns-trick-or-trap-consumers-subscriptions

**Vendor telemetry / benchmarks (label as such)**
- OpenView 2023 Product Benchmarks — https://openviewpartners.com/2023-product-benchmarks/
- Localytics app abandonment (2016–2018) — http://info.localytics.com/blog/23-of-users-abandon-an-app-after-one-use
- Interact Quiz Conversion Report — https://www.tryinteract.com/blog/quiz-conversion-rate-report/
- SurveyMonkey completion-rate data — https://www.surveymonkey.com/curiosity/survey_completion_times/ ; https://www.surveymonkey.com/curiosity/survey_questions_and_completion_rates/
- McKinsey Next in Personalization (2021) — https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-value-of-getting-personalization-right-or-wrong-is-multiplying

**Cross-references to prior briefs**
- Choice overload, defaults, social login, FITD, Quettra 77%, Amplitude 69% → decision-fatigue-research-brief.md
- Endowed progress, goal gradient, Zeigarnik/Ovsiankina, labor illusion, Liu & Wronski, Duolingo streak tests → progress-bars-research-brief.md
