# The Onboarding Quiz Isn't Really Asking For Your Data
### Why the best apps interview you — and when the interview backfires

**Format:** target 8–12 min | **Audience:** builders designing signup/onboarding flows
**Core thesis (say it 3 times):** *An onboarding question isn't data collection. It's an intervention. Asking changes the asker — the quiz is the commitment, not the data.*

Every number fact-checked against `research/onboarding-research.md`; original pattern data from `mobbin/onboarding-mobbin-audit.md` (n=35 flows + 18 screens, top iOS apps, Aug 2026). Prior-video overlap guard: no choice-overload, defaults, endowed-progress, or progress-bar beats — those are the forms and progress-bars videos. ⚠️ margin notes not spoken.

---

## [0:00 – 0:50] COLD OPEN · ~120w

**ON SCREEN:** Beach illustration. A towel, a radio, a stranger.

> Jones Beach, New York, 1972. A researcher spreads out a towel, turns on a radio, and after a while, walks off. A confederate strolls up and steals the radio.
>
> When bystanders had just been minding their own business: **four out of twenty** did anything about it.
>
> But when the researcher had first asked one question — *"would you watch my things?"* — **nineteen out of twenty** intervened. Some literally chased the thief down the beach.

**ON SCREEN:** "One question. 20% → 95%."

> Nobody was paid. Nobody was persuaded. They were just *asked* — and the asking changed what they did.
>
> That mechanism — questions as interventions — is the real reason the best apps in the world open with a quiz. And it's also why most onboarding quizzes are built completely backwards.

> ⚠️ *Moriarty 1975, JPSP. Say the raw counts (19/20 vs 4/20) — small cells, and stating them is the honesty move. Do NOT cite the French replications (Guéguen — integrity investigation).*

---

## [0:50 – 1:25] ROADMAP · ~90w

> Three parts.
>
> **One** — the science of asking: what happens in someone's head when they answer a question about themselves, and the numbers behind it — including one from a 40,000-household experiment.
>
> **Two** — the trap: why a quiz that works at three questions becomes actively destructive at fourteen. There's a specific psychological cliff, and I'll show you where it is.
>
> **Three** — we sampled 35 onboarding flows from top iOS apps ourselves. What the best ones actually do, what the manipulative ones do — and then we rebuild a bad onboarding into a good one, screen by screen.

---

## [1:25 – 3:45] ACT ONE — ASKING IS AN INTERVENTION · ~340w

**ON SCREEN:** "40,000+ households. One question."

> Here's the cleanest version of this effect ever measured. In the early nineties, researchers got access to giant consumer panels — over forty thousand households analyzed. Some households were asked a single question: *"Do you intend to buy a car in the next six months?"* That's it. No ad, no pitch, no follow-up.
>
> Among households that weren't asked, **2.4 percent** bought a car. Among households that were: **3.3 percent**.
>
> One survey question — roughly a **35 percent relative increase** in actual car purchases. It's called the **mere-measurement effect**, and it's been meta-analyzed across 116 experiments since: small effect, consistently real, strongest when people predict their *own* behavior.

**ON SCREEN:** "31% vs 4%"

> And that self-prediction detail is the key to onboarding. In 1980, Steven Sherman called residents and asked them to *predict* whether they'd volunteer three hours for the Cancer Society. People over-predict their own virtue — most said yes. Called back days later with the real request: **31 percent** of the predictors actually volunteered, versus **4 percent** of people who were simply asked directly.
>
> He called it a **self-erasing error**: people make an inaccurate prediction about themselves — and then behave until it's true.

**ON SCREEN:** A goal-selection screen: "What's your goal?" → tappable chips.

> Now look at any modern goal-selection screen with those numbers in mind. When your user taps *"Run a 5K"* or *"15 minutes a day,"* that's not a database write. That's a **self-prediction** — the exact instrument from these studies, at scale.
>
> This is the part almost everyone gets wrong about onboarding quizzes. Builders think the quiz's value is the *answers* — the personalization payload. But the research says a huge part of the value fires **the moment the question is answered**, whether or not you ever use the data. **The quiz is the commitment, not the data.**
>
> And I'll be straight with you: no public controlled study has ever isolated "personalization quiz → retention" in an app. The strongest evidence is the psychology of asking itself — small, real, and meta-analyzed — plus one company's A/B program we'll get to in a minute.

> ⚠️ *Morwitz, Johnson & Schmittlein 1993 (3.3% vs 2.4%); Wood et al. 2016 meta: d+ = 0.24 overall, 0.29 self-prediction; Sherman 1980 (31%/4%). The honesty line mirrors the "confidence gap" franchise from the progress-bars video.*

---

## [3:45 – 5:40] ACT TWO — THE COMPLETION CLIFF · ~300w

**ON SCREEN:** IKEA box, half-assembled.

> But there's a second force in every quiz, and it has a failure mode.
>
> In 2012, Norton, Mochon and Ariely had people assemble IKEA boxes, fold origami, build Lego. Then they measured what people would pay for their own creations. Builders bid **63 percent more** for a box *they* had assembled than others would pay for an identical pre-built one. Amateur origami folders valued their crumpled work near what other people would pay for an **expert's**. The IKEA effect: labor leads to love.
>
> But read the paper to the end, because the authors did the destructive version too. When people built something and then **didn't finish** — or watched it get taken apart — the effect **vanished**. Their words: labor leads to love *"only when labor results in successful completion."*

**ON SCREEN:** "Incomplete effort = pure cost." A quiz abandoned at question 14 of 20.

> Now apply that to your onboarding. Every question a user answers is invested effort — which converts into attachment **only if they finish and see what their effort built**. A quiz abandoned at question fourteen isn't a partial win. It's *worse than no quiz*: you charged the user effort and handed them nothing.
>
> So every question you add does double damage: it raises the odds of abandonment — survey telemetry shows completion falling steeply as question counts climb into the teens — *and* it raises the stakes of abandonment at the same time. That's the completion cliff.
>
> Which means the payoff screen — *"Here's your plan, built from your answers"* — is not decoration. It's the **completion artifact**. It's the moment invested effort becomes ownership. An onboarding quiz without a visible payoff screen is an IKEA box you make the user assemble and then never let them see.

> ⚠️ *Norton, Mochon & Ariely 2012, JCP — 63% figure is the boxes study; "successful completion" is near-verbatim. Length dose-response is SurveyMonkey vendor telemetry — hence "survey telemetry," no fake precision. Labor-illusion/"analyzing your answers" beats live in the progress-bars video; don't rehash.*

---

## [5:40 – 7:45] ACT THREE — THE REAL WORLD: DUOLINGO, NOOM, AND OUR 35 FLOWS · ~330w

**ON SCREEN:** Duolingo onboarding.

> The best real-world numbers come from Duolingo, because they actually publish their A/B results. Their most famous onboarding test wasn't *adding* a question — it was **deleting the signup screen**. Moving account creation from the front of the flow to later — after you'd already done a lesson — raised daily active users by about **20 percent**. Softening the wall further — turning "sign up or lose your progress" into a quiet "Later" — added **another 8 percent**.
>
> Notice what that means. The questions that *survived* at the front of their flow are the commitment questions — your goal, your daily minutes. The question they evicted was the **extraction** question: *"give us your email."* Commitment first, extraction later, after you've experienced value.

> ⚠️ *Gotthilf/First Round: "+20% DAU" — DAU, not signups; +8.2% is the separate soft-wall result. Company-published, no CI. Phrase as "Duolingo reports."*

**ON SCREEN:** "We sampled 35 onboarding flows (top iOS apps, Aug 2026)."

> So we sampled thirty-five onboarding flows from top iOS apps ourselves. Three findings you can use.
>
> **One:** the infamous quiz-then-hostage sequence — question tunnel, "your plan is ready," pay or leave — showed up in about **one in six** flows, almost entirely wellness apps. Meanwhile commerce and media apps let you straight in, and **29 percent** made accounts fully deferrable. The account as save-mechanism, not toll booth, is how confident products behave.
>
> **Two:** inside the quizzes, only **one in four** let you skip a question. And the long quizzes have started interleaving *persuasion between the questions* — social proof cards, Harvard-and-Oxford authority badges, one app literally has you **finger-sign a commitment contract** mid-quiz. The quiz has become the sales letter.
>
> **Three:** the raw permission ambush is dead — **90 percent** of notification asks we saw were primed with a benefit screen first. The floor has risen; the differentiator now is *what* you promise.
>
> And the cautionary tale: Noom rode the mega-quiz to a billion-dollar funnel — and to a **56-million-dollar settlement**. To be precise: the settlement was about auto-renewal and cancellation friction, not the quiz itself. But one behavioral scientist's documented teardown of that flow — roughly **45 minutes of questions** before meaningful value — makes the deeper point: past a certain length, a quiz stops being personalization and becomes a *filter* that only the desperate survive. Filters produce great-looking outcome stats. They just don't produce great products.

> ⚠️ *Audit: hostage 5/35=14% ("about 1 in 6" spoken); deferrable 10/35=29%; skip 4/16=25%; priming 9/10=90%; contract = Liven. Noom: $56M+$6M credits, S.D.N.Y. 2022 — mechanics not quiz; Hreha teardown labeled "one behavioral scientist's documented teardown."*

---

## [7:45 – 10:20] ACT FOUR — THE TEARDOWN · ~380w

**ON SCREEN:** Interactive demo, both phones. Fictional running app, "Stride." Highlight each beat.

> Let's rebuild one. Fictional running app — Stride. The before screen is a collage of everything we just measured going wrong.

**HIGHLIGHT: the signup wall**

> **Screen one, before:** name, email, password, confirm password — a toll booth before the user has seen anything. That's the exact screen Duolingo deleted for a 20 percent DAU gain. After: the app opens **into the quiz** — and the account ask moves to the end, reframed: *"Save your plan."* Signup as save mechanism. Same fields, opposite meaning.

**HIGHLIGHT: the question tunnel**

> **The quiz itself.** Before: fourteen questions, no progress indicator, no skip, including a gem like "What's your favorite running shoe brand?" — a question whose answer changes nothing. After: **three questions**, a visible "1 of 3," and skip on every screen. Each question passes the test from the forms video — *something visibly changes based on the answer* — and one of them is doing double duty.

**HIGHLIGHT: the goal question**

> Because question one is the commitment question: *"What's your goal?"* — run a 5K, get consistent, clear my head. The user **taps it themselves**. Research on active choice suggests a choice you physically make sticks harder than one made for you — this is the one screen where I *won't* tell you to preselect a default. This tap is Sherman's self-prediction. Let them make it.

**HIGHLIGHT: the payoff screen**

> **The payoff.** Before: quiz ends at "You're all set!" — into an empty dashboard. Effort collected, receipt never issued. After: *"Your Week 1 plan"* — three runs, on the days they chose, at the level they chose, with their goal printed at the top. This is the completion artifact. The IKEA effect needs a finished box; this is the box.

**HIGHLIGHT: the notification ask**

> **The permission ask.** Before: raw iOS dialog, no context — the ambush that's nearly extinct among top apps. After: a primer first — *"Tuesday, 7am — want a reminder before your first run?"* It quotes **their own answers back to them**. A permission request that's downstream of the user's stated goal isn't an interruption; it's service.

**HIGHLIGHT: the counter**

> Count it. Before: seventeen screens, zero skips, zero payoff, account demanded upfront. After: **three questions the product visibly uses, one self-prediction, one receipt, one earned permission** — and the account ask waiting politely at the end, holding something worth saving.

---

## [10:20 – 11:00] CLOSE · ~110w

**ON SCREEN:** The payoff screen. Hold.

> The quiz was never really about your data. Asking changes the asker — one question moved car purchases 35 percent, one question turned bystanders into thief-chasers.
>
> So here's the rule: **ask questions you'll visibly use, and show the receipt.**
>
> Three questions that shape the product beat fourteen that shape a marketing profile. And nothing you collect matters if the user never sees what their answers built.
>
> Open your onboarding tonight. For every question, ask: what changes based on the answer — and *does the user see it change?* Every question that fails both tests is just friction wearing a lab coat.

**[CTA / outro]**

---
---

# APPENDIX A — Optional expansion beats

### A1. The personalization paradox (+60 sec) — insert at top of Act Three
Aguirre et al. 2015, *Journal of Retailing*: identical personalization, opposite reactions depending on data provenance — click-through drops sharply when people realize data was collected covertly; overt asking plus trust cues flips it positive. The quiz's most underrated function: it's the *disclosure mechanism* that makes personalization feel legitimate instead of creepy. "Personalization built on surveillance gets the creepiness penalty; personalization built on asking gets a pass — the quiz IS the consent." Strongest cut material; restore first.

### A2. The contaminated-dashboard warning (+40 sec) — insert at end of Act One
Mere-measurement means asking about goals *changes* retention — so your "users who set goals retain 30% better" dashboard is doubly confounded: self-selection plus the question-behavior effect itself. Only randomizing whether the question is asked isolates it, and almost nobody runs that test. For the analytics-literate slice of the audience, this beat alone is worth the subscribe.

### A3. Superhuman: question-asking at maximum fidelity (+45 sec) — insert in Act Three after Duolingo
Superhuman reports 65%+ of human-onboarded users fully switched their email client vs ~30% self-serve — the onboarding was a 30-minute *interview*, a human asking questions and configuring the product to the answers. Self-reported, no control group — but as an existence proof of the ceiling: the more faithfully answers visibly reshape the product, the more the asking pays. They wound it down when the cost beat the benefit, which is the honest coda.

---

# APPENDIX B — Title, thumbnail, chapters

**Titles**
1. `The Onboarding Quiz Isn't Really Asking For Your Data`
2. `Why the Best Apps Interview You Before Letting You In`
3. `One Question Raised Car Sales 35% (The Science of Onboarding Quizzes)`
4. `Your Onboarding Quiz Is Backwards`

**Thumbnail:** A quiz screen ("What's your goal?") with a magnifying glass revealing the word "COMMITMENT" underneath the options. Or split: 14-question tunnel (red, "FILTER") vs 3 questions + plan (green, "RECEIPT").

**Chapters**
```
0:00  One question: 20% → 95%
0:50  What we're covering
1:25  Asking is an intervention (40,000 households)
2:50  Self-erasing errors: the goal screen is a prophecy
3:45  The IKEA effect's fine print: the completion cliff
5:40  Duolingo deleted the signup screen
6:30  We sampled 35 onboarding flows — what top apps do
7:45  Teardown: rebuilding Stride's onboarding
10:20 The rule: visible use + the receipt
```

---

# APPENDIX C — Ranked cut list

Baseline ~10:50 (≈1,630 spoken words @150wpm).

| # | Cut | Saves | Cost |
|---|---|---|---|
| 1 | The shoe-brand joke + "double duty" line (Act Four, question tunnel) | 0:10 | Minor color. |
| 2 | Sherman beat compressed to one line ("people over-predict their virtue, then behave to make it true") | 0:25 | Loses the 31/4 numbers; keep the concept — Moriarty already carries the shock. |
| 3 | The Noom paragraph | 0:35 | Loses the cautionary tale + filter concept. If cut, move the "filter" line into Act Two's cliff beat. |
| 4 | The finger-signed contract + interleaved persuasion detail (Act Three) | 0:20 | The most memeable original finding — cut last among Act Three items. |
| 5 | Deferred-signup stat (29%) | 0:10 | The Duolingo beat still carries the point. |

**Below ~9:00 you're cutting teaching.** For 8:00 flat: cut 1–5 and compress the roadmap to two sentences.

---

# APPENDIX D — Production notes

**Every number, with its source**

| Claim | Source | Tier |
|---|---|---|
| 19/20 vs 4/20 beach intervention | Moriarty 1975, *JPSP* 31(2) | SOLID (small cells — state raw counts) |
| 3.3% vs 2.4% car purchase, 40k+ households | Morwitz, Johnson & Schmittlein 1993, *JCR* 20(1) | SOLID |
| 116-test meta, d+=0.24 (0.29 self-prediction) | Wood et al. 2016, *PSPR* 20(3) | SOLID |
| 31% vs 4% volunteering | Sherman 1980, *JPSP* 39(2) | SOLID |
| IKEA effect: 63% higher bids; completion boundary | Norton, Mochon & Ariely 2012, *JCP* 22(3) | SOLID |
| Quiz completion falls steeply past ~15 questions | SurveyMonkey telemetry | CONTESTED — say "survey telemetry" |
| Duolingo +20% DAU (delayed signup), +8.2% (soft wall) | Gotthilf / First Round Review | CONTESTED — company-published; "Duolingo reports"; metric is DAU |
| Audit: hostage 14%, deferrable 29%, skip 25%, priming 90%, contract-signing (Liven), carousels 11% | Our Mobbin audit, n=35 flows, Aug 2026 | Original data — state sample honestly |
| Noom $56M settlement (auto-renewal mechanics) | S.D.N.Y. 2022 settlement + FTC dark-patterns policy | SOLID — do NOT say "quiz ruled dark pattern" |
| ~45-minute Noom onboarding, filter argument | Hreha teardown | CONTESTED — "one behavioral scientist's documented teardown" |
| Active choice > passive choice | Cioffi & Garner 1996, *PSPB* 22(2) | SOLID qualitative — do NOT state "49% vs 17%" (unverifiable) |

**Delivery notes**
- The "quiz is the commitment, not the data" thesis line lands three times: Act One close, Act Four goal beat, Close. Same words each time.
- Act One's "no controlled study exists" line is the credibility beat — drop graphics, direct to camera, matching the prior videos' honesty-franchise style.
- Never say: "quizzes convert 30–40% vs forms," any Birdbrain number, "users decide in 50ms/3 seconds/3 days," "7 friends in 10 days proved," "80%/91% demand personalization." Full blacklist in the brief.

**Demo spec** — see `demo-spec.md` in this folder; walkthrough order matches Act Four highlights.
