# Research Brief — Progress Bars & Onboarding Motivation
### Backing document for the video script. Includes corrections to the source notes you sent.

**Confidence key**
`SOLID` — primary source located, figure verified
`CONTESTED` — real source exists, but methodology or interpretation is disputed
`SHAKY` — widely repeated, source is weak, absent, or circular
`BLACKLISTED` — never repeat this claim

> **Verification status:** every number in the script and this brief was independently checked against primary sources in a second pass. Nine errors were found and corrected — including a bad multiplier derivation, an odds-vs-frequency conflation, and one on-screen card that contradicted the paper it cited. Two claims remain flagged as unverified and are marked ⚠️ inline. Everything else traces to a named source.

---

## PART 1 — Four corrections to what you sent me

These matter because they're the kind of thing a comment section catches.

### 1. Your loyalty card stat is a blend of two different studies

You wrote: *"Journal of Consumer Research found customers who received a loyalty card with 2 out of 10 punches already used were 2x more likely to complete."*

There are **two separate 2006 papers** here and the internet merges them constantly:

| | Nunes & Drèze 2006 (*JCR*) | Kivetz, Urminsky & Zheng 2006 (*JMR*) |
|---|---|---|
| Setting | Car wash | Café |
| Control | 8-stamp card | 10-stamp card |
| Treatment | **10-stamp card, 2 pre-stamped** | **12-stamp card, 2 pre-stamped** |
| Real work | 8 washes both arms | 10 purchases both arms |
| n | 300 | 108 |
| Measured | **Completion: 34% vs 19%** | **Speed: 12.7 vs 15.6 days** |
| Confidence | `SOLID` | `SOLID` |

So the "2 of 10 punches" card is Nunes & Drèze — but that one is 10 stamps vs an **8**-stamp control, not vs a 10-stamp control. And Kivetz's card is 2 of **12**, and it measured *speed*, not completion. `SOLID` (that the correction is right)

**Say it this way on camera:** "Ten stamps with two already filled in, versus eight stamps starting empty. Both need eight washes."

### 2. "2x" is an overstatement — use "nearly double"

34 ÷ 19 = **1.79×**. Accurate framings: "nearly double," "a 79% relative increase," or just say the two numbers. `SOLID` — arithmetic on the paper's own figures. "Twice as likely" is the single most-repeated distortion of this study and it's checkable in ten seconds — `BLACKLISTED` as a framing.

Also worth knowing: 34% and 19% are redemption rates across **all 150 cards handed out per arm**, not among people who started using them. Only 80 of the 300 cards were ever redeemed. `SOLID`

### 3. Hull's rats are 1934, not 1932

- **Hull 1932**, *Psychological Review* — the goal gradient hypothesis as **theory**, derived from existing maze data. Contains the prediction, not the experiment. `SOLID`
- **Hull 1934**, *Journal of Comparative Psychology* — the **actual experiment**. A straight runway with electrical contacts every six feet measuring locomotion speed. This is where "the rats ran faster near the food" comes from. `SOLID`

Nearly every blog cites 1932 for the rat experiment. If you cite the speed finding, cite 1934. (The script handles this with a small correction card, which is also a nice credibility signal.)

### 4. The LinkedIn example — I could not verify it, and you should not use it as stated

You wanted to fill in *"we can see this in action at ______"* and mentioned another video used LinkedIn "always giving you some points no matter how much you've filled out."

**I searched hard for a primary source and there isn't one.** No LinkedIn statement, no engineering blog, no contemporaneous reporting showing LinkedIn seeds the meter with artificial progress. "LinkedIn starts you at 25%" is folklore. `BLACKLISTED`

**What's verifiable, and one thing that isn't — read this before you film:**

Corroborated across sources (`SOLID`):
- **The floor is a named achievement, not a number.** The lowest rung is "Beginner." You are never shown as 0%.
- Your profile level is **only visible to you.**
- At the top rung, **the meter deletes itself** rather than displaying as complete.

⚠️ **Do not state the tier count or the exact section list on camera without checking it yourself first.** My sources conflict: one read of LinkedIn's own help page returned a **three**-tier ladder (Beginner → Intermediate → All-star, with Intermediate at 4 of 7 sections), but a January 2026 third-party write-up describes the older **five**-tier ladder (Beginner → Intermediate → Advanced → Expert → All-Star) with a different section list. LinkedIn blocks automated access, so I could not settle it. `CONTESTED` — the sources disagree and neither was settled. **Open your own LinkedIn profile and look — it's a thirty-second check.**

The good news: the point the example makes doesn't depend on the tier count. The honest line, safe either way:

> **"LinkedIn never shows you a zero. The lowest rung on the ladder is a label, not a number."**

Also flag: the famous **"All-Star profiles get 40× more opportunities"** stat traces to a now-dead LinkedIn marketing page with no methodology, no sample, and no definition of "opportunities." Don't use it. `BLACKLISTED`

**Better verifiable fills for that blank:**

| Example | What's verifiable | Confidence |
|---|---|---|
| **The car wash itself** | Strongest. Real experiment, real numbers, and it's visual. Lead with this. | `SOLID` |
| **LinkedIn** | Tiered ladder, floor is a label not a zero, meter vanishes at the top | `CONTESTED` — company help page only; brief notes it could not be settled. [grade needs Kris confirmation] |
| **Starbucks Rewards** | Redemption ladder at 25 / 60 / 100 / 200 / 300 / 400 stars — a near rung is *always* visible regardless of balance. That's a goal-gradient device by construction. | `CONTESTED` — no source cited in this brief [grade needs Kris confirmation] |
| **Duolingo** | Making the streak visible tested at **+3% DAU, +1% D14**; emphasizing it, **+1% DAU / +3% D14**. Company-published. Note the magnitudes are single digits. | `CONTESTED` — company-published, no methodology |

---

## PART 2 — The strongest findings, ranked by how confidently you can state them

### Tier 1 — State flatly on camera

| Finding | Numbers | Confidence |
|---|---|---|
| **Endowed progress (car wash)** | 34% vs 19%, n=300, χ²(1)=8.1, p<.01. Nunes & Drèze 2006 | `SOLID` |
| **Goal gradient (café)** | ~20% acceleration first→last purchase, 949 cards. Kivetz et al. 2006 | `SOLID` |
| **Progress bars ≠ completion** | Constant bar LOR=0.072, **p=.365**, **k=18** (of 32 total experiments). Villar et al. 2013 | `SOLID` (meta-analytic null) |
| **Slow-early bars hurt** | Drop-off **odds** 1.56×, LOR=0.447, p=.001, k=7. Same meta-analysis. Say "odds" — it's an odds ratio, not a rate ratio | `SOLID` |
| **Post-reward reset** | 2.2/2.1 days → 3.1/2.7 days across the card boundary, n=110, p<.01 | `SOLID` — no source named in this brief. [grade needs Kris confirmation] |
| **Goal failure is costly** | n=95,532 hotel customers, 80% missed; failure reduced purchasing **vs. matched comparison customers**. Note: the promotion was net-positive overall. Wang et al. 2016 | `SOLID` |
| **Small-area hypothesis** | Field n=907 + lab replications. Koo & Fishbach 2012 | `SOLID` |
| **Rewards vs. feedback** | Completion-contingent d=−0.36; verbal/informational d=+0.33. Meta-analysis of 128 studies (behavioral measure available in 101). Deci, Koestner & Ryan 1999 | `SOLID` |
| **First-question cost** | −1.6pp per 50 words in the first question vs −0.5pp per 50 words of total survey text = **3.2×**, not 8×. n=25,080 web surveys. Liu & Wronski 2018 | `SOLID` |
| **Monzo funnel** | 9% → 40% signup completion; one check caused 40% of all drop-off | `CONTESTED` — company-published blog post, no methodology, but placed in Tier 1 by this brief [grade needs Kris confirmation] |
| **Progress framing only fires near the goal** | 13,500 lapsed donors. At **85%** progress: 1.17% vs 0.50% donation rate — more than doubled. At **66%** and **10%**: no significant effect. Cryder, Loewenstein & Seltman 2013 | `SOLID` |

> **Optional extra beat:** that last row is the sharpest single caveat in the literature and it didn't make the script for time. If you want a sixth rule, it's *"progress framing pays off most at the end — spend your encouragement budget on the last steps, not the first."* It also sets up a natural sequel on activation vs. onboarding.

### Tier 2 — Use, but hedge the framing

- **Baymard** figures (5.1 steps / 11.3 fields average; "fields matter more than steps"). Real research, 200k+ hours, but proprietary and **not peer-reviewed, and they don't run A/B tests** — it's qualitative usability testing plus benchmarking. `CONTESTED`. Say "the leading e-commerce UX research group finds..." not "studies show."
- **Obama 2012 +5%** — practitioner post, no sample size or confidence interval. Note that Rush's write-up never mentions a progress indicator at all; the test was purely about splitting one long form into four steps ordered by field error rate. `CONTESTED` as an illustration of the step-splitting principle, `BLACKLISTED` as evidence about progress bars.
- **Fast-to-slow bars help (0.80×)** — significant only after removing an outlier. `CONTESTED`. The *negative* result (slow-to-fast hurts) is much more robust than the positive one. The script is written to lean on the negative.
- **Duolingo's "7-day streak → 3.6× more likely to complete the course"** — real, company-published, but **correlational**. Duolingo never claimed causation; everyone else does. `CONTESTED` as a correlation, `BLACKLISTED` as a causal claim.

### Tier 3 — Do not use

- **"LinkedIn starts you at 25%"** — no source exists. `BLACKLISTED`
- **"All-Star profiles get 40× more opportunities"** — dead marketing page, no methodology. `BLACKLISTED`
- **"Skeleton screens make apps feel faster"** — no peer-reviewed support. `BLACKLISTED`. The larger of the two grey-lit studies (n=136) found skeleton screens perceived as the **slowest** of three options — `SHAKY` (grey literature, unreplicated)
- **"Facebook proved skeleton screens work"** — untraceable. `BLACKLISTED`
- **"The Zeigarnik effect makes users finish checklists"** — `BLACKLISTED`. A 2025 meta-analysis of 59 publications found a pooled recall ratio of **0.99** (i.e. nothing) `SOLID`. The *Ovsiankina* effect (urge to **resume** an interrupted task, ~67% resumption rate) is the one that survives `SOLID`, and it's actually the better fit for progress bars anyway
- **"Multi-step forms convert 86% better"** — no date, no n, no methodology anywhere in the chain. `BLACKLISTED`
- **Any number for "fake progress bars reduce trust by X%"** — no such study exists. Nobody has run it. `BLACKLISTED`. The ethical case has to be made from adjacent evidence (which the script does)
- **"Limit onboarding to 7±2 steps"** — Miller's number is about short-term recall, and Nielsen explicitly rejects applying it to interfaces. `BLACKLISTED`

---

## PART 3 — Things you could build a second video on

The research surfaced several threads that are too big for this video but are strong standalone topics:

1. **"Your progress bar is lying and it's costing you"** — the honesty angle. There's a genuinely under-covered engineering point: a fake progress bar destroys your ability to signal *failure*. If the bar always advances, a stalled process looks identical to a working one.

2. **The dark pattern line.** California's statutory definition — *"a user interface designed or manipulated with the substantial effect of subverting or impairing user autonomy"* — is **effects-based, not intent-based**. Good intentions are not a defense. `SOLID`. Meanwhile, notably, "fake progress" doesn't appear as a named type in either major dark-pattern taxonomy (Brignull's or Mathur et al.'s). It's an unclassified pattern, which is an interesting hook. `CONTESTED` — neither taxonomy is cited in this brief's source list, and an absence claim needs the two documents in hand [grade needs Kris confirmation]

3. **Why gamification often fails.** A 16-week classroom study found students in the gamified course showed *less* motivation and satisfaction over time, with lower final exam scores mediated by intrinsic motivation. Pairs well with the d=−0.36 finding. Contrarian, well-evidenced, and directly relevant to builders bolting XP systems onto everything. `CONTESTED` — single classroom study, no n or effect size given here, no replication noted [grade needs Kris confirmation]

4. **The confidence gap.** Possibly the sharpest angle of all: I looked specifically for a controlled experiment on progress indicators in *app signup/onboarding* — as opposed to surveys — and **found none.** Industry confidence in this practice substantially exceeds what the literature supports. That gap is itself a video.

---

## Sources

**Primary literature**

- [Nunes & Drèze (2006), "The Endowed Progress Effect," *JCR* 32(4)](http://msbfile03.usc.edu/digitalmeasures/jnunes/intellcont/Endowed%20Progress%20Effect-1.pdf)
- [Kivetz, Urminsky & Zheng (2006), "The Goal-Gradient Hypothesis Resurrected," *JMR* 43(1)](https://home.uchicago.edu/ourminsky/Goal-Gradient_Illusionary_Goal_Progress.pdf)
- Hull (1932), *Psychological Review* 39(1), 25–43 — [doi:10.1037/h0072640](https://doi.org/10.1037/h0072640)
- Hull (1934), *J. Comparative Psychology* 17(3), 393–422 — [doi:10.1037/h0071299](https://doi.org/10.1037/h0071299)
- [Koo & Fishbach (2012), "The Small-Area Hypothesis," *JCR* 39(3)](https://academic.oup.com/jcr/article-abstract/39/3/493/1822606)
- [Villar, Callegaro & Yang (2013), "Where Am I?" *SSCR* 31(6)](https://openaccess.city.ac.uk/id/eprint/14427/4/Social%20Science%20Computer%20Review-2013-Villar-744-62.pdf)
- [Conrad, Couper, Tourangeau & Peytchev (2010), *Interacting with Computers* 22(5)](https://academic.oup.com/iwc/article-abstract/22/5/417/688424)
- [Liu & Wronski (2018), *SSCR* 36(1)](https://journals.sagepub.com/doi/10.1177/0894439317695581)
- [Wang, Lewis, Cryder & Sprigg (2016), *Marketing Science* 35(4)](https://pubsonline.informs.org/doi/10.1287/mksc.2015.0966)
- [Buell & Norton (2011), "The Labor Illusion," *Management Science* 57(9)](https://www.hbs.edu/ris/Publication%20Files/Norton_Michael_The%20labor%20illusion%20How%20operational_f4269b70-3732-4fc4-8113-72d0c47533e0.pdf)
- [Deci, Koestner & Ryan (1999), *Psychological Bulletin* 125(6)](https://home.ubalt.edu/tmitch/642/articles%20syllabus/Deci%20Koestner%20Ryan%20meta%20IM%20psy%20bull%2099.pdf)
- [Ghibellini & Meier (2025), Zeigarnik/Ovsiankina meta-analysis, *HSSC* 12:962](https://www.nature.com/articles/s41599-025-05000-w)
- [Cryder, Loewenstein & Seltman (2013), "Goal gradient in helping behavior," *JESP* 49(6), 1078–1083](https://www.cmu.edu/dietrich/sds/docs/loewenstein/GoalGradBeh.pdf)
- [Bauer, Khamitov, Isaac & Sevilla (2026), "The visual moderation effect," *JAMS* 54(2)](https://link.springer.com/article/10.1007/s11747-025-01133-1)
- [Hanus & Fox (2015), gamification longitudinal study, *Computers & Education* 80](https://www.sciencedirect.com/science/article/abs/pii/S0360131514002000)

**Practitioner / company-published**

- [Monzo — "How we cut time to sign up from 17 minutes to 4 minutes"](https://monzo.com/blog/how-we-cut-time-to-sign-up-for-monzo-us-from-17-minutes-to-4-minutes)
- [LinkedIn Help — "Your Profile level meter"](https://www.linkedin.com/help/linkedin/answer/a594698)
- [Duolingo — "How the Duolingo streak builds habit"](https://blog.duolingo.com/how-duolingo-streak-builds-habit)
- [Econsultancy — Duolingo A/B test results (Zan Gilani, Canvas Conference)](https://econsultancy.com/six-a-b-tests-used-by-duolingo-to-tap-into-habit-forming-behaviour/)
- [Kyle Rush — Obama 2012 A/B testing](https://kylerush.net/blog/optimization-at-the-obama-campaign-ab-testing/)
- [Baymard — checkout form fields](https://baymard.com/blog/checkout-flow-average-form-fields)
- [NN/g — Progress Indicators](https://www.nngroup.com/articles/progress-indicators/)
- [Viget — "A Bone to Pick with Skeleton Screens"](https://www.viget.com/articles/a-bone-to-pick-with-skeleton-screens)
- [CPPA Enforcement Advisory 2024-02 (dark patterns)](https://www.cppa.ca.gov/pdf/enfadvisory202402.pdf)
