Research Brief — Progress Bars & Onboarding Motivation
Backing document for the video script. Includes corrections to the source notes you sent.
Confidence key
SOLID — primary source located, figure verified
CONTESTED — real source exists, but methodology or interpretation is disputed
SHAKY — widely repeated, source is weak, absent, or circular
BLACKLISTED — never repeat this claim
Verification status: every number in the script and this brief was independently checked against primary sources in a second pass. Nine errors were found and corrected — including a bad multiplier derivation, an odds-vs-frequency conflation, and one on-screen card that contradicted the paper it cited. Two claims remain flagged as unverified and are marked ️ inline. Everything else traces to a named source.
PART 1 — Four corrections to what you sent me
These matter because they're the kind of thing a comment section catches.
1. Your loyalty card stat is a blend of two different studies
You wrote: "Journal of Consumer Research found customers who received a loyalty card with 2 out of 10 punches already used were 2x more likely to complete."
There are two separate 2006 papers here and the internet merges them constantly:
| Nunes & Drèze 2006 (JCR) | Kivetz, Urminsky & Zheng 2006 (JMR) | |
|---|---|---|
| Setting | Car wash | Café |
| Control | 8-stamp card | 10-stamp card |
| Treatment | 10-stamp card, 2 pre-stamped | 12-stamp card, 2 pre-stamped |
| Real work | 8 washes both arms | 10 purchases both arms |
| n | 300 | 108 |
| Measured | Completion: 34% vs 19% | Speed: 12.7 vs 15.6 days |
| Confidence | SOLID |
SOLID |
So the "2 of 10 punches" card is Nunes & Drèze — but that one is 10 stamps vs an 8-stamp control, not vs a 10-stamp control. And Kivetz's card is 2 of 12, and it measured speed, not completion. SOLID (that the correction is right)
Say it this way on camera: "Ten stamps with two already filled in, versus eight stamps starting empty. Both need eight washes."
2. "2x" is an overstatement — use "nearly double"
34 ÷ 19 = 1.79×. Accurate framings: "nearly double," "a 79% relative increase," or just say the two numbers. SOLID — arithmetic on the paper's own figures. "Twice as likely" is the single most-repeated distortion of this study and it's checkable in ten seconds — BLACKLISTED as a framing.
Also worth knowing: 34% and 19% are redemption rates across all 150 cards handed out per arm, not among people who started using them. Only 80 of the 300 cards were ever redeemed. SOLID
3. Hull's rats are 1934, not 1932
- Hull 1932, Psychological Review — the goal gradient hypothesis as theory, derived from existing maze data. Contains the prediction, not the experiment.
SOLID - Hull 1934, Journal of Comparative Psychology — the actual experiment. A straight runway with electrical contacts every six feet measuring locomotion speed. This is where "the rats ran faster near the food" comes from.
SOLID
Nearly every blog cites 1932 for the rat experiment. If you cite the speed finding, cite 1934. (The script handles this with a small correction card, which is also a nice credibility signal.)
4. The LinkedIn example — I could not verify it, and you should not use it as stated
You wanted to fill in "we can see this in action at ______" and mentioned another video used LinkedIn "always giving you some points no matter how much you've filled out."
I searched hard for a primary source and there isn't one. No LinkedIn statement, no engineering blog, no contemporaneous reporting showing LinkedIn seeds the meter with artificial progress. "LinkedIn starts you at 25%" is folklore. BLACKLISTED
What's verifiable, and one thing that isn't — read this before you film:
Corroborated across sources (SOLID):
- The floor is a named achievement, not a number. The lowest rung is "Beginner." You are never shown as 0%.
- Your profile level is only visible to you.
- At the top rung, the meter deletes itself rather than displaying as complete.
️ Do not state the tier count or the exact section list on camera without checking it yourself first. My sources conflict: one read of LinkedIn's own help page returned a three-tier ladder (Beginner → Intermediate → All-star, with Intermediate at 4 of 7 sections), but a January 2026 third-party write-up describes the older five-tier ladder (Beginner → Intermediate → Advanced → Expert → All-Star) with a different section list. LinkedIn blocks automated access, so I could not settle it. CONTESTED — the sources disagree and neither was settled. Open your own LinkedIn profile and look — it's a thirty-second check.
The good news: the point the example makes doesn't depend on the tier count. The honest line, safe either way:
"LinkedIn never shows you a zero. The lowest rung on the ladder is a label, not a number."
Also flag: the famous "All-Star profiles get 40× more opportunities" stat traces to a now-dead LinkedIn marketing page with no methodology, no sample, and no definition of "opportunities." Don't use it. BLACKLISTED
Better verifiable fills for that blank:
| Example | What's verifiable | Confidence |
|---|---|---|
| The car wash itself | Strongest. Real experiment, real numbers, and it's visual. Lead with this. | SOLID |
| Tiered ladder, floor is a label not a zero, meter vanishes at the top | CONTESTED — company help page only; brief notes it could not be settled. [grade needs Kris confirmation] |
|
| Starbucks Rewards | Redemption ladder at 25 / 60 / 100 / 200 / 300 / 400 stars — a near rung is always visible regardless of balance. That's a goal-gradient device by construction. | CONTESTED — no source cited in this brief [grade needs Kris confirmation] |
| Duolingo | Making the streak visible tested at +3% DAU, +1% D14; emphasizing it, +1% DAU / +3% D14. Company-published. Note the magnitudes are single digits. | CONTESTED — company-published, no methodology |
PART 2 — The strongest findings, ranked by how confidently you can state them
Tier 1 — State flatly on camera
| Finding | Numbers | Confidence |
|---|---|---|
| Endowed progress (car wash) | 34% vs 19%, n=300, χ²(1)=8.1, p<.01. Nunes & Drèze 2006 | SOLID |
| Goal gradient (café) | ~20% acceleration first→last purchase, 949 cards. Kivetz et al. 2006 | SOLID |
| Progress bars ≠ completion | Constant bar LOR=0.072, p=.365, k=18 (of 32 total experiments). Villar et al. 2013 | SOLID (meta-analytic null) |
| Slow-early bars hurt | Drop-off odds 1.56×, LOR=0.447, p=.001, k=7. Same meta-analysis. Say "odds" — it's an odds ratio, not a rate ratio | SOLID |
| Post-reward reset | 2.2/2.1 days → 3.1/2.7 days across the card boundary, n=110, p<.01 | SOLID — no source named in this brief. [grade needs Kris confirmation] |
| Goal failure is costly | n=95,532 hotel customers, 80% missed; failure reduced purchasing vs. matched comparison customers. Note: the promotion was net-positive overall. Wang et al. 2016 | SOLID |
| Small-area hypothesis | Field n=907 + lab replications. Koo & Fishbach 2012 | SOLID |
| Rewards vs. feedback | Completion-contingent d=−0.36; verbal/informational d=+0.33. Meta-analysis of 128 studies (behavioral measure available in 101). Deci, Koestner & Ryan 1999 | SOLID |
| First-question cost | −1.6pp per 50 words in the first question vs −0.5pp per 50 words of total survey text = 3.2×, not 8×. n=25,080 web surveys. Liu & Wronski 2018 | SOLID |
| Monzo funnel | 9% → 40% signup completion; one check caused 40% of all drop-off | CONTESTED — company-published blog post, no methodology, but placed in Tier 1 by this brief [grade needs Kris confirmation] |
| Progress framing only fires near the goal | 13,500 lapsed donors. At 85% progress: 1.17% vs 0.50% donation rate — more than doubled. At 66% and 10%: no significant effect. Cryder, Loewenstein & Seltman 2013 | SOLID |
Optional extra beat: that last row is the sharpest single caveat in the literature and it didn't make the script for time. If you want a sixth rule, it's "progress framing pays off most at the end — spend your encouragement budget on the last steps, not the first." It also sets up a natural sequel on activation vs. onboarding.
Tier 2 — Use, but hedge the framing
- Baymard figures (5.1 steps / 11.3 fields average; "fields matter more than steps"). Real research, 200k+ hours, but proprietary and not peer-reviewed, and they don't run A/B tests — it's qualitative usability testing plus benchmarking.
CONTESTED. Say "the leading e-commerce UX research group finds..." not "studies show." - Obama 2012 +5% — practitioner post, no sample size or confidence interval. Note that Rush's write-up never mentions a progress indicator at all; the test was purely about splitting one long form into four steps ordered by field error rate.
CONTESTEDas an illustration of the step-splitting principle,BLACKLISTEDas evidence about progress bars. - Fast-to-slow bars help (0.80×) — significant only after removing an outlier.
CONTESTED. The negative result (slow-to-fast hurts) is much more robust than the positive one. The script is written to lean on the negative. - Duolingo's "7-day streak → 3.6× more likely to complete the course" — real, company-published, but correlational. Duolingo never claimed causation; everyone else does.
CONTESTEDas a correlation,BLACKLISTEDas a causal claim.
Tier 3 — Do not use
- "LinkedIn starts you at 25%" — no source exists.
BLACKLISTED - "All-Star profiles get 40× more opportunities" — dead marketing page, no methodology.
BLACKLISTED - "Skeleton screens make apps feel faster" — no peer-reviewed support.
BLACKLISTED. The larger of the two grey-lit studies (n=136) found skeleton screens perceived as the slowest of three options —SHAKY(grey literature, unreplicated) - "Facebook proved skeleton screens work" — untraceable.
BLACKLISTED - "The Zeigarnik effect makes users finish checklists" —
BLACKLISTED. A 2025 meta-analysis of 59 publications found a pooled recall ratio of 0.99 (i.e. nothing)SOLID. The Ovsiankina effect (urge to resume an interrupted task, ~67% resumption rate) is the one that survivesSOLID, and it's actually the better fit for progress bars anyway - "Multi-step forms convert 86% better" — no date, no n, no methodology anywhere in the chain.
BLACKLISTED - Any number for "fake progress bars reduce trust by X%" — no such study exists. Nobody has run it.
BLACKLISTED. The ethical case has to be made from adjacent evidence (which the script does) - "Limit onboarding to 7±2 steps" — Miller's number is about short-term recall, and Nielsen explicitly rejects applying it to interfaces.
BLACKLISTED
PART 3 — Things you could build a second video on
The research surfaced several threads that are too big for this video but are strong standalone topics:
-
"Your progress bar is lying and it's costing you" — the honesty angle. There's a genuinely under-covered engineering point: a fake progress bar destroys your ability to signal failure. If the bar always advances, a stalled process looks identical to a working one.
-
The dark pattern line. California's statutory definition — "a user interface designed or manipulated with the substantial effect of subverting or impairing user autonomy" — is effects-based, not intent-based. Good intentions are not a defense.
SOLID. Meanwhile, notably, "fake progress" doesn't appear as a named type in either major dark-pattern taxonomy (Brignull's or Mathur et al.'s). It's an unclassified pattern, which is an interesting hook.CONTESTED— neither taxonomy is cited in this brief's source list, and an absence claim needs the two documents in hand [grade needs Kris confirmation] -
Why gamification often fails. A 16-week classroom study found students in the gamified course showed less motivation and satisfaction over time, with lower final exam scores mediated by intrinsic motivation. Pairs well with the d=−0.36 finding. Contrarian, well-evidenced, and directly relevant to builders bolting XP systems onto everything.
CONTESTED— single classroom study, no n or effect size given here, no replication noted [grade needs Kris confirmation] -
The confidence gap. Possibly the sharpest angle of all: I looked specifically for a controlled experiment on progress indicators in app signup/onboarding — as opposed to surveys — and found none. Industry confidence in this practice substantially exceeds what the literature supports. That gap is itself a video.
Sources
Primary literature
- Nunes & Drèze (2006), "The Endowed Progress Effect," JCR 32(4)
- Kivetz, Urminsky & Zheng (2006), "The Goal-Gradient Hypothesis Resurrected," JMR 43(1)
- Hull (1932), Psychological Review 39(1), 25–43 — doi:10.1037/h0072640
- Hull (1934), J. Comparative Psychology 17(3), 393–422 — doi:10.1037/h0071299
- Koo & Fishbach (2012), "The Small-Area Hypothesis," JCR 39(3)
- Villar, Callegaro & Yang (2013), "Where Am I?" SSCR 31(6)
- Conrad, Couper, Tourangeau & Peytchev (2010), Interacting with Computers 22(5)
- Liu & Wronski (2018), SSCR 36(1)
- Wang, Lewis, Cryder & Sprigg (2016), Marketing Science 35(4)
- Buell & Norton (2011), "The Labor Illusion," Management Science 57(9)
- Deci, Koestner & Ryan (1999), Psychological Bulletin 125(6)
- Ghibellini & Meier (2025), Zeigarnik/Ovsiankina meta-analysis, HSSC 12:962
- Cryder, Loewenstein & Seltman (2013), "Goal gradient in helping behavior," JESP 49(6), 1078–1083
- Bauer, Khamitov, Isaac & Sevilla (2026), "The visual moderation effect," JAMS 54(2)
- Hanus & Fox (2015), gamification longitudinal study, Computers & Education 80
Practitioner / company-published
- Monzo — "How we cut time to sign up from 17 minutes to 4 minutes"
- LinkedIn Help — "Your Profile level meter"
- Duolingo — "How the Duolingo streak builds habit"
- Econsultancy — Duolingo A/B test results (Zan Gilani, Canvas Conference)
- Kyle Rush — Obama 2012 A/B testing
- Baymard — checkout form fields
- NN/g — Progress Indicators
- Viget — "A Bone to Pick with Skeleton Screens"
- CPPA Enforcement Advisory 2024-02 (dark patterns)