# Your Dashboard Is a Feedback Intervention (And 38% of Those Backfire)
### The science of progress displays that retain users instead of burning them out

**Format:** target 8–12 min | **Audience:** builders designing home screens, dashboards, stats surfaces
**Core thesis (say it 3 times):** *A dashboard isn't a report. It's a feedback intervention — and feedback about the work works, while feedback about the worker backfires.*

Every number fact-checked against `research/dashboards-research.md`; original pattern data from `mobbin/dashboards-mobbin-audit.md` (n=76 screens, top apps, Aug 2026). Cross-reference guard: goal-gradient/endowed-progress/small-area beats belong to the progress-bars video — not rehashed here. ⚠️ margin notes not spoken.

---

## [0:00 – 0:55] COLD OPEN · ~130w

**ON SCREEN:** A person walking. A step counter ticking up.

> Researchers gave people pedometers for a day. Simple field experiment. The people wearing a counter **walked more** — measurably more steps than the control group.
>
> A win, right? Keep watching. The same people **enjoyed walking less**. Their end-of-day happiness was *lower*. And here's the detail that should worry you: participants who *volunteered* to wear the pedometer were hurt just the same.
>
> Then the researchers ran the kicker study with reading. While a page counter was visible, people read more. Then the counter was removed, and everyone got a free choice: keep going or stop.
>
> **48.5 percent** of the never-measured group kept reading. Of the measured group: **27.3**.

**ON SCREEN:** "The measurement worked. That's the problem."

> The counter boosted the metric and nearly halved the desire to continue. If you build dashboards, that sentence should keep you up at night.

> ⚠️ *Etkin 2016, JCR — Exp 2 (n=95, field): log-steps 7.97 vs 7.01, p=.003; enjoyment 4.82 vs 5.33, p=.030. Exp 6 (n=66): 27.3% vs 48.5%, p=.041. All figures verified against the paper.*

---

## [0:55 – 1:30] ROADMAP · ~85w

> Four parts.
>
> **One** — the biggest review of feedback ever conducted, and its most uncomfortable number. Your dashboard is a feedback intervention whether you meant it that way or not.
>
> **Two** — a moderator table from that review that reads like a dashboard design spec. Which widgets carry six times the effect of others.
>
> **Three** — what we found tallying 76 real home screens and dashboards from top apps — including the one thing consumer apps and SaaS dashboards do *completely* differently.
>
> **Four** — the teardown: rebuilding an overloaded money dashboard, widget by widget.

---

## [1:30 – 4:10] ACT ONE — FEEDBACK ABOUT THE WORK vs THE WORKER · ~400w

**ON SCREEN:** "131 papers · 607 effect sizes · 12,652 participants"

> In 1996, Kluger and DeNisi published the definitive meta-analysis of feedback interventions: 131 papers, 607 effect sizes, twelve and a half thousand participants — nearly a century of "let's show people how they're doing" research.
>
> Average effect: positive, moderate — d of 0.41. Feedback helps, on average.
>
> But averages hide the story. **More than 38 percent of the effects were negative.** In over a third of the experiments, showing people feedback about their performance made them perform *worse*. And it's not one weird lab — remove the biggest outlier researcher entirely and a third of the remaining effects are still negative.

**ON SCREEN:** Two arrows from a number: → THE WORK / → THE WORKER

> Their theory of when it backfires is beautifully simple. Feedback pulls attention to one of two places: **the task** — what happened, what to adjust — or **the self** — what does this say about *me*? The closer attention moves to the self, the worse feedback performs.
>
> **Feedback about the work works. Feedback about the worker backfires.**

**ON SCREEN:** The moderator table, rows animating in with widget mockups beside each.

> And the moderator table reads like a design spec, because every row *is a dashboard widget*.
>
> **Velocity feedback** — how you're doing versus last time — carries the biggest effect in the table: d = .55. That's your delta chip: *"up 12 percent from last week."*
>
> **Feedback that includes the next action** — d = .43. Not "your number is low," but "here's the thing to do about it."
>
> Now the bottom of the table. **Praise: d = .09.** Statistically, the confetti animation and the "Great job!!" toast are doing almost nothing. **Feedback that threatens self-esteem — comparisons, ranks, grades: .08.** And **discouraging feedback: negative** — shame-framed metrics actively hurt.
>
> There's a classroom study that makes it visceral. Schoolkids got returned work three ways: comments only, grades only, or grades *plus* comments. Comments alone improved performance. And adding a grade *next to* the comment **wiped out the comment's benefit** — the ego-number captured all the attention, and the useful information got ignored. That's what your big red score does to the helpful diagnostics sitting right under it.

> ⚠️ *K&D: >38% negative (say "more than a third — 38%"); do NOT say 23,000 participants (that's observations). Moderators: velocity .55 vs .28; solution .43 vs .25; praise .09 vs .34; self-esteem threat .08 vs .47; discouragement −.14. Butler 1988 (~200 students) — present as illustration, single pre-registration-era study.*

---

## [4:10 – 5:40] ACT TWO — THE HIDDEN COST OF COUNTING · ~230w

**ON SCREEN:** Coloring books, pedometers, page counters.

> Back to that pedometer study, because the mechanism matters. Across six experiments — coloring, walking, reading — the pattern held: measurement increases output and **drains enjoyment**. Why? Measured, an activity starts to *feel like work*. Attention shifts from the experience to the number. Intrinsic motivation — the reason they'd come back tomorrow — leaks out.
>
> But experiment four found the boundary, and it's the design rule: when reading was framed *as* work — as learning — the damage disappeared. **Quantifying things people already treat as work is safe. Quantifying what they do for joy converts play into a job.**
>
> A sales pipeline, invoices, training plans: measure freely. Leisure reading, hobby learning, meditation minutes? Every counter you add is charging interest against the reason they showed up.
>
> And the long-run version shows up in hardware: forget the famous "a third of wearables get abandoned" line — that's a consultancy survey with an undisclosed sample. The honest number is from researchers who gave **711 people Fitbits and read the actual device logs for 320 days**: about three-quarters still tracking at day 100... **16 percent** at day 320.
>
> Metrics-first products don't just plateau. They actively spend down the intrinsic motivation they were built to display.

> ⚠️ *Etkin Exp 4: moderated mediation, work-frame kills the effect. Hermsen et al. 2017, JMIR mHealth: 73.9% day 100, 16.0% day 320, n=711 objective logs. Endeavour line only as "consultancy survey" if mentioned at all.*

---

## [5:40 – 7:20] ACT THREE — WHAT 76 REAL DASHBOARDS DO · ~260w

**ON SCREEN:** Grid of home screens. Counters animate.

> We tallied 76 home screens and dashboards from top apps — 64 consumer, 12 SaaS. Three findings.
>
> **One: they're two different species.** A third of consumer home screens lead with a **single hero number** — one ring, one balance, one streak. SaaS dashboards? **Zero** of twelve. They open with a wall — and 42 percent of them show *nine or more numbers* before you scroll, while 42 percent of consumer screens show one to three. Consumer design compresses; SaaS accumulates.
>
> **Two: consumer apps frame data as progress; SaaS frames it as observation.** Goal rings, bars, "3 of 5" — on more than half of consumer screens, and on exactly **none** of the SaaS dashboards. Same with the timeframe: consumer opens on *today*; SaaS opens on the last 30 days. One is built to drive an action; the other is built to watch.
>
> **Three: top apps already obey the feedback research, whether they know it or not.** Red, shame-framed numbers? **Seven percent** of screens — positive framing outnumbers negative three to one. The "Hey Kris!" greeting the tutorials love? **Three percent** — that pixel row goes to data now. And the category that solved "what do I do next" best is learning apps: the lesson path makes the next action *structural* — the map is the CTA. Meanwhile SaaS dashboards surfaced a core-task next action exactly **zero** times.
>
> The market converged on task-focused feedback. The stragglers are the ones still shipping scoreboards.

> ⚠️ *Audit n=76 (curated top-app sample — phrase as "top apps we sampled"): hero 22/64 vs 0/12; 9+ numbers 5/12 SaaS vs 1–3 numbers 27/64 iOS; goal UI ~56% vs 0; today 57% vs 17%; red 5/76=7%; greeting 2/76=3%; SaaS core-task CTA 0/12.*

---

## [7:20 – 10:00] ACT FOUR — THE TEARDOWN · ~400w

**ON SCREEN:** Interactive demo. Fictional budgeting app "Penny." Before/after home screens.

> Let's rebuild one. Fictional budgeting app — Penny. The before screen is the overload pattern we found on real money apps: eleven numbers, three competing mental models, and a scoreboard's worth of self-focused feedback.

**HIGHLIGHT: the wall of numbers → one anchor**

> **The anchor.** Before: cash in, cash out, balance, budget left, daily average, four category totals — *eleven numbers*, no hierarchy. Every number is shouting, so none of them is heard. After: one hero — **"Safe to spend today: $42."** Everything else drops below it. That's the consumer compression pattern — one number you *feel* — and notice it isn't a grade. It's an instruction disguised as a status.

**HIGHLIGHT: the delta chip**

> **Velocity.** Before: this month's spending, alone, context-free. After: *"$118 less than last week"* — the change, not just the level. That's the top row of the moderator table — d = .55, double the effect of raw feedback. If you add only one widget from this video, add the delta.

**HIGHLIGHT: the red overspend banner**

> **The shame state.** Before: a red banner — "OVER BUDGET: –$63 😬." That's discouraging feedback about the *worker* — the row of the table that goes negative. And color isn't neutral chrome, by the way: experiments in finance found showing losses in red measurably shifts people's risk decisions — an effect that vanishes in colorblind participants. Color is a behavioral input. After: the same fact, task-framed and neutral: **"Dining is $63 over. Move $50 from Fun budget?"** — one tap. The information survived; the verdict about you didn't. Feedback with the correct next step attached: d = .43.

**HIGHLIGHT: the confetti toast**

> **The praise.** Before: "🎉 AMAZING! You logged expenses 3 days in a row!!" — praise, d = .09, the near-zero row. After: informational confirmation: *"3 days logged — your spending picture is now accurate."* Competence information beats cheerleading — it tells them what their effort *bought*.

**HIGHLIGHT: the live ticker**

> **The frequency.** Before: a real-time spending ticker, updating with every transaction. Feels alive; makes users chase noise — decision experiments found *less* frequent feedback beat constant feedback whenever the environment is noisy, because people over-react to the last data point. Daily spending is noisy. After: a **weekly rollup** with a smoothed trend. A boring cadence is an evidence-based feature.

**HIGHLIGHT: the goal ring**

> **And the goal.** Before: a savings goal at 4 percent, glaring at a brand-new user like a debt. After, two states: early on, the ring leads with what's *done* — "$120 saved" — because accumulated progress is what convinces an uncertain user this matters. Once they're committed, it flips to *"$380 to go."* Same number, framing that grows with the user.

**ON SCREEN:** Widget scorecard: work-focused 6 · worker-focused 0.

> Count it: before — eleven numbers, a grade, a shame state, and confetti. After — one anchor, one delta, one next action, one receipt, one honest cadence, one goal that grows up. All six point at the work. None point at the worker.

---

## [10:00 – 10:45] CLOSE · ~120w

**ON SCREEN:** The after screen. Hold.

> The pedometer people walked more and loved it less — and when the counter disappeared, so did they. That's the trap: a dashboard can inflate every metric it displays while draining the motivation that brings people back.
>
> So audit yours with one question per widget: **is this feedback about the work, or about the worker?** Deltas, next steps, receipts — work. Grades, ranks, shame banners, confetti — worker.
>
> More than a third of feedback interventions backfire. The fix isn't showing less. It's pointing every number at the task — because the task is the thing your user can actually do something about.

**[CTA / outro]**

---
---

# APPENDIX A — Optional expansion beats

### A1. The chart-perception rule that replicated after 26 years (+60 sec) — insert before Act Three
Cleveland & McGill, 1984: judging values encoded by **position** (bars, dots) beats length beats angle (pies) — tested with the percent-comparison task. In 2010, Heer & Bostock reran it on Mechanical Turk: same ranking. One of the few chart rules with a lab result AND an independent crowd replication — use position for anything users must compare; pies only for coarse part-to-whole. ⚠️ Only the tested slices — the full ten-step hierarchy is their *proposed* ranking; angle vs length was never separated.

### A2. Tufte vs the evidence (+45 sec) — contrarian aside for Act Three
The data-ink ratio has zero experiments behind it. What exists points the other way: Bateman's "Useful Junk?" found embellished charts equal for reading accuracy and *better* recalled 2–3 weeks later (⚠️ n=10 per recall condition — say so), and Borkin's 2,070-visualization study found minimalist charts the *least* memorable. Verdict: minimalism is an aesthetic, not a law. Don't overcorrect into clutter — but stop citing Tufte as science.

### A3. Streaks: the honest numbers + the mercy rule (+50 sec) — insert in Act Three
Our audit: streaks are a *category* strategy — 0 of 13 finance screens, 0 of 13 health screens, but half of learning apps; when present, hero-vs-badge splits 50/50. The academic legitimation is Silverman & Barasch 2023: highlighting an *intact* streak boosts engagement holding behavior constant, breaks hurt most when self-attributed — and **letting users repair a broken streak attenuates the damage**. That's the theory-backed case for streak freezes. Pair with Duolingo's own published A/Bs: streak visibility moved DAU by **1–3 percent** — the most famous retention mechanic in software is a single-digit effect, and anyone quoting "gamification +48%" is quoting a number with no stable source.

---

# APPENDIX B — Title, thumbnail, chapters

**Titles**
1. `Your Dashboard Is a Feedback Intervention (38% of Those Backfire)`
2. `The Pedometer That Ruined Walking — and What It Means for Your App`
3. `Confetti Is Worth d = 0.09 (Dashboard Psychology, With Receipts)`
4. `I Checked 76 App Dashboards Against 100 Years of Feedback Research`

**Thumbnail:** A confetti burst stamped "d = .09" vs a delta chip "↑12% vs last week" stamped "d = .55". Or the pedometer with a cracked heart.

**Chapters**
```
0:00  The pedometer that ruined walking
0:55  What we're covering
1:30  38% of feedback backfires (the 607-effect review)
2:40  The moderator table is a design spec
4:10  Measurement turns play into work
5:40  What 76 real dashboards do
7:20  Teardown: rebuilding Penny's home screen
10:00 The rule: work, not worker
```

---

# APPENDIX C — Ranked cut list

Baseline ~10:45 (≈1,625 spoken words @150wpm).

| # | Cut | Saves | Cost |
|---|---|---|---|
| 1 | The Butler classroom study (Act One) | 0:25 | Loses the most visceral illustration; the table still carries the argument. |
| 2 | The color/risk aside inside the shame-state beat | 0:15 | Loses the "color is a behavioral input" point; keep if the audience skews fintech. |
| 3 | Hermsen wearables numbers (Act Two) | 0:20 | Loses the longitudinal reality-check; Etkin still lands the mechanism. |
| 4 | Audit finding on greetings + learning-path CTA | 0:20 | Trims color from Act Three; keep the species-divide and red stats. |
| 5 | The goal-ring framing flip (Act Four) | 0:25 | Loses the most original design rule in the teardown — cut last. |

**Below ~9:00 you're cutting teaching.** For 8:30: cuts 1–4 and compress the roadmap.

---

# APPENDIX D — Production notes

**Every number, with its source**

| Claim | Source | Tier |
|---|---|---|
| Pedometer: more steps, less enjoyment (opt-in unprotected); counter cut continuation 48.5%→27.3% | Etkin 2016, *JCR* 42(6), Exp 2/3/6 | SOLID |
| 131 papers / 607 effects / 12,652 participants; d=0.41; >38% negative (33% excl. outlier) | Kluger & DeNisi 1996, *Psych Bulletin* 119(2) | SOLID — NEVER say "23,000 participants" |
| Moderators: velocity .55, solution .43, praise .09, esteem-threat .08, discouragement −.14 | K&D Table 2 | SOLID |
| Grade wipes out comment benefit | Butler 1988, *BJEP* 58 (~200 students) | SOLID as illustration |
| Work-frame kills quantification damage | Etkin Exp 4 (n=310) | SOLID |
| 711 Fitbits, device logs: 73.9% day 100, 16.0% day 320 | Hermsen et al. 2017, *JMIR mHealth* 5(10) | SOLID |
| Red losses shift risk behavior; absent in colorblind; muted in China | Bazley, Cronqvist & Mormann, *Mgmt Science* 2021 | SOLID direction — hedge exact % |
| Less frequent feedback beats constant in noisy environments | Lurie & Swaminathan 2009, *OBHDP* 108 | SOLID |
| To-date framing for uncertain commitment; to-go for committed | Koo & Fishbach 2008, *JPSP* 94(1) | SOLID — note: different paper from the small-area rule (2012) in the progress-bars video |
| Audit: hero 34% vs 0; 9+ numbers 42% of SaaS; goal UI 56% vs 0; today 57% vs 17%; red 7%; greeting 3%; SaaS core CTA 0/12; streaks 0 finance/health, ~half learning | Our Mobbin audit, n=76, Aug 2026 | Original data — say "top apps we sampled" |
| Streak-repair attenuation | Silverman & Barasch 2023, *JCR* 49(6) | SOLID |
| Duolingo streak A/Bs: 1–3% | Company-published (see progress-bars brief) | CONTESTED — "Duolingo reports" |

**Delivery notes**
- The moderator-table-as-design-spec mapping (Act One) is the centerpiece and the channel's novel contribution — build each row with a real widget mockup beside the d-value.
- "The measurement worked. That's the problem." — beat of silence after.
- Do not moralize the before screen; the overload pattern comes from real, successful apps. The tone is "the defaults are wrong, not the builders."
- Never say: "23,000 participants," "gamification +48%," "a study found a third of wearables abandoned in 6 months" (consultancy survey), "users look at dashboards for X seconds" (no source), "pie charts are objectively evil," "Tufte's research shows" (doctrine, not data), "Zeigarnik" (dead — see progress-bars brief). Full blacklist in the brief.

**Demo spec** — `demo-spec.md` in this folder; teardown beats match Act Four highlights.
