Home / Dashboards / Video script
Download raw .md

Your Dashboard Is a Feedback Intervention (And 38% of Those Backfire)

The science of progress displays that retain users instead of burning them out

Format: target 8–12 min | Audience: builders designing home screens, dashboards, stats surfaces Core thesis (say it 3 times): A dashboard isn't a report. It's a feedback intervention — and feedback about the work works, while feedback about the worker backfires.

Every number fact-checked against research/dashboards-research.md; original pattern data from mobbin/dashboards-mobbin-audit.md (n=76 screens, top apps, Aug 2026). Cross-reference guard: goal-gradient/endowed-progress/small-area beats belong to the progress-bars video — not rehashed here. ️ margin notes not spoken.


[0:00 – 0:55] COLD OPEN · ~130w

ON SCREEN: A person walking. A step counter ticking up.

Researchers gave people pedometers for a day. Simple field experiment. The people wearing a counter walked more — measurably more steps than the control group.

A win, right? Keep watching. The same people enjoyed walking less. Their end-of-day happiness was lower. And here's the detail that should worry you: participants who volunteered to wear the pedometer were hurt just the same.

Then the researchers ran the kicker study with reading. While a page counter was visible, people read more. Then the counter was removed, and everyone got a free choice: keep going or stop.

48.5 percent of the never-measured group kept reading. Of the measured group: 27.3.

ON SCREEN: "The measurement worked. That's the problem."

The counter boosted the metric and nearly halved the desire to continue. If you build dashboards, that sentence should keep you up at night.

Etkin 2016, JCR — Exp 2 (n=95, field): log-steps 7.97 vs 7.01, p=.003; enjoyment 4.82 vs 5.33, p=.030. Exp 6 (n=66): 27.3% vs 48.5%, p=.041. All figures verified against the paper.


[0:55 – 1:30] ROADMAP · ~85w

Four parts.

One — the biggest review of feedback ever conducted, and its most uncomfortable number. Your dashboard is a feedback intervention whether you meant it that way or not.

Two — a moderator table from that review that reads like a dashboard design spec. Which widgets carry six times the effect of others.

Three — what we found tallying 76 real home screens and dashboards from top apps — including the one thing consumer apps and SaaS dashboards do completely differently.

Four — the teardown: rebuilding an overloaded money dashboard, widget by widget.


[1:30 – 4:10] ACT ONE — FEEDBACK ABOUT THE WORK vs THE WORKER · ~400w

ON SCREEN: "131 papers · 607 effect sizes · 12,652 participants"

In 1996, Kluger and DeNisi published the definitive meta-analysis of feedback interventions: 131 papers, 607 effect sizes, twelve and a half thousand participants — nearly a century of "let's show people how they're doing" research.

Average effect: positive, moderate — d of 0.41. Feedback helps, on average.

But averages hide the story. More than 38 percent of the effects were negative. In over a third of the experiments, showing people feedback about their performance made them perform worse. And it's not one weird lab — remove the biggest outlier researcher entirely and a third of the remaining effects are still negative.

ON SCREEN: Two arrows from a number: → THE WORK / → THE WORKER

Their theory of when it backfires is beautifully simple. Feedback pulls attention to one of two places: the task — what happened, what to adjust — or the self — what does this say about me? The closer attention moves to the self, the worse feedback performs.

Feedback about the work works. Feedback about the worker backfires.

ON SCREEN: The moderator table, rows animating in with widget mockups beside each.

And the moderator table reads like a design spec, because every row is a dashboard widget.

Velocity feedback — how you're doing versus last time — carries the biggest effect in the table: d = .55. That's your delta chip: "up 12 percent from last week."

Feedback that includes the next action — d = .43. Not "your number is low," but "here's the thing to do about it."

Now the bottom of the table. Praise: d = .09. Statistically, the confetti animation and the "Great job!!" toast are doing almost nothing. Feedback that threatens self-esteem — comparisons, ranks, grades: .08. And discouraging feedback: negative — shame-framed metrics actively hurt.

There's a classroom study that makes it visceral. Schoolkids got returned work three ways: comments only, grades only, or grades plus comments. Comments alone improved performance. And adding a grade next to the comment wiped out the comment's benefit — the ego-number captured all the attention, and the useful information got ignored. That's what your big red score does to the helpful diagnostics sitting right under it.

K&D: >38% negative (say "more than a third — 38%"); do NOT say 23,000 participants (that's observations). Moderators: velocity .55 vs .28; solution .43 vs .25; praise .09 vs .34; self-esteem threat .08 vs .47; discouragement −.14. Butler 1988 (~200 students) — present as illustration, single pre-registration-era study.


[4:10 – 5:40] ACT TWO — THE HIDDEN COST OF COUNTING · ~230w

ON SCREEN: Coloring books, pedometers, page counters.

Back to that pedometer study, because the mechanism matters. Across six experiments — coloring, walking, reading — the pattern held: measurement increases output and drains enjoyment. Why? Measured, an activity starts to feel like work. Attention shifts from the experience to the number. Intrinsic motivation — the reason they'd come back tomorrow — leaks out.

But experiment four found the boundary, and it's the design rule: when reading was framed as work — as learning — the damage disappeared. Quantifying things people already treat as work is safe. Quantifying what they do for joy converts play into a job.

A sales pipeline, invoices, training plans: measure freely. Leisure reading, hobby learning, meditation minutes? Every counter you add is charging interest against the reason they showed up.

And the long-run version shows up in hardware: forget the famous "a third of wearables get abandoned" line — that's a consultancy survey with an undisclosed sample. The honest number is from researchers who gave 711 people Fitbits and read the actual device logs for 320 days: about three-quarters still tracking at day 100... 16 percent at day 320.

Metrics-first products don't just plateau. They actively spend down the intrinsic motivation they were built to display.

Etkin Exp 4: moderated mediation, work-frame kills the effect. Hermsen et al. 2017, JMIR mHealth: 73.9% day 100, 16.0% day 320, n=711 objective logs. Endeavour line only as "consultancy survey" if mentioned at all.


[5:40 – 7:20] ACT THREE — WHAT 76 REAL DASHBOARDS DO · ~260w

ON SCREEN: Grid of home screens. Counters animate.

We tallied 76 home screens and dashboards from top apps — 64 consumer, 12 SaaS. Three findings.

One: they're two different species. A third of consumer home screens lead with a single hero number — one ring, one balance, one streak. SaaS dashboards? Zero of twelve. They open with a wall — and 42 percent of them show nine or more numbers before you scroll, while 42 percent of consumer screens show one to three. Consumer design compresses; SaaS accumulates.

Two: consumer apps frame data as progress; SaaS frames it as observation. Goal rings, bars, "3 of 5" — on more than half of consumer screens, and on exactly none of the SaaS dashboards. Same with the timeframe: consumer opens on today; SaaS opens on the last 30 days. One is built to drive an action; the other is built to watch.

Three: top apps already obey the feedback research, whether they know it or not. Red, shame-framed numbers? Seven percent of screens — positive framing outnumbers negative three to one. The "Hey Kris!" greeting the tutorials love? Three percent — that pixel row goes to data now. And the category that solved "what do I do next" best is learning apps: the lesson path makes the next action structural — the map is the CTA. Meanwhile SaaS dashboards surfaced a core-task next action exactly zero times.

The market converged on task-focused feedback. The stragglers are the ones still shipping scoreboards.

Audit n=76 (curated top-app sample — phrase as "top apps we sampled"): hero 22/64 vs 0/12; 9+ numbers 5/12 SaaS vs 1–3 numbers 27/64 iOS; goal UI ~56% vs 0; today 57% vs 17%; red 5/76=7%; greeting 2/76=3%; SaaS core-task CTA 0/12.


[7:20 – 10:00] ACT FOUR — THE TEARDOWN · ~400w

ON SCREEN: Interactive demo. Fictional budgeting app "Penny." Before/after home screens.

Let's rebuild one. Fictional budgeting app — Penny. The before screen is the overload pattern we found on real money apps: eleven numbers, three competing mental models, and a scoreboard's worth of self-focused feedback.

HIGHLIGHT: the wall of numbers → one anchor

The anchor. Before: cash in, cash out, balance, budget left, daily average, four category totals — eleven numbers, no hierarchy. Every number is shouting, so none of them is heard. After: one hero — "Safe to spend today: $42." Everything else drops below it. That's the consumer compression pattern — one number you feel — and notice it isn't a grade. It's an instruction disguised as a status.

HIGHLIGHT: the delta chip

Velocity. Before: this month's spending, alone, context-free. After: "$118 less than last week" — the change, not just the level. That's the top row of the moderator table — d = .55, double the effect of raw feedback. If you add only one widget from this video, add the delta.

HIGHLIGHT: the red overspend banner

The shame state. Before: a red banner — "OVER BUDGET: –$63 ." That's discouraging feedback about the worker — the row of the table that goes negative. And color isn't neutral chrome, by the way: experiments in finance found showing losses in red measurably shifts people's risk decisions — an effect that vanishes in colorblind participants. Color is a behavioral input. After: the same fact, task-framed and neutral: "Dining is $63 over. Move $50 from Fun budget?" — one tap. The information survived; the verdict about you didn't. Feedback with the correct next step attached: d = .43.

HIGHLIGHT: the confetti toast

The praise. Before: " AMAZING! You logged expenses 3 days in a row!!" — praise, d = .09, the near-zero row. After: informational confirmation: "3 days logged — your spending picture is now accurate." Competence information beats cheerleading — it tells them what their effort bought.

HIGHLIGHT: the live ticker

The frequency. Before: a real-time spending ticker, updating with every transaction. Feels alive; makes users chase noise — decision experiments found less frequent feedback beat constant feedback whenever the environment is noisy, because people over-react to the last data point. Daily spending is noisy. After: a weekly rollup with a smoothed trend. A boring cadence is an evidence-based feature.

HIGHLIGHT: the goal ring

And the goal. Before: a savings goal at 4 percent, glaring at a brand-new user like a debt. After, two states: early on, the ring leads with what's done — "$120 saved" — because accumulated progress is what convinces an uncertain user this matters. Once they're committed, it flips to "$380 to go." Same number, framing that grows with the user.

ON SCREEN: Widget scorecard: work-focused 6 · worker-focused 0.

Count it: before — eleven numbers, a grade, a shame state, and confetti. After — one anchor, one delta, one next action, one receipt, one honest cadence, one goal that grows up. All six point at the work. None point at the worker.


[10:00 – 10:45] CLOSE · ~120w

ON SCREEN: The after screen. Hold.

The pedometer people walked more and loved it less — and when the counter disappeared, so did they. That's the trap: a dashboard can inflate every metric it displays while draining the motivation that brings people back.

So audit yours with one question per widget: is this feedback about the work, or about the worker? Deltas, next steps, receipts — work. Grades, ranks, shame banners, confetti — worker.

More than a third of feedback interventions backfire. The fix isn't showing less. It's pointing every number at the task — because the task is the thing your user can actually do something about.

[CTA / outro]

---

APPENDIX A — Optional expansion beats

A1. The chart-perception rule that replicated after 26 years (+60 sec) — insert before Act Three

Cleveland & McGill, 1984: judging values encoded by position (bars, dots) beats length beats angle (pies) — tested with the percent-comparison task. In 2010, Heer & Bostock reran it on Mechanical Turk: same ranking. One of the few chart rules with a lab result AND an independent crowd replication — use position for anything users must compare; pies only for coarse part-to-whole. ️ Only the tested slices — the full ten-step hierarchy is their proposed ranking; angle vs length was never separated.

A2. Tufte vs the evidence (+45 sec) — contrarian aside for Act Three

The data-ink ratio has zero experiments behind it. What exists points the other way: Bateman's "Useful Junk?" found embellished charts equal for reading accuracy and better recalled 2–3 weeks later (️ n=10 per recall condition — say so), and Borkin's 2,070-visualization study found minimalist charts the least memorable. Verdict: minimalism is an aesthetic, not a law. Don't overcorrect into clutter — but stop citing Tufte as science.

A3. Streaks: the honest numbers + the mercy rule (+50 sec) — insert in Act Three

Our audit: streaks are a category strategy — 0 of 13 finance screens, 0 of 13 health screens, but half of learning apps; when present, hero-vs-badge splits 50/50. The academic legitimation is Silverman & Barasch 2023: highlighting an intact streak boosts engagement holding behavior constant, breaks hurt most when self-attributed — and letting users repair a broken streak attenuates the damage. That's the theory-backed case for streak freezes. Pair with Duolingo's own published A/Bs: streak visibility moved DAU by 1–3 percent — the most famous retention mechanic in software is a single-digit effect, and anyone quoting "gamification +48%" is quoting a number with no stable source.


APPENDIX B — Title, thumbnail, chapters

Titles 1. Your Dashboard Is a Feedback Intervention (38% of Those Backfire) 2. The Pedometer That Ruined Walking — and What It Means for Your App 3. Confetti Is Worth d = 0.09 (Dashboard Psychology, With Receipts) 4. I Checked 76 App Dashboards Against 100 Years of Feedback Research

Thumbnail: A confetti burst stamped "d = .09" vs a delta chip "↑12% vs last week" stamped "d = .55". Or the pedometer with a cracked heart.

Chapters

0:00  The pedometer that ruined walking
0:55  What we're covering
1:30  38% of feedback backfires (the 607-effect review)
2:40  The moderator table is a design spec
4:10  Measurement turns play into work
5:40  What 76 real dashboards do
7:20  Teardown: rebuilding Penny's home screen
10:00 The rule: work, not worker

APPENDIX C — Ranked cut list

Baseline ~10:45 (≈1,625 spoken words @150wpm).

# Cut Saves Cost
1 The Butler classroom study (Act One) 0:25 Loses the most visceral illustration; the table still carries the argument.
2 The color/risk aside inside the shame-state beat 0:15 Loses the "color is a behavioral input" point; keep if the audience skews fintech.
3 Hermsen wearables numbers (Act Two) 0:20 Loses the longitudinal reality-check; Etkin still lands the mechanism.
4 Audit finding on greetings + learning-path CTA 0:20 Trims color from Act Three; keep the species-divide and red stats.
5 The goal-ring framing flip (Act Four) 0:25 Loses the most original design rule in the teardown — cut last.

Below ~9:00 you're cutting teaching. For 8:30: cuts 1–4 and compress the roadmap.


APPENDIX D — Production notes

Every number, with its source

Claim Source Tier
Pedometer: more steps, less enjoyment (opt-in unprotected); counter cut continuation 48.5%→27.3% Etkin 2016, JCR 42(6), Exp 2/3/6 SOLID
131 papers / 607 effects / 12,652 participants; d=0.41; >38% negative (33% excl. outlier) Kluger & DeNisi 1996, Psych Bulletin 119(2) SOLID — NEVER say "23,000 participants"
Moderators: velocity .55, solution .43, praise .09, esteem-threat .08, discouragement −.14 K&D Table 2 SOLID
Grade wipes out comment benefit Butler 1988, BJEP 58 (~200 students) SOLID as illustration
Work-frame kills quantification damage Etkin Exp 4 (n=310) SOLID
711 Fitbits, device logs: 73.9% day 100, 16.0% day 320 Hermsen et al. 2017, JMIR mHealth 5(10) SOLID
Red losses shift risk behavior; absent in colorblind; muted in China Bazley, Cronqvist & Mormann, Mgmt Science 2021 SOLID direction — hedge exact %
Less frequent feedback beats constant in noisy environments Lurie & Swaminathan 2009, OBHDP 108 SOLID
To-date framing for uncertain commitment; to-go for committed Koo & Fishbach 2008, JPSP 94(1) SOLID — note: different paper from the small-area rule (2012) in the progress-bars video
Audit: hero 34% vs 0; 9+ numbers 42% of SaaS; goal UI 56% vs 0; today 57% vs 17%; red 7%; greeting 3%; SaaS core CTA 0/12; streaks 0 finance/health, ~half learning Our Mobbin audit, n=76, Aug 2026 Original data — say "top apps we sampled"
Streak-repair attenuation Silverman & Barasch 2023, JCR 49(6) SOLID
Duolingo streak A/Bs: 1–3% Company-published (see progress-bars brief) CONTESTED — "Duolingo reports"

Delivery notes - The moderator-table-as-design-spec mapping (Act One) is the centerpiece and the channel's novel contribution — build each row with a real widget mockup beside the d-value. - "The measurement worked. That's the problem." — beat of silence after. - Do not moralize the before screen; the overload pattern comes from real, successful apps. The tone is "the defaults are wrong, not the builders." - Never say: "23,000 participants," "gamification +48%," "a study found a third of wearables abandoned in 6 months" (consultancy survey), "users look at dashboards for X seconds" (no source), "pie charts are objectively evil," "Tufte's research shows" (doctrine, not data), "Zeigarnik" (dead — see progress-bars brief). Full blacklist in the brief.

Demo specdemo-spec.md in this folder; teardown beats match Act Four highlights.