# Research Brief — Dashboards & Progress Displays That Drive Retention
### Backing document for the video script. Thesis: your dashboard is a feedback intervention, and feedback research predicts when it motivates vs. demotivates.

**Confidence key**
`SOLID` — primary source located, figure verified
`CONTESTED` — real source exists, but methodology or interpretation is disputed
`SHAKY` — widely repeated, source is weak, absent, or circular
`BLACKLISTED` — never repeat this claim

> Cross-references: the goal-gradient, endowed-progress, small-area, Cryder 85%-progress, and Duolingo streak A/B figures are already verified in `progress-bars-research-brief.md` — reuse those numbers verbatim, don't re-derive. Wang et al. 2016 (hotel goal failure, n=95,532) is also already used there. Amplitude retention benchmarks (B2B 3-month median 2.5%; 69% activation↔retention link) are verified in `decision-fatigue-research-brief.md`.

---

# ⚠️ PART 1 — CORRECTIONS

Five things in the working notes needed fixing before filming.

### 1. Kluger & DeNisi is NOT "~23,000 participants"

The meta-analysis is **607 effect sizes from 131 papers, based on 12,652 participants and 23,663 observations** (multiple observations per participant). Average sample per effect: **39 people**. If you say "23,000 participants" a commenter with the PDF will catch it in a minute. `SOLID` — verified against the primary PDF, p. 258.

### 2. "Over one-third decreased performance" — the exact figure is stronger

The paper's own wording: **"over 38% of the effects were negative."** And the robustness detail almost nobody quotes: 91 of the 607 effects came from one researcher (Mikulincer) who ran extreme-negative-feedback experiments (his studies alone average d = −0.39). **After excluding all of his studies, 33% of the remaining effects were still negative.** So the finding survives its biggest outlier. Say "more than a third — 38 percent" and you're exact and bulletproof. `SOLID`

### 3. The Endeavour Partners "one-third abandon" stat — here's what it actually was

The famous 2014 wearables-abandonment stat traces to **"Inside Wearables" Part 1 (Ledger & McCaffrey, Jan 2014)**: an **internet survey run in September 2013** of, verbatim, **"thousands of Americans" — the sample size is never disclosed in the report**. Claims: 1 in 10 US adults owned an activity tracker; **more than half of ever-owners no longer used it; one-third stopped within six months**. Part 2 (July 2014) discloses n > 1,700. It's a consulting firm's marketing white paper, not research. `CONTESTED` — cite as "a consultancy survey," never "a study."

**The better, honest replacement — objective device logs:** Hermsen et al. (2017), *JMIR mHealth*, gave **711 people a Fitbit Zip and read the actual data for 320 days**: **73.9% still tracking at day 100, only 16.0% at day 320; mean usage 129 days.** Top abandonment reason was mundane: dead batteries, broken or lost trackers (21.5%). `SOLID` — and it says the same thing as Endeavour with a disclosed method: attrition is the norm.

### 4. Cleveland & McGill's hierarchy is only partly tested — don't oversell it

The canonical ranking (position > length > direction > angle > area > volume > shading > saturation) is **a theoretical ordering; their experiments empirically tested only two slices of it**: position vs. length (55 subjects) and position vs. angle i.e. bar vs. pie (54 subjects). Both showed position clearly winning. But **angle vs. length was never separated — in both the original and the Heer & Bostock replication, angle did NOT perform worse than length.** State the tested part flatly ("position beats length beats angle/area"); attribute the full ten-step ladder as "their proposed ranking." `SOLID` with that framing.

### 5. "More frequent feedback = better" — the literature says the opposite, twice

- In Kluger & DeNisi's own moderator table, feedback frequency ran **backwards** (top-quartile frequency d = .32 vs. bottom-quartile .39) and the authors flag the variable as a possible artifact. At minimum: frequency is NOT a demonstrated positive moderator.
- Lurie & Swaminathan (2009) tested it directly (see §5): **less frequent feedback produced higher performance** in noisy environments. `SOLID`

---

# PART 2 — THE VERIFIED SPINE

## §1 — Feedback Intervention Theory: the video's spine

**Kluger & DeNisi (1996), *Psychological Bulletin* 119(2): 254–284.** `SOLID` — primary PDF verified.

Headline numbers:
| Figure | Value |
|---|---|
| Papers / effect sizes | 131 / 607 |
| Participants / observations | 12,652 / 23,663 |
| Mean effect of feedback on performance | **d = 0.41** (moderate positive) |
| Effects that were NEGATIVE | **>38%** (33% even after excluding the Mikulincer outlier studies) |

**The theory:** feedback shifts attention among three levels — **task learning, task motivation, and meta-task (self) processes**. Verbatim core claim: feedback effectiveness **decreases as attention moves up the hierarchy, closer to the self and away from the task.**

**The moderator table (Table 2, weighted mean d by level — the numbers to put on screen):**

| Moderator | With | Without | Dashboard translation |
|---|---|---|---|
| **Velocity feedback** (change vs. last time) | **.55** | .28 | Trends, deltas, "up 12% from last week" |
| **Correct solution provided** | **.43** | .25 | Feedback that contains the next action |
| Computerized (vs. person-delivered) | .41 | .23 | Machines don't trigger self-defense |
| Goal setting attached | .51 | .30 | Metric tied to an explicit target (p < .05 tier) |
| **Praise** | **.09** | .34 | Confetti, "great job!" toasts — near-zero effect |
| **Discouraging feedback** | **−.14** | .33 | Shame-framed metrics actively hurt |
| **Threat to self-esteem** (top vs. bottom quartile) | **.08** | .47 | Ranks, public comparison, grades |
| Task complexity (top vs. bottom quartile) | **.03** | .55 | Feedback barely helps on complex tasks |
| Physical tasks | −.11 | .36 | (fitness apps, note this one) |

**The FIT one-liner for camera:** *feedback about the work works; feedback about the worker backfires.* Velocity and next-step cues (task-level) carry the biggest positive deltas; praise, discouragement, and self-esteem threat (self-level) carry the smallest or negative ones.

**Classroom demonstration of the same law — Butler (1988), *BJEP* 58: 1–14.** ~200 fifth/sixth-graders, three feedback conditions across sessions: **comments only, grades only, grades + comments.** Interest and performance were highest with comments only; **adding a grade to the comment wiped out the comment's benefit** — the ego-involving number captured attention and the task-involving information was ignored. `SOLID` (single study, pre-registration era — present as illustration of FIT, not as a standalone law). Direct analogy: a dashboard "score" next to useful diagnostics behaves like the grade next to the comment.

## §2 — The hidden cost of measuring: Etkin

**Etkin (2016), "The Hidden Cost of Personal Quantification," *JCR* 42(6): 967–984.** Six experiments + one follow-up. `SOLID` — primary PDF verified; all stats below checked against the paper.

| Exp | Activity | n | Result |
|---|---|---|---|
| 1 | Coloring, 10 min | 105 | Measured group colored more shapes (8.68 vs. 7.02, p = .010) but enjoyed it less (4.63 vs. 5.13, p = .062); also colored **less creatively** (p = .014), fewer colors (p = .035) |
| 1b | Coloring, active control | 160 | Effect survives vs. an identical-action control (clicks changed a letter, not a counter): output p = .015, enjoyment p = .014 — it's not distraction |
| 2 | **Walking, field, all-day pedometers** | 95 (analyzed 50 vs 41) | Measured walked more (log-steps 7.97 vs. 7.01, p = .003) but **enjoyed walking less** (4.82 vs. 5.33, p = .030). **Participants opted IN to the pedometer** — self-selection didn't protect them |
| 3 | Walking, 3 conditions | 100 | Same pattern; mediated by walking **feeling like work** (p = .005); **subjective well-being dropped too** (4.59 vs. 5.18, p = .026). Optional-viewing condition identical — and **71.4% chose to look** |
| 4 | Reading, 2×3 design | 310 | Measurement → more pages read (16.14 vs. 13.44, p < .001), less enjoyment — **UNLESS reading was framed as work/learning, which killed the negative effect** (moderated mediation index = .40) |
| 5 | Reading + continued engagement | 236 (MTurk) | While counter visible: read more. **After the counter was removed: read LESS than never-measured controls** (3.75 vs. 4.20 pages, p = .034) |
| 6 | Reading + free choice | 66 | Given the choice to keep reading: **27.3% of measured participants continued vs. 48.5% of controls** (p = .041). 93.9% of the optional group chose to peek at the counter |

**Mechanism (tested, not vibes):** measurement draws attention to output → the activity feels like work → intrinsic motivation drops. Ruled out: distraction (1b, cognitive-load control in 5), doing-more fatigue (no output–enjoyment correlation in 2), stress/anxiety (6), difficulty (1, 6), need for achievement (1).

**The two beats that matter for the video:**
1. **Exp 6 is the retention finding**: the counter nearly halved voluntary continuation (48.5% → 27.3%).
2. **Exp 4 is the design rule**: quantify what users already treat as work; be very careful quantifying what they do for pleasure. A step counter on a commute app ≠ a page counter on a leisure-reading app.

**Boundary honesty:** lab/field tasks were minutes-to-a-day long; nobody has run the 12-month version. The wearables data (§Corrections #3) is the ecological rhyme, not proof of the same mechanism.

**Base-rate context — Pew, "Tracking for Health" (Fox & Duggan, Jan 2013), n = 3,014 US adults:** 69% track a health indicator, but **of those trackers ~49% track "in their heads," ~34% on paper, only ~21% use any technology**. The appetite for formal self-quantification is smaller than the industry assumes. `SOLID` (2013 — say the year).

**Qualitative abandonment color:** Lazar et al. (UbiComp 2015): 17 tech-company employees chose and bought smart devices; **~80% of the devices were abandoned within two months** — data "not useful," didn't fit self-concept, too much maintenance. Clawson et al. (UbiComp 2015) analyzed **~1,600 Craigslist listings** of secondhand trackers in one month — abandonment is common enough to have a resale market. Shih et al. (2015): college students given Fitbits — **65% stopped using them within two weeks**. All small/qualitative — use as texture, not statistics. `SOLID` as described.

## §3 — Goals: what's solid, what's contested

**Locke & Latham (2002), *American Psychologist* 57: 705–717** — the 35-year summary. Specific, difficult goals beat "do your best" with meta-analytic **d = 0.42–0.80**; goal-difficulty effects d = 0.52–0.82. Their own stated moderators: **commitment, feedback (goals need feedback to work — and feedback needs goals), ability, task complexity** (effects shrink on complex tasks — matching K&D's complexity moderator, d = .03 top quartile). `SOLID` as "the most replicated finding in organizational psychology," with moderators stated.

**The critique — Ordóñez, Schweitzer, Galinsky & Bazerman (2009), "Goals Gone Wild," *Academy of Management Perspectives* 23(1): 6–16.** Documented side effects of aggressive goal-setting: narrow focus, unethical behavior (Sears auto-repair quotas, Ford Pinto), distorted risk preferences, crowded-out intrinsic motivation. **Locke & Latham's published rebuttal in the same journal ("Has Goal Setting Gone Wild...?") disputes the scholarship but concedes the failure modes exist.** `SOLID` as a documented academic dispute — present both sides; the practical warning (dashboard targets get gamed) is uncontroversial. Their metaphor is usable: goals are *"prescription-strength medication"* — attribute it.

**Goal failure on a dashboard is expensive:** Wang et al. 2016 (n = 95,532; 80% missed the promotion goal; failures purchased less afterward vs. matched controls) — already verified in the progress-bars brief; cross-reference. Pairs with Cryder et al. 2013 (progress framing only fires near the goal): a dashboard that surfaces a goal the user will probably miss is a churn machine, not a motivator.

**Framing the same number two ways — Koo & Fishbach (2008), *JPSP* 94(1): 183–195:** identical progress can be framed **to-date ("you've done 40%") or to-go ("60% left")**. To-date motivates when **commitment is uncertain** (new users — accumulated progress signals "this matters to me"); to-go motivates when **commitment is established** (power users — remaining distance signals what to do). `SOLID` — this is the cleanest actionable rule in the goal literature for onboarding vs. engaged-state dashboards.

## §4 — Chart perception: what's actually proven

**Cleveland & McGill (1984), *JASA* 79(387): 531–554.** Position-length experiment: 55 subjects; position-angle experiment: 54 subjects; task = judge what percent the smaller value is of the larger; error metric log₂(|judged−true|+⅛). Results: **position judgments clearly more accurate than length; position clearly more accurate than angle (bar beats pie for comparisons)**. `SOLID` — with the Correction #4 framing: the full ten-rank hierarchy is proposed, not fully tested.

**The modern replication — Heer & Bostock (CHI 2010), "Crowdsourcing Graphical Perception."** Mechanical Turk, N = 24 assignments per chart across 70 judgment trials: **"the ranking of types by accuracy is consistent between the two experiments"** — position > length; their added conditions put angle (pie) and circular area worse than position, area worst. Also new: **extreme aspect ratios wreck rectangular-area (treemap) judgments**; a 26-year-old lab result replicated on the crowd. `SOLID`. Combined beat: one of the few UI design rules with a 1984 finding and a 2010 independent replication — use position (bars, dots, lines) for anything users must compare; save pies for coarse part-to-whole; treat treemap sizes as decoration.

**Chartjunk — the contrarian beat. Bateman et al. (CHI 2010), "Useful Junk?"** 20 participants (10 per recall condition), 14 charts: Nigel Holmes illustrated charts vs. plain versions of the same data. **Interpretation accuracy: no difference** (subject p = .412, categories p = .185, trend p = .818) — and value-message descriptions were actually *better* for Holmes charts (p = .003). **After 2–3 weeks: recall significantly better for embellished charts** on subject (p = .015), categories (p ≈ .000), trend (p = .042), and message (p = .020). `SOLID` as a finding, `CONTESTED` as a general rule — **n = 10 per recall condition**, single-fact editorial charts (not dashboards), and embellishment was confounded with color and imagery. Gelman's public critique exists; acknowledging it is the credibility move.

**Borkin et al. (IEEE InfoVis 2013), "What Makes a Visualization Memorable?"** 5,693 visualizations scraped → **2,070 single-panel visualizations, 261 MTurk participants**, 410 targets. Most memorable: pictograms/human-recognizable objects, more colors, low density. **Minimalist Tufte-style charts were the LEAST memorable.** Caveat you must say: **memorable ≠ comprehensible** — memorability was measured as image recognition; their 2016 follow-up shows titles and text drive message recall. `SOLID` with caveat.

**Tufte's data-ink ratio: label it as doctrine, not data.** *The Visual Display of Quantitative Information* (1983) presents the ratio as a principle with zero experiments. The empirical tests that exist (Bateman, Borkin; earlier Carswell reviews) fail to support strict minimalism. Frame: "Tufte is a great designer whose most famous rule has never won an experiment."

**Stephen Few — expert opinion, clearly labeled.** His definition is the best working one: a dashboard is a *"visual display of the most important information needed to achieve one or more objectives; consolidated and arranged on a single screen so the information can be monitored at a glance."* His principles (no gauges/dials — space-inefficient; single screen; overview first) are craft wisdom from *Information Dashboard Design* (2006/2013). `SOLID` as attribution, `SHAKY` as science — there are no controlled trials behind them.

**Red numbers — real evidence exists. Bazley, Cronqvist & Mormann, "In the Red: The Effects of Color on Investment Behavior," *Management Science* (2021).** Experiments: displaying potential losses **in red vs. black made people meaningfully more risk-averse** (~25% in secondary reporting — cite the direction from the paper, hedge the magnitude) and lowered return expectations after red price charts. Robustness: effect absent in **colorblind participants** and muted in **China**, where red means gains — it's learned association, not salience. `SOLID` for direction and robustness checks; `CONTESTED` for the exact 25%. This is the "peak signals" beat: color on a dashboard is not neutral chrome; it changes decisions.

## §5 — Streaks and feedback frequency

**Streaks are now real academic literature — Silverman & Barasch (2023), "On or Off Track: How (Broken) Streaks Affect Consumer Decisions," *JCR* 49(6): 1095–1117.** Seven studies: **highlighting an intact streak in a behavior log increases subsequent engagement vs. highlighting a broken one — holding actual past behavior constant.** The streak becomes a goal in itself. Breaks hurt most when self-attributed; **letting users "repair" a broken streak attenuates the damage.** `SOLID`. This is the academic legitimation of Duolingo's streak-freeze design (company-published mechanic, same logic).

**Duolingo's own numbers — reuse from the progress-bars brief, consistently:** making the streak visible tested at **+3% DAU / +1% D14**; emphasizing it +1% DAU / +3% D14 (company-published A/B). The "7-day streak → 3.6× course completion" figure is **correlational** and Duolingo never claimed causation. **The honest headline: the single most famous retention mechanic in consumer software moved metrics by single digits.** That's your antidote to every "48% engagement" claim.

**Feedback frequency — Lurie & Swaminathan (2009), *OBHDP* 108: 315–329.** Four newsvendor experiments; Exp 1: **76 students**, ordering decisions over 30 rounds, feedback every 1, 3, or 6 rounds. **Less frequent feedback → higher profits** (118.6 vs. 113.5 thousand francs, p < .01 for 6-round vs. every-round). The interaction is the design rule: **the harm appears only in high-variance (noisy) environments** (105.2 vs. 94.9, p < .05); with low variance, frequency didn't matter (all ≈132). Mechanism: frequent feedback → overweighting the most recent data point → demand-chasing. `SOLID`. Dashboard translation: real-time counters on noisy metrics make users chase noise; daily/weekly rollups and smoothing are evidence-backed, not cowardice.

**Gamification overall — the real effect sizes:** Sailer & Homner (2020) meta-analysis, *Educational Psychology Review* 32: 77–112: cognitive g = 0.49, motivational **g = 0.36**, behavioral **g = 0.25** — small, heterogeneous, and **the motivational/behavioral effects were unstable in the high-rigor subsplit**. Pairs with Hanus & Fox 2015 (gamified classroom → *lower* motivation over 16 weeks; already in progress-bars brief). `SOLID`.

---

# PART 3 — HOOK CANDIDATES

1. **"Psychology's biggest-ever review of feedback found that 38% of the time, giving people feedback made them WORSE."** (607 effects, d = .41 average, >38% negative; survives outlier removal at 33%.) Then: "your dashboard is a feedback intervention."
2. **The pedometer that ruined walking.** Etkin Exp 2–3: people walked more and enjoyed it less, felt less happy at day's end — and the ones who *chose* the pedometer were hurt just the same. Exp 6 kicker: the counter cut voluntary continuation from 48.5% to 27.3%. "The measurement worked. That's the problem."
3. **Praise is worth d = 0.09.** The confetti animation your PM loves sits in the least effective row of the moderator table; a plain "you improved vs. last week" (velocity, d = .55) is ~6× the effect size.
4. **"Duolingo published its streak experiments. The effects are 1 to 3 percent."** vs. the "gamification increases engagement 48%" folklore (vendor numbers that appear as 29%, 30%, 47%, and 48% depending on the listicle).
5. **The chart hierarchy that replicated after 26 years** — 1984 lab result, 2010 Mechanical Turk replication, same ranking. Rare good news: one dashboard rule you can trust.
6. **Red numbers make people risk-averse — except in China.** (Bazley et al.: gone in colorblind users, reversed-culture muted. Your color palette is a behavioral intervention.)

# PART 4 — NOVEL-ANGLE CANDIDATES

1. **Map K&D's moderator table onto dashboard widgets** (the video's likely centerpiece): velocity metrics = deltas/trends (.55); correct-solution cues = "here's what to do next" (.43); computerized delivery (.41); praise = celebration animations (.09); discouraging framing = shame states (−.14); self-esteem threat = leaderboards/percentile ranks (.08). Nobody in the dashboard-design content space has done this mapping with the actual numbers.
2. **Etkin's work-frame boundary as a product rule:** quantification is safe for activities users already frame as work (sales pipeline, invoices, workouts-as-training) and dangerous for intrinsically enjoyable ones (reading, hobby learning, casual play). Predicts *which* apps' dashboards backfire.
3. **The frequency argument for boring dashboards:** Lurie & Swaminathan + K&D's artifact-flagged frequency moderator → real-time displays of noisy metrics cause noise-chasing; weekly digests are the evidence-based default. Contrarian vs. every "live dashboard" product page.
4. **To-date for new users, to-go for committed users** (Koo & Fishbach): the same progress number should flip framing as commitment grows — a concrete, testable UI rule almost no product implements.
5. **Streak repair as theory-backed mercy:** Silverman & Barasch's attenuation result is the academic case for streak freezes and grace periods — loss-aversion mechanics need an escape valve or breaks convert to churn.
6. **The confidence gap (mirror of the progress-bars brief):** there is no controlled experiment on *dashboard design → retention* as such. Everything here is adjacent evidence (feedback, quantification, perception). Saying so on camera is the brand.

# PART 5 — THE BLACKLIST

Never say these on camera:

| Claim | Why |
|---|---|
| "Feedback always helps / people always want feedback" | K&D: >38% of 607 effects negative. |
| "Kluger & DeNisi studied 23,000 participants" | 12,652 participants; 23,663 is the observation count. |
| "Gamification increases engagement by 48%" | Vendor telemetry (Gigya-era); circulates as 29%, 30%, 47%, and 48% with no stable methodology. Sailer & Homner's real meta-analytic figures are g = 0.25–0.49 and unstable at high rigor. |
| "A study found one-third of wearables are abandoned in 6 months" | Consultancy internet survey, sample size undisclosed in the report. Say "a consultancy survey"; use Hermsen (n=711, device logs) for a real number. |
| "90% of BI dashboards go unused after six months" | "Estimated" with no traceable primary source. Real sourced figure: BI adoption stuck ~22–25% (Gartner/BARC) — different claim. |
| "Users only look at a dashboard for X seconds" | No primary source found in any form. |
| "People love seeing their data" | Etkin's six experiments cut against it; Pew 2013: only ~21% of health trackers used any technology. |
| "Cleveland & McGill proved the full position>length>angle>area>color hierarchy" | Only position-vs-length and position-vs-angle were tested; angle vs. length never separated (in the replication either). |
| "Pie charts are objectively evil" | Overclaim — angle underperforms position for comparisons; part-to-whole gist is fine. |
| "Tufte's research shows maximizing data-ink improves comprehension" | There is no such research; Bateman and Borkin point the other way. Tufte is theory/aesthetics. |
| "Embellished charts are proven better" (overcorrection) | Bateman n = 10 per recall condition, editorial charts, confounded embellishment. It questions minimalism; it doesn't crown decoration. |
| "Duolingo streaks cause 3.6× completion" | Correlational; Duolingo never claimed causation. Their causal A/Bs moved 1–3%. |
| "The Zeigarnik effect explains streaks/progress displays" | 2025 meta-analysis: pooled recall ratio 0.99 (null). Ovsiankina (resumption) is the survivor. (Already blacklisted in the progress-bars brief — stay consistent.) |
| "Red means danger is hardwired, so use red for losses" | Bazley et al. show it's a *learned* association (absent in colorblind users, muted in China) — and it changes risk behavior, which may not be what you want. |
| "Goals Gone Wild proved goal setting doesn't work" | It's a critique of over-prescription; Locke & Latham's d = .42–.80 core stands and the authors of the critique don't dispute it. |

# PART 6 — PRIMARY SOURCES

**Feedback**
- Kluger & DeNisi (1996), *Psych. Bulletin* 119(2) — https://mrbartonmaths.com/resourcesnew/8.%20Research/Marking%20and%20Feedback/The%20effects%20of%20feedback%20interventions.pdf (also doi:10.1037/0033-2909.119.2.254)
- Butler (1988), *BJEP* 58 — https://bpspsychub.onlinelibrary.wiley.com/doi/abs/10.1111/j.2044-8279.1988.tb00874.x
- Lurie & Swaminathan (2009), *OBHDP* 108 — https://marketing.business.uconn.edu/wp-content/uploads/sites/724/2014/08/is-timely-information-always-better.pdf

**Quantification & wearables**
- Etkin (2016), *JCR* 42(6) — https://marketing.wharton.upenn.edu/wp-content/uploads/2016/10/Etkin-Jordan-11-12-15-Hidden-Cost.pdf (published: https://academic.oup.com/jcr/article-abstract/42/6/967/2358309)
- Endeavour Partners, Inside Wearables Pt. 1 (2014) — https://medium.com/@endeavourprtnrs/inside-wearable-how-the-science-of-human-behavior-change-offers-the-secret-to-long-term-engagement-a15b3c7d4cf3
- Hermsen et al. (2017), *JMIR mHealth* 5(10):e164 — https://mhealth.jmir.org/2017/10/e164/
- Lazar et al. (UbiComp 2015) — https://dl.acm.org/doi/10.1145/2750858.2804288
- Clawson et al. (UbiComp 2015) — https://dl.acm.org/doi/10.1145/2750858.2807554
- Pew, Tracking for Health (2013) — https://www.pewresearch.org/internet/2013/01/28/tracking-for-health/

**Goals**
- Locke & Latham (2002), *Am. Psychologist* 57 — https://pubmed.ncbi.nlm.nih.gov/12237980/
- Ordóñez et al. (2009), Goals Gone Wild — https://www.hbs.edu/faculty/Pages/item.aspx?num=35109
- Locke & Latham (2009) rebuttal — https://journals.aom.org/doi/10.5465/amp.2009.37008000
- Koo & Fishbach (2008), *JPSP* 94(1) — https://pubmed.ncbi.nlm.nih.gov/18211171/

**Perception & chart design**
- Cleveland & McGill (1984), *JASA* 79(387) — http://euclid.psych.yorku.ca/www/psy6135/papers/ClevelandMcGill1984.pdf
- Heer & Bostock (CHI 2010) — https://idl.cs.washington.edu/files/2010-MTurk-CHI.pdf
- Bateman et al. (CHI 2010), Useful Junk? — https://sites.stat.columbia.edu/gelman/communication/Bateman2010.pdf
- Borkin et al. (2013), IEEE TVCG 19(12) — http://web.mit.edu/zoya/www/docs/InfoVis_borkin-128.pdf
- Bazley, Cronqvist & Mormann, In the Red, *Mgmt Science* — https://scholar.smu.edu/business_marketing_research/28/
- Few, *Information Dashboard Design* (2nd ed., 2013) — expert opinion; definition via https://www.perceptualedge.com/files/Dashboard_Design_Course.pdf

**Streaks & gamification**
- Silverman & Barasch (2023), *JCR* 49(6) — https://academic.oup.com/jcr/article-abstract/49/6/1095/6623414
- Sailer & Homner (2020), *Ed. Psych. Review* 32 — https://eric.ed.gov/?id=EJ1245270
- Duolingo streak A/Bs — see progress-bars brief (Econsultancy/Gilani + Duolingo blog links there)

**Cross-referenced briefs**
- `~/Downloads/progress-bars-research-brief.md` (goal gradient, endowed progress, small-area, Cryder 85%, Wang 2016, Duolingo, Zeigarnik-null)
- `~/Downloads/decision-fatigue-research-brief.md` (Amplitude retention benchmarks, defaults, Baymard)
