Product evidence
Reading kept getting better up to 18pt in the only large eye-tracking experiment on the question — and most apps still ship 14px. Size is the cheapest UX win you have.
The research, distilled
In a 104-person experiment with an eye-tracker, objective readability and comprehension kept improving as text size grew, all the way to 18pt — the largest size tested. Not preference: measured reading. The platforms quietly agree; iOS ships body text at 17pt and Material at 16sp, while most product screens still set 14px because it "fits more in."
Rello, Pielot & Marcos (2016), CHI '16 — eye-tracking, n=104Positive display polarity (dark text on a light ground) produced better acuity and proofreading performance for both younger and older adults, and the follow-up work found the mechanism: bright backgrounds contract the pupil, which sharpens the image on the retina. Dark mode is a legitimate preference and a battery saver on OLED at high brightness — but for long reading, light wins on the measurements.
Piepenbrock, Mayr, Mund & Buchner (2013), Ergonomics; mechanism follow-up 2014WCAG's threshold was derived from 1980s vision models and typical-vision-loss arithmetic — not from experiments on modern screens. It is still the right floor to enforce: it is the standard, it is enforceable, and it catches real failures. But know its weak spot: the formula is least reliable for light text on dark backgrounds, which is why dark-mode pairs that "pass" can still bloom and smear. Meet the number, then check dark pairs with your eyes.
W3C, Understanding SC 1.4.3 — provenance documented; graded as a standard, not an experimentWe went looking for the studies behind the pairing rulebook — serif-with-sans, matched x-heights, the two-typeface limit — and found none. Controlled legibility differences between mainstream typefaces are small to nonexistent. The pairing rules are aesthetic convention, which is fine, as long as nobody bills them as science. Spend the energy where the experiments are: size, contrast, and line length.
Our source-check of the pairing literature — judgment call, labeled as suchThe anchor study
Researchers sat 104 people in front of an eye-tracker and had them read at a range of text sizes. Instead of asking what looked nice, they measured objective readability and comprehension. Both kept improving as the text grew — all the way to 18pt, the largest size tested, with no plateau. The platforms quietly agree: iOS ships body text at 17pt and Material at 16sp. Most product screens still set 14px because it "fits more in."
Reading kept getting better all the way to the largest size tested.
no plateau · n=104, eye-tracked · measured readability and comprehension, up to 18ptwhat most product screens still set — under Material's 16sp and the 18pt the eye-tracking kept rewarding
Concepts in action
The same paragraph at the sizes and contrasts real apps ship. The ratio badge is the actual WCAG calculation for the pair you picked — watch what passes, what fails, and what passes-but-blooms in dark mode.
You ran 4 times this week for a total of 26 kilometers — up 12% on last week. Your longest run was Thursday's 9.4 km, and your average pace improved to 5:42 per kilometer. Keep this rhythm for two more weeks and you'll be ready to start the 10k plan.
How to read this: ratios are computed with the WCAG 2.x relative-luminance formula for the exact colors shown. The research result on size (better reading up to 18pt) is from Rello et al. 2016; the dark-mode caveat — pairs that pass the formula but bloom — is the documented weak spot of the formula itself, which is why the verdict line treats dark passes more skeptically.
Apply it to your app
This prompt turns the research above into an audit-then-fix session for Claude Code, Cursor, or similar. It maps your styles first, so nothing changes until it understands your system.
You are a senior product engineer applying published reading research to my app's typography. Grounding: in the largest eye-tracking experiment on text size (Rello, Pielot & Marcos 2016, n=104), objective readability and comprehension improved continuously up to 18pt with no plateau — so 16px is a floor for body text, not a target. Dark-on-light is measurably more legible for all age groups (Piepenbrock et al. 2013; the mechanism is pupil contraction). WCAG's 4.5:1 is a derived floor, not an experimental result, and it is least reliable for light-on-dark pairs — so dark-mode contrast needs a human check even when the math passes. BEFORE YOU CHANGE ANYTHING, map my system and wait: 1. List every text style in the codebase (font-size, weight, color, where used) — design tokens if they exist, else grep the styles. 2. Identify which styles carry body content people must actually read, versus labels and captions. 3. Note my dark-mode implementation, if any. THEN audit and show me the plan before coding: - Flag every body-content style under 16px, and propose 16-18px with adjusted line-height (1.5-1.65). - Compute the WCAG contrast ratio for every text/background pair in both themes. Flag everything under 4.5:1 — no exemptions for placeholder or "secondary" text people must read. - In dark mode, list pairs that pass 4.5:1 but have low luminance separation — those need a visual check on a real device, and likely a step down in brightness for large light text. - Do NOT touch typeface choices or add fonts — pairing rules have no experimental backing, and this session is about the variables that do: size, contrast, line length (45-75 characters). - Preserve intentional hierarchy: if bumping body text makes it collide with a heading size, scale the heading, don't shrink the body. THEN implement with my design tokens, smallest diff possible, behind a flag if the system supports it. Show me the diff before applying. Close with a table: style, before, after, ratio before/after, and which screens it touches.
From Build With Kris
Subscribe to catch the teardown when it drops — the eye-tracking result, the polarity evidence, and the type lab, live.