Software, mid-2026 · every number verified or cut
Writing the code became the easy part this year. Verifying it, distributing it, and pricing it became the hard part. This report is what the 2025–26 evidence actually says about all four — with the famous numbers we refused to print listed by name.
Free · from the Science-Based Software research libraryEvery claim was traced to its primary source where one exists. PRIMARY means we fetched and read the original study or dataset. SECONDARY means credible reporting we could not trace to a fetched primary — treat the exact number as softer than the direction.
Six widely-circulated 2026 "facts" failed our source-check entirely — including a viral outage story and a famous share-of-code stat. They get their own dark page at the end, so you can stop repeating them too.
The "19% slower" study defined the debate — and then its own authors took the number off the table. What's left is a more useful, less quotable picture.
In the July 2025 randomized trial, 16 experienced open-source maintainers predicted AI would make them 24% faster, measured 19% slower, and believed afterward they had been 20% faster — a forty-point gap between perception and reality.PRIMARY Then, in February 2026, METR published a follow-up: with newer tools, the original cohort measured –18% with a confidence interval crossing zero, new recruits –4% — and the team disclosed that 30–50% of participants were withholding exactly the tasks AI suits least, biasing the design. They now treat the 19% figure as historical.PRIMARY
Meanwhile the largest field experiment — 4,867 developers across Microsoft, Accenture, and a Fortune 100 firm — found +26% completed tasks, with the gains concentrated in less-experienced developers.PRIMARY DORA 2025 (~5,000 respondents) squares the circle: 90% use AI daily, most report gains, delivery instability rose with adoption, and 30% barely trust the output. Their verdict: AI is an amplifier of whatever system it lands in.PRIMARY
The clearest 2026 signal isn't in anyone's opinion — it's in the merge data. Generation is no longer the bottleneck; verification is.
The flagship datapoint cuts both ways: Bun's ~1M-line Zig-to-Rust port ran 11 days on ~50 continuously-running Claude Code workflows, at an estimated ~$165k of computePRIMARY — and drew the year's defining criticism when Zig's creator called the result "unreviewed slop." Million-line rewrites became a budget question; trusting them stayed an open one.
| The old default | What the 2026 evidence supports | |
|---|---|---|
| Typing code was the job; speed meant keystrokes. | → | Specifying is the job. Spec-first loops (constitution → spec → plan → tasks → verify) went mainstream; GitHub's Spec Kit passed 120k stars. |
| Review was a courtesy pass at the end. | → | Review is a production function. Agentic PRs wait 5.3× longer and merge at a third of the rate — staff it, SLA it, and keep AI PRs small. |
| "Compiles and passes tests" meant done. | → | Security scanning lives inside the agent loop. ~45% of generated code carries an OWASP flaw at >95% syntax correctness — the two are unrelated. |
| Refactoring happened when someone had time. | → | Duplication and churn are CI gates. Assistants default to paste-over-refactor; the debt is measurable at ecosystem scale. |
| Hiring valued lines-per-day. | → | Hiring values judgment-per-token: decomposition, specification, review, and system design. TypeScript became GitHub's #1 language partly because types pair better with agents. |
The distribution math that launched a decade of content-led startups broke this year — measurably, from three independent directions.
The developer-distribution version is starker: Stack Overflow's monthly questions fell to 2008-era levels — the Q&A discovery path is simply gone. And a new store-shaped surface opened: OpenAI's Apps SDK puts apps inside ChatGPT's 800M weekly users, built on MCP.PRIMARY
For the first time since 2008, web-checkout economics are a real design decision on mobile — in the US, the EU, and Japan, for different reasons.
| Front | Where it landed | |
|---|---|---|
| US · Apple | → | External purchase links are legal; the fee is being litigated. The 2025 injunction forced 0%-commission link-outs; the Ninth Circuit later allowed a fee capped at "genuinely necessary" costs — Apple has proposed up to 15%.PRIMARY |
| US · Google | → | Play allowed external links and alternative billing from late 2025 after Epic's injunction was upheld; the settlement was filed March 2026.PRIMARY |
| EU · DMA | → | €500M anti-steering fine; the Core Technology Fee became a 5% commission; alternative app stores are live — and Japan's new law brought Epic Games Store and AltStore there in January 2026.PRIMARY |
The subscription-app dataset (115,000+ apps, $16B+ tracked revenue) tells one story in two numbers: ~15,000 new subscription apps launch monthly — up from ~2,000 in 2022 — while median MRR growth sits at 5.3%. The top decile grew 306%. In vendor telemetry, with its limits: the gap is craft.
The strangest 2025–26 finding for anyone writing marketing copy: saying "AI" out loud measurably hurts.
In controlled experiments, identical products described as using "artificial intelligence" drew lower purchase intent in every tested case than the same products described as "high tech" — mediated by reduced emotional trust, strongest for high-risk purchases.PRIMARY The climate agrees: 50% of US adults are now more concerned than excited about AI (37% in 2021), and Gartner projects over 40% of agentic-AI projects canceled by end-2027, coining "agent washing" for the rest.PRIMARY
Two MIT results complete the picture. The NANDA report found 95% of enterprise GenAI pilots produce no P&L impact — the 5% that work are embedded in workflows with memory and feedback loops, not chat windows.PRIMARY And the "Your Brain on ChatGPT" EEG study (n=54, preprint — directional, not definitive) measured "cognitive debt": LLM users showed the weakest brain connectivity and carried the deficit into unassisted work.PRIMARY
Regulatory noise was constant; four things actually changed what shipping software requires.
| Rule | What it means in your product | |
|---|---|---|
| EU AI Act, Article 50 — in force Aug 2, 2026 | → | Chatbots must disclose they're AI; synthetic media must be labeled and machine-readably marked. High-risk obligations slipped to Dec 2027 — transparency did not. Fines reach €15M / 3% of turnover.PRIMARY |
| FTC click-to-cancel vacated (July 2025) | → | California became the US baseline: AB 2863 requires cancel-in-the-same-medium, affirmative consent with records, and renewal reminders. ROSCA enforcement never paused.PRIMARY |
| CPPA enforcement | → | Broken opt-outs now carry real fines: Honda $632,500; Tractor Supply a record $1.35M for a non-functioning Global Privacy Control signal. Regulators are testing behavior, not policy text.PRIMARY |
| European Accessibility Act (June 2025) | → | WCAG AA is market-entry table stakes in the EU. Enforcement so far is audits, notices, and court orders — a French court ordered Carrefour to make its site and app fully accessible in June 2026.SECONDARY |
Each of these circulated widely this year. We went looking for the primary source; here is what we found instead.
The story appears only in a single low-authority source chain — no primary confirmation, no major-outlet reporting. A perfect cautionary tale, which is exactly why it should have a source and doesn't.
FAILED CHECK · UNCONFIRMED STORYCirculates only in secondary aggregations. The verifiable executive claims are far more modest: ~30% at Microsoft (Nadella, April 2025) and ~25–30% at Google per Pichai's earlier statements.
FAILED CHECK · SECONDARY AGGREGATORS ONLYAppears only in vendor blogs citing Gartner — no locatable Gartner press release or report contains it. Generative-UI tooling is real; this measurement is not.
FAILED CHECK · UNLOCATABLE CITATIONTraces to 1990s–2000s catalog studies. The 2026 preregistered lab work finds 9-endings improve price image while signaling lower quality, and a preregistered 2022 study of 4,788 purchase decisions failed to reproduce left-digit effects. Test on your own product.
STALE EVIDENCE · MODERN TESTS ARE MIXEDVendor marketing data from billing platforms, not experiments. The pattern is cheap to test and directionally plausible — run your own holdout before quoting anyone's numbers.
VENDOR DATA · NO EXPERIMENTAlso its opposite: "the Copilot study proves +26% for everyone." Both flatten studies that disagree by design — and the slowdown study's authors have withdrawn its headline number themselves. The honest 2026 answer is: it depends on task, tenure, and codebase, and perception is not measurement.
OVERCLAIM IN BOTH DIRECTIONSWant this run against your actual codebase?The Science-Based Software App Audit is one paste into your coding agent: 38 evidence-graded checks, a report like this one, and a fix plan it implements diff by diff.
sciencebasedsoftware.com/audit · $79Sources: METR (2025 RCT + Feb 2026 update) · Cui et al., SSRN 4945566 · DORA 2025 · GitClear 2026 · Stack Overflow 2025 · LinearB 2026 · Veracode 2025 · SparkToro/Similarweb 2026 · Ahrefs · Pew Research · Cloudflare · RevenueCat State of Subscription Apps 2026 · WSU/Cicek et al. · Gartner · MIT NANDA · MIT Media Lab · EU AI Act Article 50 · California AB 2863 · CPPA enforcement actions · court reporting on Epic v. Apple/Google. Confidence labels mark whether we fetched the primary source or relied on credible secondary reporting.