Science-Based Software

The 2026 Best Practices Report

Software, mid-2026 · every number verified or cut

Writing the code became the easy part this year. Verifying it, distributing it, and pricing it became the hard part. This report is what the 2025–26 evidence actually says about all four — with the famous numbers we refused to print listed by name.

Free · from the Science-Based Software research library
The 2026 Report1
In this reportSeven questions · every answer sourced

Seven questions 2026 actually answered.

How we verified

Every claim was traced to its primary source where one exists. PRIMARY means we fetched and read the original study or dataset. SECONDARY means credible reporting we could not trace to a fetched primary — treat the exact number as softer than the direction.

What we refused to print

Six widely-circulated 2026 "facts" failed our source-check entirely — including a viral outage story and a famous share-of-code stat. They get their own dark page at the end, so you can stop repeating them too.

The 2026 Report2
Question 01Did AI actually make developers faster?

The most famous answer of 2025 withdrew itself.

The "19% slower" study defined the debate — and then its own authors took the number off the table. What's left is a more useful, less quotable picture.

In the July 2025 randomized trial, 16 experienced open-source maintainers predicted AI would make them 24% faster, measured 19% slower, and believed afterward they had been 20% faster — a forty-point gap between perception and reality.PRIMARY Then, in February 2026, METR published a follow-up: with newer tools, the original cohort measured –18% with a confidence interval crossing zero, new recruits –4% — and the team disclosed that 30–50% of participants were withholding exactly the tasks AI suits least, biasing the design. They now treat the 19% figure as historical.PRIMARY

Randomized trials · read with their own caveats

Meanwhile the largest field experiment — 4,867 developers across Microsoft, Accenture, and a Fortune 100 firm — found +26% completed tasks, with the gains concentrated in less-experienced developers.PRIMARY DORA 2025 (~5,000 respondents) squares the circle: 90% use AI daily, most report gains, delivery instability rose with adoption, and 30% barely trust the output. Their verdict: AI is an amplifier of whatever system it lands in.PRIMARY

So what do you do?
  • Measure AI productivity with telemetry, never self-report — perception ran ~40 points optimistic in the one place both were measured.
  • Route AI adoption by task and seniority: juniors on standard tasks gain most; experts in mature codebases can lose.
  • Treat every AI-productivity number as time-stamped. The tools change faster than the studies.
The 2026 Report3
Question 02Code got cheap — what got expensive?

Everything after the diff.

The clearest 2026 signal isn't in anyone's opinion — it's in the merge data. Generation is no longer the bottleneck; verification is.

+81%code duplication per million changed lines, 2023–2026, across 623M analyzed changesPRIMARY
13%→3.8%share of changed lines that are refactors — paste replaced reworkPRIMARY
–35%cross-file connectivity — AI writes islandsPRIMARY
5.3×longer first-review pickup for agentic-AI pull requests (8.1M PRs)SECONDARY
32.7%merge rate for agentic PRs vs 84.5% unassistedSECONDARY
~45%of AI-generated code failing security tests with an OWASP Top 10 flawSECONDARY

The flagship datapoint cuts both ways: Bun's ~1M-line Zig-to-Rust port ran 11 days on ~50 continuously-running Claude Code workflows, at an estimated ~$165k of computePRIMARY — and drew the year's defining criticism when Zig's creator called the result "unreviewed slop." Million-line rewrites became a budget question; trusting them stayed an open one.

The year's one-line summary Stack Overflow's 49,000-developer survey: 84% use AI, 46% actively distrust its accuracy, and the #1 complaint is "almost-right" answers that create debugging work. The market's bottleneck is verification UX, not generation quality.
The 2026 Report4
Question 02 · continuedThe practice shift, old → new

The 2026 workflow, in one table.

The old defaultWhat the 2026 evidence supports
Typing code was the job; speed meant keystrokes.→Specifying is the job. Spec-first loops (constitution → spec → plan → tasks → verify) went mainstream; GitHub's Spec Kit passed 120k stars.
Review was a courtesy pass at the end.→Review is a production function. Agentic PRs wait 5.3× longer and merge at a third of the rate — staff it, SLA it, and keep AI PRs small.
"Compiles and passes tests" meant done.→Security scanning lives inside the agent loop. ~45% of generated code carries an OWASP flaw at >95% syntax correctness — the two are unrelated.
Refactoring happened when someone had time.→Duplication and churn are CI gates. Assistants default to paste-over-refactor; the debt is measurable at ecosystem scale.
Hiring valued lines-per-day.→Hiring values judgment-per-token: decomposition, specification, review, and system design. TypeScript became GitHub's #1 language partly because types pair better with agents.
So what do you do?
  • Plan the verification harness before the agent fleet — the Bun criticism was about review, not capability.
  • Add duplication/churn gates and scheduled refactoring capacity to CI now, while the debt is young.
  • Put SAST and dependency scanning in the loop the agent runs, not just the pipeline humans watch.
The 2026 Report5
Question 03Where did the users go?

Search still answers. It just stopped sending.

The distribution math that launched a decade of content-led startups broke this year — measurably, from three independent directions.

68%of US Google searches end without a click (was 60% in 2024, ~45% in 2016)PRIMARY
–58%top-result CTR when an AI Overview is present (Ahrefs re-measurement, Dec 2025)PRIMARY
8 vs 15of 100 AI-Overview pages get any result click, vs pages without one (Pew, 68,879 queries)PRIMARY
1 in 100AI-Overview pages produce a click on the AI's own citation linksPRIMARY
904:1OpenAI's crawl-to-referral ratio (Anthropic ~10,300:1; Google search historically ~14:1)PRIMARY
0.32%of web traffic referred by ChatGPT at its all-time high — real, tiny, growingSECONDARY

The developer-distribution version is starker: Stack Overflow's monthly questions fell to 2008-era levels — the Q&A discovery path is simply gone. And a new store-shaped surface opened: OpenAI's Apps SDK puts apps inside ChatGPT's 800M weekly users, built on MCP.PRIMARY

So what do you do?
  • Treat organic search as a citation surface, not a traffic pipe — move acquisition weight to owned channels: email, community, product-led loops.
  • Optimize to be the cited source (entity clarity, quotable stats, schema) and expect the click to be rare even when cited.
  • Make your product legible to LLMs — clean docs, llms.txt, an MCP server — and set an explicit AI-crawler policy: block, allow, or charge.
The 2026 Report6
Question 04What happened to the app stores?

The 30% toll became a negotiation.

For the first time since 2008, web-checkout economics are a real design decision on mobile — in the US, the EU, and Japan, for different reasons.

FrontWhere it landed
US · Apple→External purchase links are legal; the fee is being litigated. The 2025 injunction forced 0%-commission link-outs; the Ninth Circuit later allowed a fee capped at "genuinely necessary" costs — Apple has proposed up to 15%.PRIMARY
US · Google→Play allowed external links and alternative billing from late 2025 after Epic's injunction was upheld; the settlement was filed March 2026.PRIMARY
EU · DMA→€500M anti-steering fine; the Core Technology Fee became a 5% commission; alternative app stores are live — and Japan's new law brought Epic Games Store and AltStore there in January 2026.PRIMARY
So what do you do?
  • Model LTV under 0–15% link-out fees versus 15–30% IAP — per market, since the answer now differs by region.
  • Build the web-checkout funnel now, even if you don't switch: the option value is real and the rails finally exist.
The 2026 Report7
Question 05What's happening to subscription money?

Supply exploded. The median stood still.

The subscription-app dataset (115,000+ apps, $16B+ tracked revenue) tells one story in two numbers: ~15,000 new subscription apps launch monthly — up from ~2,000 in 2022 — while median MRR growth sits at 5.3%. The top decile grew 306%. In vendor telemetry, with its limits: the gap is craft.

5×hard paywalls out-convert freemium by day 35 (10.7% vs 2.1%)PRIMARY
>60%of conversions happen inside week one — the first session is the businessPRIMARY
+41%revenue per payer for AI-powered apps…PRIMARY
30–36%…and that much faster churn — novelty monetizes, then leavesPRIMARY
41→52%AI startup gross margins 2024→2026, vs ~80% classic SaaSSECONDARY
+126%YoY growth in credit-based pricing — the AI-monetization bridgeSECONDARY
So what do you do?
  • Front-load real value into the first minutes — over half of short-trial cancellations land on day zero, and week one decides most of the revenue.
  • Price AI features with COGS in the model from day one: caps, tiers, routing to cheaper models. SaaS's 80%-margin reflexes bankrupt AI products.
  • Default to hybrid pricing — a seat floor plus a usage or credit meter. Pure seats underprice heavy users; pure usage kills budget predictability.
The 2026 Report8
Question 06Do users even want "AI-powered"?

The label is now a conversion tax.

The strangest 2025–26 finding for anyone writing marketing copy: saying "AI" out loud measurably hurts.

In controlled experiments, identical products described as using "artificial intelligence" drew lower purchase intent in every tested case than the same products described as "high tech" — mediated by reduced emotional trust, strongest for high-risk purchases.PRIMARY The climate agrees: 50% of US adults are now more concerned than excited about AI (37% in 2021), and Gartner projects over 40% of agentic-AI projects canceled by end-2027, coining "agent washing" for the rest.PRIMARY

Controlled experiments + large surveys

Two MIT results complete the picture. The NANDA report found 95% of enterprise GenAI pilots produce no P&L impact — the 5% that work are embedded in workflows with memory and feedback loops, not chat windows.PRIMARY And the "Your Brain on ChatGPT" EEG study (n=54, preprint — directional, not definitive) measured "cognitive debt": LLM users showed the weakest brain connectivity and carried the deficit into unassisted work.PRIMARY

So what do you do?
  • Market the outcome, not the model. Disclose AI where the law requires (see Question 07) — but stop leading with it in consumer copy.
  • Put AI inside existing workflows — inline, structured, stateful. Reserve open chat for genuinely open-ended intent.
  • If your product teaches or trains, design AI as scaffold-then-fade — features that do everything may erode the competence your retention depends on.
The 2026 Report9
Question 07Which new rules actually bind?

The rules that bind are the ones with invoices.

Regulatory noise was constant; four things actually changed what shipping software requires.

RuleWhat it means in your product
EU AI Act, Article 50 — in force Aug 2, 2026→Chatbots must disclose they're AI; synthetic media must be labeled and machine-readably marked. High-risk obligations slipped to Dec 2027 — transparency did not. Fines reach €15M / 3% of turnover.PRIMARY
FTC click-to-cancel vacated (July 2025)→California became the US baseline: AB 2863 requires cancel-in-the-same-medium, affirmative consent with records, and renewal reminders. ROSCA enforcement never paused.PRIMARY
CPPA enforcement→Broken opt-outs now carry real fines: Honda $632,500; Tractor Supply a record $1.35M for a non-functioning Global Privacy Control signal. Regulators are testing behavior, not policy text.PRIMARY
European Accessibility Act (June 2025)→WCAG AA is market-entry table stakes in the EU. Enforcement so far is audits, notices, and court orders — a French court ordered Carrefour to make its site and app fully accessible in June 2026.SECONDARY
So what do you do?
  • Ship AI-disclosure UX and content marking for anything EU-facing now; put high-risk classification on a 2027 track.
  • Build cancellation to California's standard everywhere — symmetric, reminded, recorded. It's also just the trust-positive design.
  • Click your own "Do Not Sell" link with GPC on and verify trackers actually stop — that exact test is what the fines were for.
The 2026 Report10
What we refused to printFailed our source-check · named so you can stop repeating them

Six famous 2026 "facts" with nothing underneath.

Each of these circulated widely this year. We went looking for the primary source; here is what we found instead.

"Amazon's AI-assisted deploy caused a 6-hour outage and 6.3M lost orders"

The story appears only in a single low-authority source chain — no primary confirmation, no major-outlet reporting. A perfect cautionary tale, which is exactly why it should have a source and doesn't.

FAILED CHECK · UNCONFIRMED STORY

"75% of new Google code is AI-generated"

Circulates only in secondary aggregations. The verifiable executive claims are far more modest: ~30% at Microsoft (Nadella, April 2025) and ~25–30% at Google per Pichai's earlier statements.

FAILED CHECK · SECONDARY AGGREGATORS ONLY

"30% of new apps will use AI-driven adaptive UI by 2026" (attributed to Gartner)

Appears only in vendor blogs citing Gartner — no locatable Gartner press release or report contains it. Generative-UI tooling is real; this measurement is not.

FAILED CHECK · UNLOCATABLE CITATION

"Charm pricing lifts sales 24–35%"

Traces to 1990s–2000s catalog studies. The 2026 preregistered lab work finds 9-endings improve price image while signaling lower quality, and a preregistered 2022 study of 4,788 purchase decisions failed to reproduce left-digit effects. Test on your own product.

STALE EVIDENCE · MODERN TESTS ARE MIXED

"Pause-instead-of-cancel: 51.7% acceptance, 60–80% reactivation"

Vendor marketing data from billing platforms, not experiments. The pattern is cheap to test and directionally plausible — run your own holdout before quoting anyone's numbers.

VENDOR DATA · NO EXPERIMENT

"The AI slowdown study proves AI makes developers slower"

Also its opposite: "the Copilot study proves +26% for everyone." Both flatten studies that disagree by design — and the slowdown study's authors have withdrawn its headline number themselves. The honest 2026 answer is: it depends on task, tenure, and codebase, and perception is not measurement.

OVERCLAIM IN BOTH DIRECTIONS
The 2026 Report11
The 2026 checklistOne pass over your own product

Eight questions to ask your product this quarter.

  • Is verification a first-class function — review SLAs, small AI PRs, scanning in the agent loop?Question 02 — the new bottleneck
  • Do CI gates catch duplication and churn before they compound?Question 02 — the measurable debt
  • Does your acquisition model survive search sending 30% fewer clicks?Question 03 — owned channels
  • Are you legible to LLMs — clean docs, llms.txt, an MCP server — with a deliberate crawler policy?Question 03 — the new surface
  • Have you modeled web checkout against IAP for each region you sell in?Question 04 — the opened stores
  • Does a first-session user reach real value in minutes — and does your pricing know its COGS?Question 05 — day zero decides
  • Does your copy sell the outcome instead of the model?Question 06 — the label tax
  • Do disclosure, cancellation, opt-outs, and accessibility meet the 2026 bar — verified by clicking, not by policy text?Question 07 — the rules with invoices

Sources: METR (2025 RCT + Feb 2026 update) · Cui et al., SSRN 4945566 · DORA 2025 · GitClear 2026 · Stack Overflow 2025 · LinearB 2026 · Veracode 2025 · SparkToro/Similarweb 2026 · Ahrefs · Pew Research · Cloudflare · RevenueCat State of Subscription Apps 2026 · WSU/Cicek et al. · Gartner · MIT NANDA · MIT Media Lab · EU AI Act Article 50 · California AB 2863 · CPPA enforcement actions · court reporting on Epic v. Apple/Google. Confidence labels mark whether we fetched the primary source or relied on credible secondary reporting.

The 2026 Report12