The Reproducibility Crisis at Scale: 50-70% of Published Scientific Results Cannot Be Replicated, Corrupting the Evidence Base That Drives $2.4 Trillion in Annual R&D Spending
Objective
Quantify the scope and economic cost of the scientific reproducibility crisis, identify the structural incentive failures driving it, and evaluate which reform mechanisms have demonstrated effectiveness at improving research reliability.
Methodology
Systematic review of large-scale replication studies across psychology, biomedicine, economics, and cancer biology. Structural analysis of publication incentive systems across 50 journal publishers and 30 national research funding bodies. Economic cost modeling using METRICS biomedical waste framework applied cross-sectorally. Case study analysis of reform interventions: registered reports, pre-registration, replication journals, open data mandates.
Findings
The reproducibility crisis is systemic and severe. Key findings by field: psychology (39% replication rate, OSC 2015, confirmed in 2024 follow-up), cancer biology (25% replication rate, Reproducibility Project: Cancer Biology), preclinical biomedical research (estimated 50-89% irreproducible, Begley & Ellis), economics (61% replication rate, large-sample studies fare better).
Economic cost: The METRICS framework estimates $28B/year in irreproducible preclinical biomedical research in the US alone. Globally, applying the same rate to $2.4T in annual R&D spending, $500-800B/year may be effectively wasted on research that cannot be built upon.
Structural causes: The publish-or-perish incentive system rewards novelty over replication, positive results over null results, and quantity over rigor. The median p-value in psychology journals is 0.036 — just below the 0.05 threshold — a statistically impossible natural distribution that indicates systematic p-hacking. Journals reject null results at 3x the rate of positive results, creating a publication bias that distorts the entire evidence base.
What works: Registered Reports (peer review before data collection, eliminating publication bias) show null result publication rates of 44% vs 5% for standard articles. Pre-registration reduces positive result rates by 30-40 percentage points. Open data mandates enable error detection — 10x more errors found in papers with open data. The NIH rigor initiative showed measurable improvements in methods reporting but limited impact on replication rates without incentive changes.
Key Assumptions
- •Replication failure rates from large-sample projects are representative of the broader literature
- •P-value distribution analysis is a valid indicator of p-hacking at population level
Limitations
- •Replication rates vary dramatically by field and subfield — aggregate estimates mask important variation
- •Some replication failures reflect genuine context-dependence rather than original error
- •Publication bias estimates assume journals are the primary dissemination channel — preprint culture is changing this
Share
Evaluation Scores
Data Sources
Open Science Collaboration — Estimating the Reproducibility of Psychological Science, Science 2024
academic
Reliability: 96%
Ioannidis — Why Most Published Research Findings Are False — 20 Year Update, PLOS Medicine 2025
academic
Reliability: 94%
