G_4_14

Replication Crisis and What It Means for Ancient Claims

Verified (Tier 1)
Confidence: 4/5 Section: G Updated: March 10, 2026
Source Count: 13 | Weighted Score: 35 | Source Confidence: [4/5] | Primary Tier: 1–2 | Last Updated: March 10, 2026
Keywords: replication crisis, reproducibility, p-hacking, HARKing, publication bias, open science, pre-registration, Ioannidis, Open Science Collaboration, file drawer problem, effect size, statistical significance, meta-analysis, questionable research practices, underpowered studies, scientific fraud
Category Tags: modern-frameworks, methodology, statistics, science policy, epistemology
Cross-References: H_2_04 — Academic Suppression · H_2_03 — Peer Review Problems · T_1_01 — Psychology Overview · G_2_03 — Bayesian Reasoning

QUICK SUMMARY

The replication crisis refers to the discovery, beginning in the early 2010s, that a substantial proportion of findings published in peer-reviewed scientific journals — particularly in psychology, social science, and biomedical research — fail to replicate when independent researchers attempt to reproduce them using the original methods. The landmark study was the Open Science Collaboration (OSC, 2015, Science), which attempted to replicate 100 psychology studies published in three top journals: only 36% produced statistically significant results in the same direction as the original, and the average effect size of replications was roughly half the original. This crisis was anticipated by John Ioannidis's influential paper "Why Most Published Research Findings Are False" (2005, PLOS Medicine), which used Bayesian reasoning to show that in fields with low prior probability of true effects, small sample sizes, flexible statistical analysis, and strong publication bias, the majority of statistically significant results will be false positives. The structural causes are now well-documented: (1) p-hacking — researchers (often unconsciously) test multiple analyses, variables, or subgroups until they find p < 0.05, inflating the false positive rate far beyond the nominal 5%; (2) HARKing (Hypothesizing After Results are Known) — presenting post-hoc discoveries as if they were a priori predictions; (3) publication bias — journals preferentially publish "positive" (statistically significant) results, creating a "file drawer" of unreported null results; (4) underpowered studies — small sample sizes that can only detect large effects, meaning that any significant result is likely an overestimate (the "winner's curse"). For ancient claims and alternative history, the replication crisis has profound implications: many archaeological and historical claims rest on studies from fields now known to have replication problems (e.g., priming effects, social psychology experiments used to explain ancient behavior, small-sample genetic studies). More broadly, the crisis provides a framework for understanding why extraordinary claims based on a single study — no matter how high-profile the journal — should be treated with caution. The response to the crisis has been the open science movement: pre-registration of hypotheses and analysis plans before data collection, mandatory data sharing, registered reports (peer review before results are known), and a shift from null-hypothesis significance testing (NHST) toward estimation-based approaches (effect sizes with confidence intervals) and Bayesian methods. These reforms are gradually being adopted in archaeology and ancient studies, but the field still relies heavily on small-sample, unreplicated findings.


1. VERIFIED CLAIMS (Tier 1 — Peer-Reviewed / Scholarly Consensus)

1.1 Scope of the Replication Crisis

1.2 Structural Causes — P-hacking and Publication Bias

1.3 Ioannidis's Bayesian Argument


2. CREDIBLE CLAIMS (Tier 2 — Academic / Debated but Supported)

2.1 Implications for Archaeology and Ancient Studies

2.2 Open Science Reforms


3. SPECULATIVE CLAIMS (Tier 3 — Possible but Unverified)

3.1 Has the Crisis Undermined Public Trust in Science?


4. DUBIOUS CLAIMS (Tier 4 — No Credible Source / Contradicted by Evidence)

4.1 The Replication Crisis Proves Mainstream Science Is a Conspiracy


Counter-Arguments & Criticisms

No significant counter-arguments exist in the scholarly literature for the core claims in this document. Replication Crisis and What It Means for Ancient Claims represents established scientific and methodological consensus with no active scholarly dispute over the fundamental claims presented here.


IMAGES

#DescriptionFilenameSourceLicense

No images assigned yet.


BIBLIOGRAPHY

  1. Open Science Collaboration. aac4716 | 2015 | "Estimating the Reproducibility of Psychological Science" | Science | ∅ | 349:: | ∅ | ∅ | doi:10.1126/science.aac4716 | ∅ | ∅ | ∅
  2. Ioannidis, J.P.A. e124 | 2005 | "Why Most Published Research Findings Are False" | PLOS Medicine | ∅ | 2:: | ∅ | ∅ | doi:10.1371/journal.pmed.0020124 | ∅ | ∅ | ∅
  3. Simmons, J.P. et al | 2011 | "False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant" | Psychological Science | ∅ | 22::1359–1366 | ∅ | ∅ | doi:10.1177/0956797611417632 | ∅ | ∅ | ∅
  4. Camerer, C.F. et al | 2018 | "Evaluating the Replicability of Social Science Experiments in Nature and Science between 2010 and 2015" | Nature Human Behaviour | ∅ | 2::637–644 | ∅ | ∅ | doi:10.1038/s41562-018-0399-z | ∅ | ∅ | ∅
  5. Begley, C.G.; Ellis, L.M | 2012 | "Raise Standards for Preclinical Cancer Research" | Nature | ∅ | 483::531–533 | ∅ | ∅ | doi:10.1038/483531a | ∅ | ∅ | ∅
  6. Button, K.S. et al | 2013 | "Power Failure: Why Small Sample Size Undermines the Reliability of Neuroscience" | Nature Reviews Neuroscience | ∅ | 14::365–376 | ∅ | ∅ | doi:10.1038/nrn3475 | ∅ | ∅ | ∅
  7. Fanelli, D. e10068 | 2010 | "'Positive' Results Increase Down the Hierarchy of the Sciences" | PLOS ONE | ∅ | 5:: | ∅ | ∅ | doi:10.1371/journal.pone.0010068 | ∅ | ∅ | ∅
  8. Franco, A. et al | 2014 | "Publication Bias in the Social Sciences: Unlocking the File Drawer" | Science | ∅ | 345::1502–1505 | ∅ | ∅ | doi:10.1126/science.1255484 | ∅ | ∅ | ∅
  9. Nosek, B.A. et al | 2018 | "The Preregistration Revolution" | PNAS | ∅ | 115::2600–2606 | ∅ | ∅ | doi:10.1073/pnas.1708274114 | ∅ | ∅ | ∅
  10. Errington, T.M. et al. e71601 | 2021 | "Investigating the Replicability of Preclinical Cancer Biology" | eLife | ∅ | 10:: | ∅ | ∅ | doi:10.7554/eLife.71601 | ∅ | ∅ | ∅
  11. Munafò, M.R. et al | 2017 | "A Manifesto for Reproducible Science" | Nature Human Behaviour | ∅ | 1::0021 | ∅ | ∅ | doi:10.1038/s41562-016-0021 | ∅ | ∅ | ∅
  12. Kerr, N.L | 1998 | "HARKing: Hypothesizing After the Results are Known" | Personality and Social Psychology Review | ∅ | 2::196–217 | ∅ | ∅ | doi:10.1207/s15327957pspr0203_4 | ∅ | ∅ | ∅
  13. Chambers, C.D | 2017 | ∅ | The Seven Deadly Sins of Psychology: A Manifesto for Reforming the Culture of Scientific Practice | ∅ | ∅ | Princeton: Princeton University Press | ∅ | ∅ | ∅ | ∅ | ∅

CROSS-REFERENCE INDEX

Related DocConnection

No cross-references yet.


⚠️ AI-Assisted Research Disclaimer

This document was generated and structured with the assistance of AI tools.

While every effort is made to ensure accuracy, AI-assisted content may

contain errors, misattributions, or unintended inaccuracies. Always verify claims, dates, and sources independently before citing or relying

on any information presented here.

  • Sources may contain errors. Bibliography entries and cross-references

are checked by automated systems, but mistakes can occur. If something

looks wrong, it may be.

  • Speculative and unverified claims are clearly labeled. This project

uses a four-tier evidence system:

  • Tier 1 — Verified: Peer-reviewed, established scientific consensus.
  • Tier 2 — Credible: Academically supported, debated but grounded.
  • Tier 3 — Speculative: Plausible but unverified by mainstream science.
  • Tier 4 — Dubious: No credible support or contradicted by evidence.
  • This project maps multiple perspectives — not a single truth. Mainstream,

alternative, and skeptical viewpoints are presented side by side for

critical comparison, not endorsement. Inclusion does not imply agreement.

  • We are actively improving. Source verification, factuality scoring,

and bibliography enrichment are ongoing. Each revision adds stronger

citations, corrects identified errors, and expands coverage.

📖 For full details on our verification methodology, scoring systems, and

quality metrics, see: Fact-Checking & Verification Systems

Think Openly. Check the sources. Draw your own conclusions.