The SCORE Initiative: How Much Can We Trust Social Science Research?

If you work in psychology or social science research, you have probably heard about the replicability crisis. For years, the debate has relied on individual studies and conflicting opinions. But in April 2026, the journal Nature simultaneously published four studies offering the most ambitious and systematic evaluation to date. The project is called SCORE, and its results compel us to reflect on how we do science.

What is SCORE and why does it matter?

SCORE stands for Systematizing Confidence in Open Research and Evidence. It is a project funded by DARPA with approximately $8 million, coordinated by the Center for Open Science (COS). What makes it unique is its scale: 865 collaborators evaluated nearly 3,900 scientific claims from the social and behavioral sciences. This is not a laboratory study or a narrative review: it is an industrial-scale effort in scientific verification.

The results were published in volume 652 of Nature, on April 2, 2026, across four complementary articles. Each addresses a different dimension of scientific reliability: reproducibility, replicability, and robustness. Together, they provide a comprehensive diagnosis of the state of evidence in our disciplines.

The three Rs of scientific credibility

Before diving into the results, it is worth clarifying three concepts that are often confused:

  • Reproducibility: Can the same results be obtained from the same data and analysis code? This is the most basic test: if another researcher runs your script with your data, do they arrive at the same numbers?
  • Replicability: Are similar results obtained when the study is repeated with new data? Here, an independent sample is collected and the same procedure is applied to see if the effect holds.
  • Robustness: Do the conclusions hold when the same data are analyzed with alternative statistical methods? A robust result does not depend on a particular analytical decision.

An important contribution of SCORE is demonstrating that these three dimensions are relatively independent of each other. A study can be reproducible but not replicable, or replicable but not robust. Scientific credibility is multidimensional, and evaluating it requires examining each R separately.

Reproducibility: the problem of code and data

The study by Miske and colleagues examined 600 articles from the social and behavioral sciences. The most alarming finding was the starting point: only 20% of articles included both the data and analysis code necessary to attempt to reproduce the results. That is, in eight out of ten published studies, it was not even possible to verify whether the numbers in the article were correct.

But here comes the most interesting part. When data and code were fully available, the reproducibility rate reached 91%. However, when materials were partial and analysts had to reconstruct the procedures, the rate plummeted to 38%. The difference between sharing or not sharing your materials is not a nuance: it is the difference between your work being verifiable or not.

Replicability: half of the effects do not survive

The study by Tyner and colleagues attempted to replicate 164 social science articles. The results confirm what previous studies had suggested, but at an unprecedented scale: approximately half of the originally statistically significant effects did not reach statistical significance in the replication, a problem directly tied to the tyranny of the p-value.

Even more revealing is what happened with effect sizes. On average, replicated effects were less than half the size originally reported. This means that even when an effect does replicate, its true magnitude is considerably smaller than the original study suggested. For those of us who design studies or calculate sample sizes, this finding has direct implications: if we base our power calculations on published effects, we are probably underestimating the required sample size.

Robustness: when the statistical method matters more than we think

Aczel and colleagues evaluated the robustness of 100 articles by subjecting each conclusion to alternative statistical analyses. The result: only 74% of statistically significant conclusions remained significant when different methods were used (tools like p-curve and evidential value help detect these issues).

This means that one in four conclusions depends on the researcher's particular methodological choice. It is not that the original analysis was necessarily poorly done, but that the conclusion was not solid enough to withstand a different but equally valid analytical approach. For reviewers and readers, this underscores the importance of authors justifying their analytical decisions and, whenever possible, presenting sensitivity analyses.

Journal policies work

The study by Brodeur and colleagues, carried out within the Institute for Replication, provides a note of optimism. Their data show that in journals that require sharing data and code, 85% of studies were reproducible. Furthermore, they documented a notable shift in community practices: while in 2014 only 59% of articles shared data or code, between 2021 and 2023 that figure rose to approximately 90%.

This finding is fundamental because it demonstrates that editorial policies have a real impact. It is not just about good intentions: when journals require transparency, researchers adopt it, and the resulting science is more reliable.

What does this mean for your research?

If you are preparing a study, thesis, or article in psychology or social sciences, the SCORE results have concrete practical implications:

  • Always share your data and code. The difference between 91% and 38% reproducibility speaks for itself. Use repositories like OSF or GitHub. Many journals already require it, and those that do not will value it positively.
  • Do not blindly trust published effect sizes. If you are planning a study based on effects from the literature, consider that the true magnitude could be half or less. Calculate your sample size with conservative estimates.
  • Include sensitivity analyses. Demonstrate that your conclusions do not depend on a particular analytical decision. Try different model specifications, exclusion criteria, or statistical methods and report the results.
  • Pre-register your studies. In a context where half of effects do not replicate, pre-registration on OSF protects your credibility by demonstrating that you have not manipulated your hypotheses or analyses after seeing the data.
  • View transparency as a competitive advantage. The highest-impact journals are adopting open data policies. Anticipating these requirements strengthens your manuscript and speeds up the review process.

A more honest science, not a science in crisis

It is tempting to read these results as a condemnation of the social sciences. But a more accurate reading is different: the scientific community is investing unprecedented resources in evaluating and improving its own practices. SCORE is not an accusation: it is a diagnosis. And like any good diagnosis, it is the first step toward improvement.

The data show that the tools to produce more reliable science already exist: sharing data, sharing code, pre-registering, and applying sensitivity analyses. It is not about revolutionizing methodology, but about adopting practices that have already proven their effectiveness. The science that survives verification is not perfect science; it is transparent science.

If you need guidance on how to implement these practices in your research, from pre-registration to preparing data and code repositories, I can help you strengthen your work to withstand any scrutiny.

References

  • Tyner, S. et al. (2026). Replicability of social and behavioural science claims in a large-scale assessment. Nature, 652, 143-150.
  • Aczel, B. et al. (2026). Robustness of social and behavioural science conclusions. Nature, 652, 135-142.
  • Miske, O. et al. (2026). Reproducibility of social and behavioural science articles. Nature, 652, 126-134.
  • Brodeur, A. et al. (2026). Replication games: reproducibility and replicability across the social sciences. Nature, 652, 151-156.
  • Center for Open Science. SCORE Project. https://www.cos.io/

Keep reading

All blog articles