TL;DR
Reproducibility assessments in social and behavioral sciences are often labor-intensive and challenging to scale. A large language model (LLM) was developed to automate these assessments, analyzing 76 published studies.
✦ Why It Matters
Engineers and researchers can leverage LLMs to streamline reproducibility assessments, enhancing efficiency and scalability in research validation.
Key Takeaways
Full Summary
Reproducibility is crucial in social and behavioral sciences, where independent researchers typically reanalyze data to verify published findings. This process is resource-intensive and not easily scalable.
Researchers developed a large language model (LLM) to automate reproducibility assessments, applying it to 76 studies with predefined claims. The LLM was able to recover original effect sizes in 41% of cases, using a tolerance of +/-0.05 in Cohen's d, a measure of effect size.
Additionally, it aligned with the original study's qualitative conclusions in 96% of instances. In comparison, human reanalysts achieved a 34% recovery rate for effect sizes and 74% for qualitative conclusions.
These findings suggest that LLMs can serve as efficient tools for systematic auditing of empirical results in this field.
Related