Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science
Univ Ottawa, Canada; Inst Replicat, Canada.
Univ Ottawa, Canada; Inst Replicat, Canada.
Univ Cambridge, England.
Univ Ottawa, Canada; Inst Replicat, Canada.
Show others and affiliations
2026 (English)In: Proceedings of the National Academy of Sciences of the United States of America, ISSN 0027-8424, E-ISSN 1091-6490, Vol. 123, no 22, article id e2524747123Article in journal (Refereed) Published
Abstract [en]

Large Language Models (LLMs) such as ChatGPT are transforming how scientists conduct and validate research, offering promise as tools to improve scientific reproducibility. However, computational reproducibility and error detection remain expensive and labor-intensive. We experimentally test how collaboration between researchers and LLM assistants influences the reproduction of quantitative social science findings across different levels of AI autonomy. We randomly assigned 288 researchers to 103 teams working under three conditions: human-only, AI-assisted (using ChatGPT as a collaborative tool), or AI-led (ChatGPT operating with minimal human oversight). Teams reproduced published results from leading social science journals, detected coding errors, and proposed robustness checks. Human-only and AI-assisted teams achieved comparable reproduction rates (94% vs. 91%) and performed similarly on most outcomes, except human-only teams identified significantly more major coding errors. Both substantially outperformed AI-led teams, which achieved only a 37% reproduction rate, detected fewer errors across all categories, proposed weaker robustness checks, and required more time. This autonomous approach, however, likely represents only a lower bound of AI capabilities. Despite rapid model advances, expert human judgment currently remains indispensable for reliable empirical verification. While AI assistance did not degrade most outcomes, it provided no measurable advantages and was associated with reduced detection of major errors. However, the 37% autonomous reproduction rate indicates that AI could provide value in settings where scale or cost constraints preclude human review of papers, even though general-purpose LLMs offer no immediate advantages for human-supervised verification.

Place, publisher, year, edition, pages
NATL ACAD SCIENCES , 2026. Vol. 123, no 22, article id e2524747123
Keywords [en]
AI; reproducibility; large language models
National Category
Computer graphics and computer vision
Identifiers
URN: urn:nbn:se:liu:diva-226926DOI: 10.1073/pnas.2524747123ISI: 001821087300001Scopus ID: 2-s2.0-105040657951OAI: oai:DiVA.org:liu-226926DiVA, id: diva2:2094598
Note

Funding Agencies|Coefficient Giving project "Benchmarking LLM agents on real-world tasks: Reproducibility"; Alfred P. Sloan Foundation Foundation [G-2023-22326]; University of Toronto; University of Ottawa; University of Cornell; University of Tilburg; Leverhulme Early Career Research Fellowship [ECF-2022-761]; Australian Research Council [DP200102935]; Ministero dell'Universita e della Ricerca

Available from: 2026-08-24 Created: 2026-08-24 Last updated: 2026-08-24

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Search in DiVA

By author/editor
Batinovic, Lucija
By organisation
Disability Research DivisionFaculty of Arts and Sciences
In the same journal
Proceedings of the National Academy of Sciences of the United States of America
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 8 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf