Large Language Model-Supported Data Collation: Addressing Accuracy and Reproducibility

Sofie Sellén, Erik Sjögren

Quantitative Medicine · 2026

ABSTRACT Manual data extraction from scientific and regulatory documents is time-consuming and limits the efficiency of research workflows. Large language models (LLMs) offer a potential solution, yet their stochastic nature challenges the absolute accuracy and reproducibility required for scientific data collation. This study utilizes a case example to exemplify the practical challenges regarding the feasibility and reliability of an artificial intelligence (AI)-supported workflow to extract immunogenicity data (antidrug antibody (ADA) incidence) of 50 monoclonal antibodies (mAbs) from diverse regulatory documents within the R programming environment.

The primary objective was to demonstrate the trade-off between stochasticity and reproducibility inherent to current Large Language Models (LLMs) in a scientific life-science context. We explored 3 workflows: a grounded Google search, a direct portable document format (PDF) analysis, and a preprocessed document workflow that converted source materials into concise Markdown reports. The grounded search and the direct PDF analysis proved unsuitable due to data misassociation errors, hallucination, and a lack of reproducibility stemming from dynamic search results.

The preprocessed document workflow successfully resolved data misassociation and achieved accurate extraction of core ADA incidence. However, the central aim of a fully reproducible workflow was not met. Despite using fixed model parameters and identical, highly structured inputs, the output exhibited day-to-day variability in the presentation and occasional omission of key contextual data.

This inherent stochasticity, even when minimal, demonstrates that fully automated data collation by a LLM-driven workflow is not yet reliable for research workflows requiring absolute consistency. In conclusion, while the AI-supported R-based workflow offers a powerful tool to accelerate initial data gathering, this study highlights that the LLM’s lack of guaranteed reproducibility necessitates meticulous verification to ensure the accuracy and completeness demanded by scientific and research standards.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.

Paperis - Large Language Model-Supported Data Collation: Addressing Accuracy and Reproducibility