ทันเอไอ

Global AI & tech news, in your language · every story source-checked

🌐Follow
← Back to news
LINE Facebook X
AI Models & ResearchVerified

Research Finds LLMs Can Help Review Disease Models, but Still Need Human Oversight

Researchers developed a pipeline that uses LLMs to extract data from 536 disease-spread modeling papers, achieving article-level accuracy of up to about 81.67%, though performance varied widely with data complexity.

📅 29 Aug 2026, 01:33
Research Finds LLMs Can Help Review Disease Models, but Still Need Human Oversight

A study titled “Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models” tested the use of Large Language Models, or LLMs, to assist with systematic literature reviews (SLRs), which typically require considerable time to find, read, and classify academic papers using consistent criteria.

The research team developed a multi-step pipeline to extract model-related information from 536 peer-reviewed papers. All focused on agent-based modeling, in which individual agents have their own behaviors and interactions, to study disease spread. The system’s results were then compared with a human-conducted SLR.

At the article level, GPT-4.1 achieved accuracy of about 77.95%, while GPT-5.0 reached about 81.67%. However, these overall figures do not mean the models performed equally well across every type of data. Field-level accuracy ranged widely from 32.40% to 100.00%, with complex fields or those requiring subjective interpretation proving less reliable.

The study also found that agreement among multiple LLMs could serve as an initial signal of output quality. Strong disagreement may indicate that some answers contain hallucinations—plausible-looking information that does not match the source. This could help identify fields that need closer human review, though model consensus should not be treated as proof that an answer is correct.

The findings support using LLMs to reduce the burden of data extraction rather than fully replace researchers, particularly for tasks requiring contextual interpretation or judgment. The paper was published on arXiv under cs.AI and is scheduled for presentation at Winter Simulation Conference 2026.

Why it matters
Faster reviews of large volumes of disease research could help Thai researchers and public-health agencies keep up with new knowledge. However, the wide accuracy range shows that experts must verify the results before using them to make decisions.
#LLM#งานวิจัย AI#แบบจำลองการแพร่โรค#การทบทวนวรรณกรรม
Sources (rewritten & summarized from): arXiv cs.AI · insiderly.ai · arxiv.org · arxiv.org

Comments

Loading comments…

No sign-up needed · comments are auto-filtered and moderated