LLM Assistance Improves Diagnostic Accuracy in Medicine and What the Evidence Means for Safe AI Integration

The medRxiv preprint “The Effect of LLM Assistance on Diagnostic Accuracy: A Meta-Analysis” by Isabel Tretow, Moritz Schwebel, Stefan Feuerriegel, Theresa Treffers, and Isabell M. Welpe examines a question that is moving from speculation to practice: whether large language models actually help physicians diagnose more accurately. Based on 15 studies, 43 effect sizes, 498 physicians, and 7,274 case evaluations, the authors find that LLM support improves diagnostic accuracy, but they also show that the size of that benefit depends heavily on context, model choice, and workflow design. For anyone tracking the rapid expansion of medical AI, the study is important because it moves the discussion away from hype and toward conditions for safe, effective use.

The headline result is encouraging but not simplistic. The pooled effect size is positive, with Hedges’ g at 0.202 and a confidence interval that excludes zero, which means that physicians using LLM assistance performed better than those working without it. That improvement was seen across several models, including GPT-4, AMIE, and MedFound-DX-PA, and across different medical fields and career stages. At the same time, the authors report substantial heterogeneity, which is a warning sign that not every setting benefits equally and that the real value of LLMs depends on how they are integrated into diagnostic work.

A particularly interesting finding is that the best results were not tied to pure automation, but to collaboration. The paper emphasizes that LLMs are most likely to matter in real clinical practice as assistive tools under physician supervision, not as stand-alone diagnosticians. That distinction matters because it points to a broader design principle for AI in high stakes environments: technology should augment human judgment rather than replace it. The study also identifies four key determinants of effectiveness, namely the LLM model, the medical field, the physician’s career stage, and the response format used to present or solicit diagnostic input.

For clinical governance, the risk section is just as important as the positive effect. The authors note that LLMs can introduce automation bias, anchoring effects, and hallucinated advice, all of which may compromise patient safety if users over trust the output. They also report that 6 of the 15 studies had high risk of bias, and 4 had applicability concerns, often because the same physicians were tested with and without LLM support or because some cases may already have been present in model training data. This makes a strong case for stricter oversight, better validation, and careful boundary setting before clinical deployment, especially in settings where the cost of error is high.

The methodological design is another strength of the paper. By using a three level meta analytic model, the authors were able to account for multiple effect sizes within the same experiment, which makes the findings more credible than a simple pooled comparison would have been. The robustness checks also support the stability of the overall result, with no clear sign of publication bias and similar estimates across sensitivity analyses. Even so, the paper is careful to note that most studies were controlled experiments rather than real world deployments, so ecological validity remains an open question.

The practical implication is that AI in medicine should be governed as a structured capability, not adopted as a shortcut. Hospitals, medical schools, and health system leaders should focus on the conditions that make LLM support useful, including prompt design, response format, clinician training, and escalation rules for uncertainty. The most useful takeaway for broader AI adoption is that effectiveness comes from disciplined human AI collaboration, not from the model alone. The full article “The Effect of LLM Assistance on Diagnostic Accuracy: A Meta-Analysis” by Isabel Tretow, Moritz Schwebel, Stefan Feuerriegel, Theresa Treffers, and Isabell M. Welpe is available here.