Generative Artificial Intelligence in Internal Medicine: Applications in Clinical Reasoning, Differential Diagnosis, and Clinical Decision-Making
Systematic review
DOI:
https://doi.org/10.64911/zy8pc019Keywords:
Generative Artificial Intelligence; Large Language Models; Clinical Decision Support Systems; Diagnostic Accuracy; Internal MedicineAbstract
Objective: To synthesize the available evidence on the diagnostic accuracy, clinical decision-support performance, advantages, and limitations of generative AI in adult internal medicine.
Methodology: This systematic review was conducted in the Department of Medicine, MMC/BKMC, Mardan, from January to June 2026. Following PRISMA 2020 guidelines, we searched PubMed/MEDLINE, Embase, Web of Science, Scopus, and preprint servers for studies published from January 2022 to September 2026. Eligible studies evaluated generative AI in clinical reasoning, differential diagnosis, or clinical decision-making in adult internal medicine. A total of 58 studies met the eligibility criteria, and 25 were selected for detailed synthesis based on their relevance and methodological contribution.
Results: Generative AI performed strongly on selected clinical reasoning and diagnostic tasks. Across complex clinical vignettes, LLMs identified the correct diagnosis within the top 5 or top 10 differential diagnoses in approximately 60–83% of cases. They achieved or exceeded passing performance on several medical licensing and internal medicine board-style examinations. In structured OSCE-style evaluations, purpose-built conversational systems performed favorably compared with physicians across several assessment domains. However, randomized trials evaluating general-purpose LLMs in realistic clinical workflows demonstrated smaller and less consistent improvements, particularly in interactive diagnostic tasks. Key limitations included hallucination, demographic bias, automation bias, variation across specialties and tasks, and inconsistent evaluation methods. Only approximately 5% of studies in a broader systematic review of healthcare LLM applications used real patient-care data.
Conclusion: Current evidence supports using generative AI as an adjunct to, rather than a replacement for, physician clinical reasoning. Safe implementation requires structured workflows, appropriate human oversight, retrieval-augmented generation, standardized evaluation, and further validation using real-world clinical data.
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.





