O013 - ARE ARTIFICIAL INTELLIGENCE MODELS READY FOR CLINICAL DECISION-MAKING? A COMPARATIVE STUDY BASED ON ESPEN GUIDELINES
O013
ARE ARTIFICIAL INTELLIGENCE MODELS READY FOR CLINICAL DECISION-MAKING? A COMPARATIVE STUDY BASED ON ESPEN GUIDELINES
A. G. Çapar1, M. Kip1,*, H. Altun1, G. Kendirli1, E. Başmısırlı1, N. İnanç1, M. Khosravi2
1Nutrition and Dietetics, Nuh Naci Yazgan University, Kayseri, Türkiye, 2University of Isfahan, Isfahan, Iran, Islamic Republic Of
Rationale: The increasing use of artificial intelligence in clinical decision-making has raised important questions regarding its accuracy and reliability compared to human experts.
Methods: In this study, the accuracy of 47 case-based multiple-choice questions, developed by three researchers in the field, was assessed based on the ESPEN nutrition guidelines. All questions were answered independently by two expert human raters in the field of clinical nutrition and three artificial intelligence models.
Results: No statistically significant differences were observed among the artificial intelligence models according to the McNemar test (p > 0.05). Nevertheless, Gemini (72.3%) and Copilot (70.2%) exhibited higher accuracy rates compared to ChatGPT (55.3%). In contrast, a high level of agreement was observed between the human raters (89.4%). A paired t-test demonstrated that human raters significantly outperformed artificial intelligence models, with a mean accuracy of 94.7% compared to 66.0% for artificial intelligence systems (p < 0.001). This indicates a substantial performance gap favoring human evaluators. For calculation-based questions, human raters achieved significantly higher accuracy than the artificial intelligence models (94.3% and 61.0%, respectively; p < 0.001), with a large effect size (Cohen’s d = 0.86).
Conclusion: While artificial intelligence systems performed similarly to each other, their accuracy was notably lower than that of human raters. Additionally, calculation-based questions appear particularly challenging for current AI systems. These findings highlighting potential limitations in their problem-solving capabilities and indicating a substantial practical gap in performance.
Disclosure of Interest: None declared