O001 - COMPARATIVE EVALUATION OF ARTIFICIAL INTELLIGENCE SYSTEMS AND CLINICAL PHARMACISTS IN MEDICATION ADMINISTRATION VIA ENTERAL FEEDING TUBES: A SCENARIO-BASED STUDY

Linked sessions

O001

COMPARATIVE EVALUATION OF ARTIFICIAL INTELLIGENCE SYSTEMS AND CLINICAL PHARMACISTS IN MEDICATION ADMINISTRATION VIA ENTERAL FEEDING TUBES: A SCENARIO-BASED STUDY

N. Ozdemir Ayduran1, E. Aras Atik2,*, K. Tecen3

1Department of Clinical Pharmacy, Gazi University Faculty of Pharmacy, Ankara, 2Department of Clinical Pharmacy, Atatürk University Faculty of Pharmacy, Erzurum, 3Department of Clinical Pharmacy, Anadolu University Faculty of Pharmacy, Eskişehir, Türkiye

 

Rationale:  

Medication administration via enteral feeding tubes is complex and may lead to adverse outcomes if performed inappropriately. While artificial intelligence (AI) is increasingly used in healthcare, its role in specialized pharmaceutical care remains unclear. This study compared the accuracy, risk scores, and inter-rater agreement of ChatGPT, Gemini, Claude, DeepSeek, and two experienced clinical pharmacists in enteral medication administration scenarios.

Methods: Ten clinical scenarios covering key aspects of enteral drug administration were developed, each with four response options including “I don’t know.” All participants responded to the same prompt, and answers were evaluated for accuracy, risk scores, and agreement using Cohen’s and Fleiss’ kappa. Analyses were performed using IBM SPSS Statistics (Version 31.0) and R (Version 4.5.1).

Results: ChatGPT, Gemini, and Claude achieved perfect accuracy (10/10), while DeepSeek answered 8/10 scenarios correctly. Among pharmacists, one achieved 10/10 and the other 9/10. Agreement among AI systems was almost perfect (Fleiss’ κ = 0.806), while agreement between pharmacists was substantial (Cohen’s κ = 0.787). Overall agreement across all participants was high (Fleiss’ κ = 0.801). DeepSeek’s incorrect responses had risk scores of 3 and 2 (total = 5), whereas the pharmacist’s single incorrect response had a risk score of 4. 

Conclusion: AI systems and clinical pharmacists demonstrated comparable accuracy, with slightly higher consistency among AI responses. However, variation in risk scores indicates that the clinical impact of errors may differ, underscoring the importance of expert oversight in complex medication decisions.

Disclosure of Interest: None declared