PD495 - ACCURACY OF AI-GENERATED NUTRITION ADVICE IN DIABETES: COMPARING THE RELIABILITY OF TWO METHODOLOGIES
PD495
ACCURACY OF AI-GENERATED NUTRITION ADVICE IN DIABETES: COMPARING THE RELIABILITY OF TWO METHODOLOGIES
R. Kleine1,*, L. Aragão2, C. Moraes3, T. Otranto Rossi4, V. Golfieri5, N. L. Zandonadi de Oliveira6, R. F. Manoel4, C. Sennersten1, E. Sonestedt1
1Kristianstad University, Kristianstad, Sweden, 2Oswaldo Cruz Foundation, Rio de Janeiro, 3Sao Francisco University, Sao Paulo, 4State University of Campinas, 5Victoria Golfieri Women's Nutrition Clinic, Campinas, 6Primary Health Care Unit, Sao Paulo, Brazil
Rationale: Type 2 diabetes (T2D) is one of the most prevalent chronic diseases worldwide and patients are increasingly using large language models (LLMs) to obtain nutritional advice. This study aimed to assess the reliability of a new methodology compared to current methods to assess AI generated nutrition advice in diabetes.
Methods: ChatGPT-5 was prompted to generate 45 frequently asked questions in diabetes. Each question was answered with a maximum of 50 words. Responses were then fragmented into individual claims, with each sentence containing a single piece of information and rewritten for standalone clarity. Six registered dietitians (RDs) with ≥3 years of clinical experience in chronic disease were divided into two groups (n=3) to assess response accuracy using the Standards of Care in Diabetes 2026 as the primary reference. Group 1 evaluated full responses using a 4-point Likert scale (1 = not at all accurate; 4 = fully accurate), representing the reference method. Group 2 assessed the responses which have been fragmented using a binary scale based on the accuracy of the statement (Yes/No), representing the proposed method. All assessors received prior training. Inter-rater reliability was evaluated using Fleiss’ kappa (p < 0.05). To enable comparison, claim-level scores were converted to equivalent Likert-scale values.
Results: Fleiss’ kappa coefficients were 0.07 (p > 0.05) for the method of reference and −0.01 (p > 0.05) for the proposed methodology, indicating very low inter-rater agreement in both approaches. Therefore, the proposed method did not improve reliability of assessing accuracy of AI generated advice compared with conventional method.
Conclusion: Further research is needed to develop reliable, systematic, and resource-efficient method for evaluating the accuracy of AI-generated nutrition advice.
Disclosure of Interest: None declared