PD048 - AUTOMATED DIGITIZATION OF ESPEN CLINICAL NUTRITION GUIDELINES: AN LLM-BASED KNOWLEDGE GRAPH APPROACH

Linked sessions

PD048

AUTOMATED DIGITIZATION OF ESPEN CLINICAL NUTRITION GUIDELINES: AN LLM-BASED KNOWLEDGE GRAPH APPROACH

 

F. Menegoni1, I. Buttignon1,*, P. Pisano2,3,4, E. FERRARI3, S. Bo5, A. Devecchi5, E. Mazzetto6

1Human Technology eXcellence, HTX SRL, Trieste, 2Economic and Statistics , 3Highest Lab, University of Torino , Torino, Italy, 4Network Science Institute, Northeastern University , Boston , United States, 5Department of Medical Sciences, University of Torino , 6Clinical nutrition unit, Molinette hospital ,Città della Salute e della Scienza , Torino, Italy

 

Rationale: ESPEN has published 14 clinical nutrition guidelines with over 920 recommendations. Patients with comorbidities require cross-referencing multiple guidelines, a time-consuming and error-prone task.

 

Methods: Thirteen guideline PDFs were processed via a hybrid pipeline: text extraction with a vision-language model, deterministic recommendation block detection, and structured extraction with a large language model (LLM) producing one IF-THEN rule per recommendation, stored in a knowledge graph. Three LLMs of increasing scale (24B, 120B, 671B parameters) were compared. Validation assessed coverage against official recommendation counts and entity-level accuracy on all 43 Cancer Nutrition guideline recommendations, comparing extracted conditions and actions against expert annotations using BERTScore with a biomedical backbone.

 

Results: The knowledge graph contains 1,029 conditions, 707 actions, and 921 recommendations. The 24B model matched the 120B’s extraction yield, while the 671B reasoning model extracted 28% fewer due to structured output failures. Coverage reached 99.9% (921/922 official recommendations). Entity-level validation showed high semantic fidelity for actions (BERTScore-F1=0.899), with 79.1% of recommendations correct or partial match. Condition omission was low: only 6 of 93 expert-annotated conditions were missed. The dominant error was granularity mismatch, where the system decomposed conditions into atomic units while experts used composite phrases, without information loss.

 

Conclusion:  A hybrid pipeline combining deterministic detection with single-recommendation LLM extraction achieves near-complete guideline coverage with high semantic fidelity. A smaller non-reasoning model performs comparably to much larger ones. Prospective clinical validation with domain experts is needed to assess real-world utility for decision support.

Disclosure of Interest: None declared