LB010 - NEXT-GENERATION AUTOMATED INTAKE MONITORING IN A HOSPITAL SETTING: FROM VOLUMETRY TO VISION-LANGUAGE MODELS

Linked sessions

LB010

NEXT-GENERATION AUTOMATED INTAKE MONITORING IN A HOSPITAL SETTING: FROM VOLUMETRY TO VISION-LANGUAGE MODELS

N. Schregenberger1, R. Fernandez2, K. Bogner3, T. Meyer4, C. M. Kiss5,*

1Dietetic Department, University Department of Geriatric Medicine FELIX PLATTER, Basel, Switzerland, 2Nutrition and Dietetics, Universidad Isabel I, Burgos, Spain, 3Health IT Department, 4Acute Geriatric Medicine, 5Clinical Nutrition, University Department of Geriatric Medicine FELIX PLATTER, Basel, Switzerland

 

Rationale: Monitoring nutritional intake in hospitalized patients is essential for identifying individuals at risk of malnutrition. AI‑supported systems now enable continuous and high‑precision intake monitoring, and the emergence of Vision‑Language Models (VLMs) marks a shift toward fully automated, context‑aware assessment.
This study compares a VLM‑based model with the current volume‑based model and validates both against a reference method.

Methods: The Foodscanner (Nutrai GmbH, Switzerland) captures pre‑ and post‑consumption plate images, from which food intake (%) was estimated by a volume-based and VLM-based model. The reference method was the mean visual estimation of three observers using a 11-point scale (0%, 10%, … 100 % eaten), their ratings showed high correlation with the weighing method in prior validation work.
Inter‑rater concordance (IRC) was quantified via the intraclass correlation (ICC). Accuracy relative to reference was evaluated using the ±12.5 percent points (pp) threshold, mean absolute error (MAE), and rate of major clinical errors (>±25 pp). To account for unbalanced intake-category distribution, reliability of aggregated ratings was assessed using an unweighted average across intake levels. Sensitivity for detecting reduced intake (<75%) was also determined.

Results: The dataset comprised 394 plates over eight days, of which 54% were empty. The IRC was excellent (ICC = 0.928). Switching to the VLM-based model raised Foodscanner accuracy from 79% to 86% (unweighted 53% to 68%) and nearly halved the MAE to 4.7 pp (unweighted 9.5 pp). Major clinical errors dropped from 9.8% to 2.8%, and sensitivity for detecting reduced intake reached 90%.

Conclusion: Transitioning to the VLM-based model delivered substantial gains across all metrics, markedly improving accuracy and reliability of automated dietary intake assessment with further advances anticipated as VLMs continue to evolve.

Disclosure of Interest: None declared