Reproducibility And Accuracy of Large Language Model Vision Apis for Carbohydrate Estimation from Food Photographs: A Four-Model Batch Comparison with Implications for Automated Insulin Dosing

Tim Street

SSRN Electronic Journal · 2026

Aims/hypothesis: We aimed to characterize the within-image reproducibility of carbohydrate estimates from four large language model (LLM) vision APIs and to quantify the clinical risk for insulin dosing, stratifying accuracy by reference value quality.

Methods: Thirteen food photographs were each submitted 500-560 times to four LLM vision APIs (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview) using an identical structured prompt adapted from the iAPS automated insulin delivery system (26,904 total queries, temperature 0.01). The primary outcome was within image variation (coefficient of variation [CV], range, distributional normality). Secondary outcomes included accuracy against reference values for nine images, stratified by quality tier (packet label, weighed/measured, portioned, or visual estimate). Clinical risk was translated at an insulin-to-carbohydrate ratio of 1:10.

Results: Median within-image CV was 2.4% (Claude), 8.4% (GPT-5.4), 10.3% (Gemini 3.1 Pro) and 11.0% (Gemini 2.5 Pro), translating to median insulin dosing uncertainties of 0.9, 2.3, 2.9 and 4.7 U respectively. All 52 model--image distributions were non-normal (Shapiro--Wilk p<0.05). On the five strong-reference images (packet label or weighed), Claude achieved the lowest mean absolute error (8.7 g, 95% CI 8.6--8.9; 100% within 20 g) and 0% of queries would have caused an insulin overdose exceeding 2 U, compared with 12--37% for the other models (all pairwise p<10^-20). Claude also achieved the lowest MAE across all nine reference images (11.6 g, 86.4% within 20 g). Conclusions/

interpretation: LLM vision APIs pose two distinct clinical risks: systematic overestimation bias and stochastic within-image variability, both invisible to end users. These findings support multi-query ensemble approaches, uncertainty communication and mandatory user confirmation before LLM-derived values are used for insulin dosing.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.