Cross-Model and Cross-Version Evaluation of Large Language Models in Breast Cancer Tumor Board Decision Support
ONUR ALKAN, Fatih Atalah, Arda Işık
SSRN Electronic Journal · 2026
Background: Multidisciplinary breast tumor boards require integrating clinical guidelines and patient factors under time pressure, making decision support tools attractive. This study compared the performance of ChatGPT-5.2, ChatGPT-4o, and Gemini 3 in responding to simulated breast tumor board cases, evaluating response quality across five domains and readability in Turkish and English.
Methods: Five synthetic breast cancer cases were submitted to three large language models. Three evaluators independently scored outputs on accuracy, completeness, clarity, audience suitability, and risk of misinformation using a 5-point Likert scale. We calculated intraclass correlation coefficients for reliability and used the Friedman test for model comparisons. Readability was assessed using the Ateşman index (Turkish) and Flesch/Gunning Fog indices (English).
Results: Ninety evaluations were included; inter-rater reliability was moderate to good (intraclass correlation coefficient: 0.510-0.789). Gemini scored highest in accuracy (4.00 ± 0.99), ChatGPT-4o in completeness (4.33 ± 0.47) and clarity (4.47 ± 0.23), and ChatGPT-5.2 in specialist suitability (4.40 ± 0.40). Differences were not statistically significant (p > 0.05). English responses scored higher than Turkish ones. Turkish readability varied significantly by model (p = 0.015); English readability was comparable.
Conclusions: The three models performed similarly, each showing relative domain strengths. English responses scored higher clinically, and Turkish readability varied by model. Language-specific validation is essential before clinical integration.