A multidimensional benchmarking framework for large language models in oncologic decision making


Creative Commons License

Halici M., SALTÜRK S., Sayin I., Ertan B., BALCI İ. C., Cepni K., ...More

Scientific Reports, vol.16, no.1, 2026 (SCI-Expanded, Scopus)

  • Publication Type: Article / Article
  • Volume: 16 Issue: 1
  • Publication Date: 2026
  • Doi Number: 10.1038/s41598-026-61195-1
  • Journal Name: Scientific Reports
  • Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, Chemical Abstracts Core, EMBASE, MEDLINE, Directory of Open Access Journals, Zoological Record, Academic Search Ultimate (EBSCO), Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
  • Keywords: Clinical decision support, Composite performance score, Large language models, Non-small cell lung cancer, Oncology
  • Open Archive Collection: AVESIS Open Access Collection
  • Yıldız Technical University Affiliated: Yes

Abstract

Large language models (LLMs) are increasingly explored as clinical decision support tools in oncology; however, reliance on isolated metrics has limited the development of multi-dimensional evaluation frameworks. This comparative observational study utilized five stepwise, clinically realistic non-small cell lung cancer scenarios reflecting real-world diagnostic, therapeutic, and follow-up decision-making. Open-ended clinical questions were answered by three LLMs (Gemini 2.5 Pro, GPT-5, and Claude Opus 4.1) via their official APIs and compared with evidence-based reference answers. Model outputs were evaluated using expert-rated clinical accuracy and explainability, alongside operational metrics including cost, response time, and generative efficiency. All dimensions were integrated into an expert-weighted Composite Performance Score (CPS). Across 30 clinical questions, significant inter-model differences were observed for all metrics (p < 0.001). GPT-5 achieved the highest accuracy, explainability, and generative efficiency, while Gemini 2.5 Pro demonstrated the lowest cost and Opus 4.1 the fastest response times. Integrated analysis yielded the highest CPS for GPT-5, followed by Gemini 2.5 Pro and Opus 4.1 (Kendall’s W = 0.87). A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice. Nevertheless, the use of LLMs in this domain should remain clinician-supervised.