| Title: | Accuracy and knowledge base evaluation of ChatGPT-4o, Gemini-2.0-Flash, and DeepSeek-V3 in metabolic and bariatric surgery : an expert-rated blinded study |
|---|
| Authors: | ID Hany, Mohamed (Author) ID Zidan, Mohamed H. (Author) ID Parmar, Chetan (Author) ID Shahabi, Shahab (Author) ID Altabbaa, Hashem (Author) ID Pintar, Tadeja (Research coworker) |
| Files: | PDF - Presentation file, download (1,01 MB) MD5: DB4CDCE74AEE21152963F8D95468AE89
URL - Source URL, visit https://link.springer.com/article/10.1007/s11695-026-08562-z
|
|---|
| Language: | English |
|---|
| Typology: | 1.01 - Original Scientific Article |
|---|
| Organization: | UKC LJ - Ljubljana University Medical Centre
|
|---|
| Abstract: | Background: Large language models (LLMs) are increasingly applied in medicine; however, their accuracy in guideline-driven, high-stakes specialties, such as metabolic and bariatric surgery (MBS), remains uncertain. This study evaluates the performance of ChatGPT-4o, Gemini 2.0 Flash, and DeepSeek-V3 in generating guideline-concordant responses to MBS clinical questions. Methods: Thirty standardized, guideline-based MBS questions were presented to each model. Responses were randomized in order, anonymized (blinded as Model A/B/C), and evaluated by 93 MBS experts using a validated 0–3 scale (0 = inaccurate; 3 = fully guideline-concordant). A repeated-measures ANOVA with Bonferroni correction tested model differences; reliability was assessed with Cronbach’s α and intraclass correlation coefficients (ICC). Results: DeepSeek-V3 achieved the highest mean score (2.44 ± 0.40), followed by ChatGPT-4o (1.79 ± 0.46) and Gemini 2.0 Flash (1.63 ± 0.47) (p < 0.001). Fully guideline-concordant ratings (score = 3) were most frequent for DeepSeek (80%) vs. ChatGPT (0%) and Gemini (3.3%). Internal consistency was excellent (α > 0.90), and inter-rater reliability was strong (ICC > 0.88). When mapped against the QUEST evaluation framework, the study addressed Quality and Understanding but did not fully capture Expression, Safety, or Trust dimensions. Conclusions: DeepSeek-V3 outperformed ChatGPT-4o and Gemini 2.0 Flash in generating guideline-concordant responses in MBS. These results highlight the need for ongoing, domain-focused validation before clinical use. |
|---|
| Keywords: | large language models, metabolic and bariatric surgery, artificial intelligence evaluation, ChatGPT, DeepSeek-V3, Gemini |
|---|
| Publication status: | Published |
|---|
| Publication version: | Version of Record |
|---|
| Year of publishing: | 2026 |
|---|
| Number of pages: | str. 2160–2171 |
|---|
| Numbering: | Vol. 36 |
|---|
| PID: | 20.500.12556/DiRROS-30747  |
|---|
| UDC: | 616-089 |
|---|
| ISSN on article: | 1708-0428 |
|---|
| DOI: | 10.1007/s11695-026-08562-z  |
|---|
| COBISS.SI-ID: | 273536771  |
|---|
| Note: | Nasl. z nasl. zaslona;
Opis vira z dne 30. 3. 2026;
Sodelavka pri raziskavi iz Slovenije: Tadeja Pintar;
|
|---|
| Publication date in DiRROS: | 01.07.2026 |
|---|
| Views: | 145 |
|---|
| Downloads: | 109 |
|---|
| Metadata: |  |
|---|
|
:
|
Copy citation |
|---|
| | | | Share: |  |
|---|
Hover the mouse pointer over a document title to show the abstract or click
on the title to get all document metadata. |