Digitalni repozitorij raziskovalnih organizacij Slovenije

Izpis gradiva
A+ | A- | Pomoč | SLO | ENG

Naslov:Accuracy and knowledge base evaluation of ChatGPT-4o, Gemini-2.0-Flash, and DeepSeek-V3 in metabolic and bariatric surgery : an expert-rated blinded study
Avtorji:ID Hany, Mohamed (Avtor)
ID Zidan, Mohamed H. (Avtor)
ID Parmar, Chetan (Avtor)
ID Shahabi, Shahab (Avtor)
ID Altabbaa, Hashem (Avtor)
ID Pintar, Tadeja (Sodelavec pri raziskavi)
Datoteke:.pdf PDF - Predstavitvena datoteka, prenos (1,01 MB)
MD5: DB4CDCE74AEE21152963F8D95468AE89
 
URL URL - Izvorni URL, za dostop obiščite https://link.springer.com/article/10.1007/s11695-026-08562-z
 
Jezik:Angleški jezik
Tipologija:1.01 - Izvirni znanstveni članek
Organizacija:Logo UKC LJ - Univerzitetni klinični center Ljubljana
Povzetek:Background: Large language models (LLMs) are increasingly applied in medicine; however, their accuracy in guideline-driven, high-stakes specialties, such as metabolic and bariatric surgery (MBS), remains uncertain. This study evaluates the performance of ChatGPT-4o, Gemini 2.0 Flash, and DeepSeek-V3 in generating guideline-concordant responses to MBS clinical questions. Methods: Thirty standardized, guideline-based MBS questions were presented to each model. Responses were randomized in order, anonymized (blinded as Model A/B/C), and evaluated by 93 MBS experts using a validated 0–3 scale (0 = inaccurate; 3 = fully guideline-concordant). A repeated-measures ANOVA with Bonferroni correction tested model differences; reliability was assessed with Cronbach’s α and intraclass correlation coefficients (ICC). Results: DeepSeek-V3 achieved the highest mean score (2.44 ± 0.40), followed by ChatGPT-4o (1.79 ± 0.46) and Gemini 2.0 Flash (1.63 ± 0.47) (p < 0.001). Fully guideline-concordant ratings (score = 3) were most frequent for DeepSeek (80%) vs. ChatGPT (0%) and Gemini (3.3%). Internal consistency was excellent (α > 0.90), and inter-rater reliability was strong (ICC > 0.88). When mapped against the QUEST evaluation framework, the study addressed Quality and Understanding but did not fully capture Expression, Safety, or Trust dimensions. Conclusions: DeepSeek-V3 outperformed ChatGPT-4o and Gemini 2.0 Flash in generating guideline-concordant responses in MBS. These results highlight the need for ongoing, domain-focused validation before clinical use.
Ključne besede:large language models, metabolic and bariatric surgery, artificial intelligence evaluation, ChatGPT, DeepSeek-V3, Gemini
Status publikacije:Objavljeno
Verzija publikacije:Objavljena publikacija
Leto izida:2026
Št. strani:str. 2160–2171
Številčenje:Vol. 36
PID:20.500.12556/DiRROS-30747 Novo okno
UDK:616-089
ISSN pri članku:1708-0428
DOI:10.1007/s11695-026-08562-z Novo okno
COBISS.SI-ID:273536771 Novo okno
Opomba:Nasl. z nasl. zaslona; Opis vira z dne 30. 3. 2026; Sodelavka pri raziskavi iz Slovenije: Tadeja Pintar;
Datum objave v DiRROS:01.07.2026
Število ogledov:148
Število prenosov:110
Metapodatki:XML DC-XML DC-RDF
:
Kopiraj citat
  
Objavi na:Bookmark and Share


Postavite miškin kazalec na naslov za izpis povzetka. Klik na naslov izpiše podrobnosti ali sproži prenos.

Gradivo je del revije

Naslov:Obesity surgery
Založnik:Springer
ISSN:1708-0428
COBISS.SI-ID:3145492 Novo okno

Licence

Licenca:CC BY 4.0, Creative Commons Priznanje avtorstva 4.0 Mednarodna
Povezava:http://creativecommons.org/licenses/by/4.0/deed.sl
Opis:To je standardna licenca Creative Commons, ki daje uporabnikom največ možnosti za nadaljnjo uporabo dela, pri čemer morajo navesti avtorja.

Nazaj