<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://dirros.openscience.si/IzpisGradiva.php?id=30747"><dc:title>Accuracy and knowledge base evaluation of ChatGPT-4o, Gemini-2.0-Flash, and DeepSeek-V3 in metabolic and bariatric surgery</dc:title><dc:creator>Hany,	Mohamed	(Avtor)
	</dc:creator><dc:creator>Zidan,	Mohamed H.	(Avtor)
	</dc:creator><dc:creator>Parmar,	Chetan	(Avtor)
	</dc:creator><dc:creator>Shahabi,	Shahab	(Avtor)
	</dc:creator><dc:creator>Altabbaa,	Hashem	(Avtor)
	</dc:creator><dc:creator>Pintar,	Tadeja	(Sodelavec pri raziskavi)
	</dc:creator><dc:subject>large language models</dc:subject><dc:subject>metabolic and bariatric surgery</dc:subject><dc:subject>artificial intelligence evaluation</dc:subject><dc:subject>ChatGPT</dc:subject><dc:subject>DeepSeek-V3</dc:subject><dc:subject>Gemini</dc:subject><dc:description>Background: Large language models (LLMs) are increasingly applied in medicine; however, their accuracy in guideline-driven, high-stakes specialties, such as metabolic and bariatric surgery (MBS), remains uncertain. This study evaluates the performance of ChatGPT-4o, Gemini 2.0 Flash, and DeepSeek-V3 in generating guideline-concordant responses to MBS clinical questions. Methods: Thirty standardized, guideline-based MBS questions were presented to each model. Responses were randomized in order, anonymized (blinded as Model A/B/C), and evaluated by 93 MBS experts using a validated 0–3 scale (0 = inaccurate; 3 = fully guideline-concordant). A repeated-measures ANOVA with Bonferroni correction tested model differences; reliability was assessed with Cronbach’s α and intraclass correlation coefficients (ICC). Results: DeepSeek-V3 achieved the highest mean score (2.44 ± 0.40), followed by ChatGPT-4o (1.79 ± 0.46) and Gemini 2.0 Flash (1.63 ± 0.47) (p &lt; 0.001). Fully guideline-concordant ratings (score = 3) were most frequent for DeepSeek (80%) vs. ChatGPT (0%) and Gemini (3.3%). Internal consistency was excellent (α &gt; 0.90), and inter-rater reliability was strong (ICC &gt; 0.88). When mapped against the QUEST evaluation framework, the study addressed Quality and Understanding but did not fully capture Expression, Safety, or Trust dimensions. Conclusions: DeepSeek-V3 outperformed ChatGPT-4o and Gemini 2.0 Flash in generating guideline-concordant responses in MBS. These results highlight the need for ongoing, domain-focused validation before clinical use.</dc:description><dc:date>2026</dc:date><dc:date>2026-07-01 14:02:52</dc:date><dc:type>Neznano</dc:type><dc:identifier>30747</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
