Article
Comparison of ChatGPT-5, Gemini 2.5 Pro, and DeepSeek-V3.1 chatbot responses to dental avulsion injuries
View article page
Özge İlter ErORCID, Özkan AdıgüzelORCID, Makbule TaşyürekORCID, Hatice Ortaç ORCID
Cite
İlter Er Ö, Adıgüzel Ö, Taşyürek M, Ortaç H. Comparison of ChatGPT-5, Gemini 2.5 Pro, and DeepSeek-V3.1 chatbot responses to dental avulsion injuries. H Sci Mon. 2025;3(1):e251201. https://doi.org/10.5577/hsm.e251201
- eISSN
- 3023-6819
- Received
- 2025-12-18
- Published
- 2025-12-25
- Pages
- e251201
Abstract
Aim: The aim of this study was to compare three large language models (LLMs); ChatGPT-5 (OpenAI), Gemini 2.5 Pro (Google), and DeepSeek-V3.1 (Hangzhou DeepSeek AI), in terms of their ability to provide information on questions asked about avulsion, a type of traumatic dental injury.
Methods: In the study, 25 open-ended questions prepared based on the guidelines of the International Association of Dental Traumatology (IADT) were directed to large language models in their "Think / DeepThink / Reasoning" modes. The responses were evaluated by two experienced experts using a five-point Likert scale for accuracy and a three-point Likert scale for completeness (exhaustiveness). Readability was analyzed using the Flesch Reading Ease Score (FRES), Flesch–Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), Simplified Measure of Gobbledygook (SMOG), and Coleman–Liau Index (CLI). Inter-rater reliability was assessed, and ANOVA/Kruskal-Wallis and post-hoc Dunn-Bonferroni tests were used for group comparisons, based on the data's normality.
Results: The results showed that all three models demonstrated high evaluator agreement in terms of both accuracy and completeness. Significant differences were observed between the models in terms of accuracy scores (p=0.016). In subgroup analyses, the ChatGPT-5 model's accuracy score was significantly higher than that of the DeepSeek-V3.1 model (p=0.020). Significant differences were observed between the models in terms of completeness scores (p=0.006). In the subgroup analyses, the completeness scores of the ChatGPT-5 and Gemini 2.5 Pro models were significantly higher than those of the DeepSeek-V3.1 model (p=0.007 and p=0.046, respectively). In terms of readability, the DeepSeek-V3.1 model achieved higher FRES scores and lower FKGL, GFI, and SMOG scores. Overall, ChatGPT-5 demonstrated the highest performance in terms of accuracy and completeness, while DeepSeek-V3.1 achieved the highest readability values.
Conclusion: The results show that ChatGPT-5 is the most reliable model in terms of clinical accuracy and content integrity, whereas DeepSeek-V3.1 produces more linguistically understandable responses. However, no model can completely replace human expertise. These systems should be considered tools to support clinical decision-making processes and should be used under physician supervision, especially in pediatric emergencies.
Keywords
License
© 2025 The Author(s). Published by Zeus Publishing.