ss logo
Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
Filter by Categories
Case Report
Editorial
Original Research Article
Review Article
Technical Note
ss logo
Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
Filter by Categories
Case Report
Editorial
Original Research Article
Review Article
Technical Note
ss logo
Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
Filter by Categories
Case Report
Editorial
Original Research Article
Review Article
Technical Note
View/Download PDF

Translate this page into:

Original Research Article
ARTICLE IN PRESS
doi:
10.25259/JADPR_51_2025

Performance of chatbots for clinical decision-making in orthodontics and dentofacial orthopedics

Orthodontist, Private Practice, Hamedan, Iran,
Université de Médecine et Paramédical de Clermont Auvergne, Clermont-Ferrand, France,
Department of Orthodontics, School of Dentistry, Qom University of Medical Sciences, Qom, Iran,
Department of Pediatric Dentistry, School of Dentistry, Loma Linda University, Loma Linda, California, United States.

*Corresponding author: Parisa Doroudgar, Université de Médecine et Paramédical de Clermont Auvergne, Clermont-Ferrand, France. parisa.doroudgar@yahoo.com

Licence
This is an open-access article distributed under the terms of the Creative Commons Attribution-Non Commercial-Share Alike 4.0 License, which allows others to remix, transform, and build upon the work non-commercially, as long as the author is credited and the new creations are licensed under the identical terms.

How to cite this article: Soheilifar S, Doroudgar P, Balaghi E, Soheilifar S, Rokhshad R. Performance of chatbots for clinical decision-making in orthodontics and dentofacial orthopedics. J Adv Dental Pract Res. doi: 10.25259/JADPR_51_2025

Abstract

Objectives:

Artificial intelligence utilization in medical fields has been expanded in recent years. The aim of this study was to compare the performance of three chatbots in answering clinical questions raised from an orthodontic and orthopedic clinical practice guideline (CPG).

Material and Methods:

Two orthodontists generated 60 questions from the CPG of the American Association of Orthodontics: 20 true-false questions, 20 multichoice questions, and 20 long-answer questions. Five researchers asked each question to ChatGPT-4, Copilot, and Gemini. The answers to true-false and multichoice questions were graded as wrong or right, and accuracy was analyzed using one-way analysis of variance (ANOVA). Answers to long-answer questions were graded from 0 to 10 in four categories (comprehensiveness, scientific accuracy, clarity, and relevance) and analyzed by Kruskal– Wallis test. A p < 0.05 was assumed statistically significant.

Results:

The overall accuracy of the three chatbots in answering true-false and multichoice questions was 0.7–0.9 and 0.5–0.8, respectively. There was no significant difference between the three chatbots in answering true-false questions, while multichoice questions were answered differently (p < 0.035), with Copilot being slightly more accurate. For long answer questions, median grades were 8–9 for comprehensiveness, 10 for scientific accuracy, 10 for clarity, and 10 for relevance. Only for relevance did we find statistically significant differences between the chatbots (p < 0.029).

Conclusion:

All three chatbots have substantial potential in answering orthodontic clinical questions, while performance varied greatly depending on question type.

Keywords

Accuracy
Artificial Intelligence
Chatbots
Clinical Guideline
Orthodontics

INTRODUCTION

The utilization of artificial intelligence (AI) in many different fields, including orthodontics, has grown exponentially over the past decade.[1,2] AI describes a machine’s ability to replicate human cognitive function to perform tasks that are pertaining to human intelligence.[3-5] Large language models (LLMs) employ machine learning and, specifically, deep learning (DL), for automated feature extraction in text and speech without manual intervention. They assist in thinking, interpreting, strategizing, and generating human-like ideas and create the basis for various natural language processing tasks.[4-6] Chatbots are one of the notable applications of LLMs[7] and include models such as OpenAI’s ChatGPT, Google Gemini, and Microsoft Copilot.[5,8-10] Despite the growing advantages of chatbots, there remain challenges associated with chatbots answering clinical practice guideline (CPG) questions. In orthodontics, LLMs have been employed for answering patients’ questions[4,7,10,11] as well as for detecting and classifying malocclusion from orthodontic imagery.[3,6,12] Moreover, case assessment summaries provided by AI have been shown to potentially guide clinicians to acquire rapid and accurate evaluations and reduce human errors.[1,2,13,14]

The aim of this study was to assess the performance of three chatbots – ChatGPT-4, Google Gemini, and Microsoft Copilot – for answering orthodontics and dentofacial orthopedics questions reflecting the CPG catalog of the American Association of Orthodontics (AAO).[15] Our hypothesis was that different chatbots have significantly different accuracy in answering these questions.[1,16]

MATERIAL AND METHODS

Question generation

Two researchers independently generated a list of questions based on the AAO CPG Catalog for Orthodontics and Dentofacial Orthopedics [Supplementary Material]. Both researchers were orthodontic specialists with more than 10 years of experience. Three question types were employed: (1) Twenty multiple-choice questions, (2) Twenty binary “True or False” questions, and (3) Twenty long-answer questions. The number of questions was based on previous studies assessing chatbots in the field of orthodontics and orthognathic surgery.[8,17] Each orthodontist generated 10 questions for each question type (thirty questions by each orthodontist, totally sixty questions). Questions from both researchers were then compared to prevent duplicate questions. Each orthodontist provided an answer to each question and repeated this 1 week later. They further compared their answers. All questions were answered identically by both experts at both time points.

Supplementary Material

Prompting and answering

Five researchers asked each question in a new conversation window with one of the three chatbots: ChatGPT-4 (OpenAI, San Francisco, USA), Gemini (Gemini Trust, New York, USA), and Copilot (Microsoft, Washington, USA) from September 11th to September 17th, 2024. After entering each prompt and generating the answers, the history of search was deleted using the prompt “Delete the chat history.”[18] Each question was asked and then specifically added “be specific and incorporate evidence-based medical guidelines.” The answers generated by the chatbots were independently assessed and compared to the AAO CPG by the two researchers who had generated the questions. The answers to true-false and multiple-choice questions were scored zero (wrong answer) or one (right answer). Long answers were evaluated by scoring them according to a previous study[8] with each answer graded in four categories: (1) Comprehensiveness, (2) scientific accuracy, (3) clarity, and (4) relevance, from minimum (0) to maximum (10) [Table 1]. Any disagreements among the evaluators were resolved through discussion. Each answer was re-assessed by both researchers after 2 weeks to evaluate intra-rater reliability. Intra-rater reliability was 95%.

Table 1: Scoring for long answers.
Criteria 0 1–2 3–4 5–7 8–9 10
Comprehensiveness Response is not at all comprehensive Very few key points are included A basic level of key points is included Many key points are included, but still some key information is missing Almost all key points are included All key points are included
Scientific accuracy Response completely inaccurate Response mostly inaccurate Response presents a basic level of accuracy Response exhibits a good level of scientific accuracy, but some key points are not presented accurately Almost all key points are presented accurately All key points are presented accurately
Clarity Response not clear at all, leaving the assessor very confused Response obscures most key points Response obscures many key points Response is mostly clear, but some key points are not Almost all key points are presented clearly All key points are presented clearly
Relevance Response not relevant at all Many irrelevant points included Some irrelevant points included Most points included are relevant Almost all points included are relevant All points included are relevant

Statistical analysis

All statistical analyses were performed using the Statistical Package for the Social Sciences statistics for Windows, version 27 (IBM, Armonk, USA). We assessed performances (accuracy for true-false and multichoice questions, scoring from 0 to 10 for long answers, see above) for each query (as each chatbot had been inquired by five teams).

Accuracy of true-false and multichoice questions was compared using one-way analysis of variance (ANOVA) after confirming normal distribution. In case of statistically significant difference, pairwise comparison was conducted by Bonferroni post hoc test. Agreement between different rounds of prompts (5 times) was assessed by Fleiss Kappa agreement.

Median and quartiles of grades for each category of long answer questions were calculated. The data were not normally distributed (p< 0.05), and nonparametric Kruskal–Wallis analysis was used to compare answers received. To analyze the overall performance of the chatbots in answering long-answer questions, the data received by the five researchers were combined and the overall median of grades was calculated. In case of statistically significant difference, pairwise comparison was conducted by Mann–Whitney U-test. Agreement between different rounds of prompts was assessed by the Intraclass correlation coefficient (ICC). A p = 0.05 was considered to be statistically significant.

RESULTS

The overall accuracy of chatbots in answering true-false questions was between 0.7 and 0.9 [Table 2], without statistically significant differences between chatbots (p > 0.05/ANOVA). Agreement among different rounds of prompting was fair for ChatGPT-4 and substantial for Gemini and Copilot. The overall agreement of the three chatbots was moderate.

Table 2: Average accuracy of each chatbot and Fleiss Kappa agreement between rounds of prompts in answering true-false and multichoice questions.
Chatbots True-false questions Mean±SD (Minimum–Maximum) Fleiss Kappa agreement between 5 rounds of prompting p-value Multichoice questions Mean±SD (Minimum–Maximum) Fleiss Kappa agreement between 5 rounds of prompting Kappa p-value
ChatGPT4 0.82±0.06 (0.75–0.90) 0.255 p<0.001 0.63±0.11 (0.5–0.75) 0.421 p<0.001
Gemini 0.80±0.61 (0.70–0.85) 0.621 p<0.001 0.63±0.06 (0.55–0.7) 0.356 p<0.001
Copilot 0.79±0.065 (0.70–0.85) 0.638 p<0.001 0.76±0.05 (0.7–0.8) 0.589 p<0.001
Comparison of three chatbots One-way ANOVA p: 0.761 Fleiss Kappa agreement: 0.516 p<0.001 One-way ANOVA p: 0.038* Fleiss Kappa agreement: 0.454 p<0.001

SD: Standard deviation, ANOVA: Analysis of variance. *Statistically significant. Results based on one-way ANOVA and Fleiss Kappa agreement

The overall accuracy of chatbots in answering multichoice questions was between 0.5 and 0.8 [Table 2], with statistically significant differences between the chatbots (p < 0.038). Twoby-two comparison of the groups by the Bonferroni post hoc test did not show significant pairwise differences, though. Agreement among different rounds of prompting was fair for Gemini and moderate for ChatGPT4 and Copilot. The overall agreement of the three chatbots was moderate.

For long-answer questions, grading was performed in four categories. The median grade of comprehensiveness of answers received from chatbots was from 8 to 9 [Table 3]. No statistically significant difference was observed (Kruskal-Wallis/p = 0.934). ICC analysis between different rounds of prompting showed that for the three chatbots, there was an excellent consistency [Table 3].

Table 3: Median, 25 percentile, and 75 percentiles of comprehensiveness, scientific accuracy, clarity, and relevancy of each chatbot (results based on Kruskal-Wallis) and ICC results between different rounds of prompts.
Scoring criteria Median 25 percentiles 75 percentiles p-value Kruskal- Wallis ICC between different rounds of prompts p-value ICC
Comprehensiveness
Chat-GPT4 8 7 10 0.934 0.945 0.001
Gemini 9 7 10 0.941 0.001
Copilot 8 7 10 0.950 0.001
Scientific accuracy
Chat-GPT4 10 10 10 0.127 0.491 0.021
Gemini 10 10 10 0.966 0.001
Copilot 10 10 10 0.898 0.001
Clarity
Chat-GPT4 10 10 10 0.953 0.627 0.001
Gemini 10 10 10 0.577 0.004
Copilot 10 10 10 0.845 0.001
Relevancy
Chat-GPT4 10 10 10 0.029* 0.991 0.001
Gemini 10 10 10 1 0
Copilot 10 10 10 0.961 0.001

ICC: Intraclass correlation coefficient. *Statistically significant

The median grade of scientific accuracy of answers received from chatbots was 10 [Table 3]. No statistically significant difference was observed (p = 0.127). ICC analysis between different rounds of prompting showed a moderate consistency for ChatGPT-4 and excellent consistency for Gemini and Copilot [Table 3].

The median grade of clarity of answers received from chatbots was 10 [Table 3]. No statistically significant difference was observed (p = 0.953). ICC analysis between different rounds of prompting showed a good consistency for ChatGPT-4, moderate consistency for Gemini, and excellent consistency for Copilot [Table 3].

The median relevancy of answers received from chatbots was 10 [Table 3]. A statistically significant difference was observed in the total relevance of answers between the three chatbots (p < 0.029). Two-by-two comparison using Mann–Whitney U-test showed that there was no statistically significant difference between ChatGPT-4 and Gemini and between ChatGPT-4 and Copilot but between Gemini and Copilot (p < 0.013), with Gemini being slightly more relevant. ICC analysis between different rounds of prompting showed that for the three chatbots, there was an excellent consistency.

DISCUSSION

This study evaluated the performance of three advanced AI-driven chatbots (ChatGPT-4, Google Gemini, and Microsoft Copilot) in responding to orthodontics and dentofacial orthopedics questions derived from the AAO CPG. By systematically comparing chatbot responses across true-false, multichoice, and long-answer question formats, we aimed to assess their accuracy, consistency, and clinical applicability. The null hypothesis of this study was that there were no significant differences in the performance of ChatGPT-4, Google Gemini, and Microsoft Copilot in responding to CPG questions for orthodontics and dentofacial orthopedics across different answer question types. Any observed variations in accuracy, comprehensiveness, clarity, and relevance would hence be attributable to random chance or external factors rather than inherent differences in the capabilities of the chatbots. Our results revealed that while all three chatbots performed reasonably well, there were notable differences depending on the question format. We hence reject our hypothesis.

For true-false questions, accuracy scores ranged from 0.7 to 0.9, with ChatGPT-4 consistently achieving a slightly higher mean accuracy than the other two models. However, the differences among the chatbots were not statistically significant (p < 0.761). Variations in performance between individual researchers’ prompts highlight the role of phrasing and context in influencing chatbot responses.

For multichoice questions, Copilot showed a marginally higher accuracy (mean 0.76), with a statistically significant overall difference among chatbots (p < 0.038), although pairwise comparisons did not reach significance. This marginal advantage of Copilot in handling multichoice questions may be related to the design and algorithm of this type of AI. These results also reflect that the handling of structured options might be more refined in certain LLMs, making them more suitable for format-specific clinical assessments. Notably, the lower accuracy in multichoice and binary responses (ranging down to 50%) suggests that some responses likely included clinically incorrect or inappropriate recommendations. Such inaccuracies, if unrecognized, could pose risks in real-world clinical contexts.

In long answer questions, all three chatbots achieved high scores for comprehensiveness, scientific accuracy, and clarity (median scores of 8–10), but relevance showed a statistically significant difference (p < 0.029), with Gemini performing slightly, although (again) not significantly, better than Copilot (p < 0.13). This aligns with previous observations in this study, where variations in chatbot performance appeared more pronounced in certain question formats. While accuracy and comprehensiveness were generally comparable across all three models, this difference in relevance underscores the potential impact of underlying AI model design and training data sources. The superior relevance of Gemini’s responses may be linked to its access to real-time internet data, allowing it to incorporate more current and contextually appropriate information.

Given the probabilistic nature of LLMs, repeating questions multiple times can help mitigate randomness and enhance reliability. Our use of repeated prompting with consistent wording across researchers allowed us to measure intra- and inter-rater consistency. While high consistency was observed in long answer formats (ICC values >0.9), lower consistency in true-false and multichoice questions (Fleiss Kappa values between 0.25 and 0.64) suggests that chatbot outputs can vary with slight changes in context, emphasizing the need for caution in clinical decision-making.

The study’s findings underscore the potential of AI chatbots in supporting orthodontic clinical practice. High scores for scientific accuracy and relevance reflect their ability to align responses with established guidelines. However, the lack of significant performance differences raises questions about whether advancements in generative AI models are translating into meaningful improvements in their practical applications. These findings are consistent with a pilot study assessing chatbot performance in pediatric dentistry.[18] The study found that while chatbots such as ChatGPT-4 demonstrated acceptable consistency (Cronbach’s alpha >0.7) and accuracy, they performed significantly worse than pediatric dentists and general clinicians, who achieved much higher accuracy levels. These results highlight the current limitations of chatbots in substituting human expertise, particularly in diagnostic decision-making, but suggest their value as adjunct tools for educational purposes and patient information.

Similar challenges were observed in a comparative analysis of LLMs (Google Bard, ChatGPT, and Microsoft Bing) in orthodontics, where Bing Chat scored highest in comprehensiveness and relevance, outperforming other models such as ChatGPT-4 and ChatGPT-3.5.[8] However, the study also highlighted the occasional lack of clarity and accuracy in responses across all models, emphasizing the risks associated with relying solely on these tools without critical human oversight.

The evaluation of long-answer questions in this study aligns with findings from research in orthognathic surgery.[17] Chatbots were assessed on their ability to provide high-quality, reliable, and original responses to patient-centered questions. Notably, ChatGPT-4 demonstrated high originality, while other models like OpenEvidence excelled in reliability and adherence to quality standards. Notably, the readability of responses was often challenging, requiring at least a college-level education. These findings further underscore the necessity of tailoring chatbot outputs to match user literacy and comprehension levels, particularly in patient-facing applications.

Another study evaluated the use of GPT-4 for oncology guidelines,[19] also highlighting the capability of AI tools to assist clinicians by synthesizing information from complex guidelines. The studies also converge on the importance of tailoring AI applications to specific clinical scenarios. In both studies, the ability to align responses with established guidelines reflects the significant potential of these tools for evidence-based practice. However, they also underscore that while AI tools can enhance efficiency, their occasional inaccuracies necessitate cautious integration into clinical workflows. Similar observations have been made in radiology, where DL models have shown promise in diagnostic imaging while still requiring human oversight.[20] Notably, systematic and narrative reviews of AI in healthcare emphasize the importance of adapting AI tools to diverse clinical tasks and decision-making environments,[21,22] while ethical implications of AI, including transparency and bias, remain critical across domains.[23]

This study noted statistical differences in performance across question types but did not assess the nature of errors. Incorporating detailed error analysis in future evaluations could yield deeper insights into AI limitations and inform targeted improvements, as has been demonstrated in studies on AI-based clinical decision support systems.[22] Further limitations include the dataset of questions, which – while being representative – was confined to specific CPGs and may not capture the full spectrum of orthodontic practice. In addition, the study evaluated chatbot responses over a fixed time period, which may not account for updates or changes in AI algorithms. Future studies should expand the scope of evaluation to include other AI-driven tools and explore longitudinal performance trends. Investigating the impact of contextual variations in prompts and integrating patient-centric queries may further elucidate the capabilities and limitations of these chatbots in real-world scenarios. Future studies also should explore the impact of augmentative techniques on chatbot performance in orthodontics. By addressing these areas, researchers can further elucidate the capabilities and limitations of AI tools, paving the way for their effective integration into diverse clinical practices.

CONCLUSION

The comparative analysis of ChatGPT-4, Google Gemini, and Microsoft Copilot highlights their substantial potential in assisting orthodontic clinicians with evidence-based queries. While differences in performance between the bots were minimal, the question type clearly affected the accuracy of the AI answers. Continued advancements in AI technology and refinement of chatbot training datasets could further enhance their utility in clinical orthodontics.

Author’s contributions:

SS, PD, EB, SS, and RR: Wrote the main manuscript text. PD: Supervised the manuscript. All authors reviewed the manuscript.

Ethical approval:

Institutional Review Board approval is not required since this is a software generated data and no experimental subject either human or animals have been directly involved.

Declaration of patient consent:

Patient’s consent not required as there are no patients in this study.

Conflicts of interest:

There are no conflicts of interest.

Use of artificial intelligence (AI)-assisted technology for manuscript preparation:

The authors confirm that they have used artificial intelligence (AI) for their comparison and data, but have written the manuscript by themselves without using AI.

Financial support and sponsorship: Nil.

References

  1. , , . Application of artificial intelligence in orthodontics: Current state and future perspectives. Healthcare (Basel). 2023;11:2760.
    [CrossRef] [PubMed] [Google Scholar]
  2. , , , , , , et al. Artificial intelligence and its clinical applications in orthodontics: A systematic review. Diagnostics (Basel). 2023;13:3677.
    [CrossRef] [PubMed] [Google Scholar]
  3. , , . Artificial intelligence in dentistry: Current applications and future perspectives. Quintessence Int. 2020;51:248-57.
    [Google Scholar]
  4. , , , , , , et al. Where is the artificial intelligence applied in dentistry? Systematic review and literature analysis. Healthcare (Basel). 2022;10:1269.
    [CrossRef] [PubMed] [Google Scholar]
  5. , . Can artificial intelligence models serve as patient information consultants in orthodontics? BMC Med Inform Decis Mak. 2024;24:211.
    [CrossRef] [PubMed] [Google Scholar]
  6. , , , , , , et al. Artificial intelligence techniques: Analysis, application, and outcome in dentistry-a systematic review. Biomed Res Int. 2021;2021:9751564.
    [CrossRef] [PubMed] [Google Scholar]
  7. , , , , . Overview of chatbots with special emphasis on artificial intelligence-enabled ChatGPT in medical science. Front Artif Intell. 2023;6:1237704.
    [CrossRef] [PubMed] [Google Scholar]
  8. , , . Evidence-based potential of generative artificial intelligence large language models in orthodontics: A comparative study of ChatGPT, google bard, and microsoft bing. Eur J Orthod. 2025;48:cjae017.
    [CrossRef] [PubMed] [Google Scholar]
  9. , , , , , . The breakthrough of large language models release for medical applications: 1-Year timeline and perspectives. J Med Syst. 2024;48:22.
    [CrossRef] [PubMed] [Google Scholar]
  10. , , , , , , et al. The performance of artificial intelligence models in generating responses to general orthodontic questions: ChatGPT vs Google Bard. Am J Orthod Dentofac Orthop. 2024;165:652-62.
    [CrossRef] [PubMed] [Google Scholar]
  11. , , , . Artificial intelligence in orthodontics: Where are we now? A scoping review. Orthod Craniofac Res. 2021;24:6-15.
    [CrossRef] [PubMed] [Google Scholar]
  12. , , , , . An artificial intelligence system using machine-learning for automatic detection and classification of dental restorations in panoramic radiography. Oral Surg Oral Med, Oral Pathol Oral Radiol. 2020;130:593-602.
    [CrossRef] [PubMed] [Google Scholar]
  13. , . Trends and application of artificial intelligence technology in orthodontic diagnosis and treatment planning-a review. Appl Sci. 2022;12:11864.
    [CrossRef] [Google Scholar]
  14. , , , , , . Retrieval-Augmented Large Language Models for Adolescent Idiopathic Scoliosis Patients in Shared Decision-Making. Proceedings of the 14th ACM International Conference on Bioinformatics In: Computational Biology, and Health Informatics. . p. :1-10.
    [CrossRef] [Google Scholar]
  15. . Clinical Practice Guideline for Orthodontics and Dentofacial Orthopedics. . St. Louis: American Association of Orthodontists; Available from: https://www2.aaoinfo.org/practice-management/cpg [Last accessed on 2025 Nov 07]
    [Google Scholar]
  16. , , , , , . Content analysis of AI-generated (ChatGPT) responses concerning orthodontic clear aligners. Angle Orthod. 2024;94:263-72.
    [CrossRef] [PubMed] [Google Scholar]
  17. , , . A comparative analysis of AI-based chatbots: Assessing data quality in orthognathic surgery related patient information. J Stomatol Oral Maxillofac Surg. 2024;125:101757.
    [CrossRef] [PubMed] [Google Scholar]
  18. , , , , , . Accuracy and consistency of chatbots versus clinicians for answering pediatric dentistry questions: A pilot study. J Dent. 2024;144:104938.
    [CrossRef] [PubMed] [Google Scholar]
  19. , , , , , , et al. GPT-4 for information retrieval and comparison of medical oncology guidelines. NEJM AI. 2024;1:AIcs2300235.
    [CrossRef] [Google Scholar]
  20. , , , , , , et al. Deep learning in radiology. Acad Radiol. 2018;25:1472-80.
    [CrossRef] [PubMed] [Google Scholar]
  21. , . Artificial intelligence and digital pathology: Challenges and opportunities. J Pathol Inform. 2018;9:38.
    [CrossRef] [PubMed] [Google Scholar]
  22. , . Clinical decision support in the era of artificial intelligence. JAMA. 2018;320:2199-200.
    [CrossRef] [PubMed] [Google Scholar]
  23. , , , , , . Operationalising AI ethics: Barriers, enablers and next steps. AI Soc. 2023;38:411-23.
    [CrossRef] [Google Scholar]
Show Sections