AI Chatbots such as ChatGPT, Microsoft Copilot, and Gemini are increasingly being used by citizens to search for medical information. However, the accuracy and safety of AI-powered medical answers, especially in the field of congenital heart in children, have not been fully evaluated.
Objective: Compare the quality of the answers of three AI platforms (ChatGPT Go, Gemini Pro, and Microsoft Copilot free) to frequently asked questions by parents of children with congenital heart disease.
Research methods: A cross-sectional comparative descriptive study. A set of 86 real-world questions from parents of children with congenital heart disease was included in three AI platforms. Two pediatric cardiologists independently assessed the answers using the Likert scale 1–5 based on five criteria: correctness, completeness, clarity, medical appropriateness, and safety recommendations. Statistical analysis using mean, standard deviation, and McNemar testing.
Results: A total of 258 responses were evaluated. The average score of the chatbots ranges from 4.38 to 4.76 points, with a median of 5 points. When using a ≥3 cut-off point, all of the three chatbots' answers were satisfactory (100%). When applying a higher standard (≥4 points), the rate is from 91.9% to 100%. The difference between the three AI platforms was not statistically significant (p > 0.05). Compatibility between the two doctors assessed at an acceptable level (Kappa >=0.7)
Conclusion: AI chatbots are capable of providing basic advice on congenital heart disease to patients with a relatively high level of reliability, completeness, and safety. There is no meaningful difference between the answers of the 3 AI chatbots on the same question.