Articles T67 Số CDD10 - HN Bệnh viện Nhi Đồng 2 - HN Trường Đại học Nguyễn Tất Thành

RELIABILITY OF AI CHATBOTS RESPONSES TO QUESTIONS FROM PARENTS OF CHILDREN WITH CONGENITAL HEART DISEASE

Trinh Huu Tung1,2,3,4, Truong Viet Dung5,6, Nguyen Thi Ngoc Phuong1,3, Phan Thanh Tho1,3, Tran Van Hung1,3, Tran Tuan Kiet1,3
1 Children's Hospital 2
2 Nguyen Tat Thanh University
3 Bệnh viện Nhi Đồng 2
4 Trường Đại Học Nguyễn Tất Thành
5 Tam Anh University
6 Đại học Tâm Anh
DOI: 10.52163/yhc.v67icd10.5869
0 Views
0 Downloads
Abstract

AI Chatbots such as ChatGPT, Microsoft Copilot, and Gemini are increasingly being used by citizens to search for medical information. However, the accuracy and safety of AI-powered medical answers, especially in the field of congenital heart in children, have not been fully evaluated.

Objective: Compare the quality of the answers of three AI platforms (ChatGPT Go, Gemini Pro, and Microsoft Copilot free) to frequently asked questions by parents of children with congenital heart disease.

Research methods: A cross-sectional comparative descriptive study. A set of 86 real-world questions from parents of children with congenital heart disease was included in three AI platforms. Two pediatric cardiologists independently assessed the answers using the Likert scale 1–5 based on five criteria: correctness, completeness, clarity, medical appropriateness, and safety recommendations. Statistical analysis using mean, standard deviation, and McNemar testing.

Results: A total of 258 responses were evaluated. The average score of the chatbots ranges from 4.38 to 4.76 points, with a median of 5 points. When using a ≥3 cut-off point, all of the three chatbots' answers were satisfactory (100%). When applying a higher standard (≥4 points), the rate is from 91.9% to 100%. The difference between the three AI platforms was not statistically significant (p > 0.05). Compatibility between the two doctors assessed at an acceptable level (Kappa >=0.7)

Conclusion: AI chatbots are capable of providing basic advice on congenital heart disease to patients with a relatively high level of reliability, completeness, and safety. There is no meaningful difference between the answers of the 3 AI chatbots on the same question.

References
[1]
Topol EJ (2019). High-performance medicine: the convergence of human and artificial intelligence. Nat Med.;25:44–56. Google Scholar
[2]
Rajkomar A, Dean J, Kohane I (2019). Machine learning in medicine. N Engl J Med.;380:1347–58. Google Scholar
[3]
Kung TH, et al (2023). Performance of ChatGPT on USMLE. PLOS Digit Health.;2:e0000198. Google Scholar
[4]
Nadarzynski T, et al (2019). Acceptability of AI-led chatbots in healthcare. Digit Health.;5:1–12. Google Scholar
[5]
Ayers JW, et al (2023). Comparing physician and AI chatbot responses to patient questions. JAMA Intern Med. 2023;183:589–96. Google Scholar
[6]
Hallucinations in neural machine translation. ACL. 2018. Google Scholar
[7]
Shaw J, et al (2020). Risks of AI in pediatric healthcare. Pediatrics.;145:e20194000. Google Scholar
[8]
McCoy LG, et al (2023). Chatbots and health misinformation. NPJ Digit Med.;6:1–9. Google Scholar
[9]
Singhal K, et al (2023). Large language models encode clinical knowledge. Nature.;620:172–80. Google Scholar
[10]
Gilson A, et al (2023). How does ChatGPT perform on medical questions? Med Educ.;57:565–73. Google Scholar
[11]
Hirosawa T, et al (2023). Evaluation of ChatGPT in clinical decision support. JMIR Med Educ;9:e48023. Google Scholar
[12]
Koo TK, Li MY (2016). A guideline of selecting and reporting ICC. J Chiropr Med.;15:155–63 Google Scholar