Back to Search Start Over

Influence of Model Evolution and System Roles on ChatGPT’s Performance in Chinese Medical Licensing Exams: Comparative Study

Authors :
Shuai Ming
Qingge Guo
Wenjun Cheng
Bo Lei
Source :
JMIR Medical Education, Vol 10, Pp e52784-e52784 (2024)
Publication Year :
2024
Publisher :
JMIR Publications, 2024.

Abstract

Abstract BackgroundWith the increasing application of large language models like ChatGPT in various industries, its potential in the medical domain, especially in standardized examinations, has become a focal point of research. ObjectiveThe aim of this study is to assess the clinical performance of ChatGPT, focusing on its accuracy and reliability in the Chinese National Medical Licensing Examination (CNMLE). MethodsThe CNMLE 2022 question set, consisting of 500 single-answer multiple choices questions, were reclassified into 15 medical subspecialties. Each question was tested 8 to 12 times in Chinese on the OpenAI platform from April 24 to May 15, 2023. Three key factors were considered: the version of GPT-3.5 and 4.0, the prompt’s designation of system roles tailored to medical subspecialties, and repetition for coherence. A passing accuracy threshold was established as 60%. The χ2 ResultsGPT-4.0 achieved a passing accuracy of 72.7%, which was significantly higher than that of GPT-3.5 (54%; PPPP ConclusionsGPT-4.0 passed the CNMLE and outperformed GPT-3.5 in key areas such as accuracy, consistency, and medical subspecialty expertise. Adding a system role insignificantly enhanced the model’s reliability and answer coherence. GPT-4.0 showed promising potential in medical education and clinical practice, meriting further study.

Details

Language :
English
ISSN :
23693762
Volume :
10
Database :
Directory of Open Access Journals
Journal :
JMIR Medical Education
Publication Type :
Academic Journal
Accession number :
edsdoj.7fbe0f24610842deab04cf87eeb6a15d
Document Type :
article
Full Text :
https://doi.org/10.2196/52784