Back to Search Start Over

The Evaluation of Generative AI Should Include Repetition to Assess Stability.

Authors :
Zhu L
Mou W
Hong C
Yang T
Lai Y
Qi C
Lin A
Zhang J
Luo P
Source :
JMIR mHealth and uHealth [JMIR Mhealth Uhealth] 2024 May 06; Vol. 12, pp. e57978. Date of Electronic Publication: 2024 May 06.
Publication Year :
2024

Abstract

The increasing interest in the potential applications of generative artificial intelligence (AI) models like ChatGPT in health care has prompted numerous studies to explore its performance in various medical contexts. However, evaluating ChatGPT poses unique challenges due to the inherent randomness in its responses. Unlike traditional AI models, ChatGPT generates different responses for the same input, making it imperative to assess its stability through repetition. This commentary highlights the importance of including repetition in the evaluation of ChatGPT to ensure the reliability of conclusions drawn from its performance. Similar to biological experiments, which often require multiple repetitions for validity, we argue that assessing generative AI models like ChatGPT demands a similar approach. Failure to acknowledge the impact of repetition can lead to biased conclusions and undermine the credibility of research findings. We urge researchers to incorporate appropriate repetition in their studies from the outset and transparently report their methods to enhance the robustness and reproducibility of findings in this rapidly evolving field.<br /> (©Lingxuan Zhu, Weiming Mou, Chenglin Hong, Tao Yang, Yancheng Lai, Chang Qi, Anqi Lin, Jian Zhang, Peng Luo. Originally published in JMIR mHealth and uHealth (https://mhealth.jmir.org), 06.05.2024.)

Details

Language :
English
ISSN :
2291-5222
Volume :
12
Database :
MEDLINE
Journal :
JMIR mHealth and uHealth
Publication Type :
Academic Journal
Accession number :
38688841
Full Text :
https://doi.org/10.2196/57978