Back to Search
Start Over
External Validation of Mortality Prediction Models for Critical Illness Reveals Preserved Discrimination but Poor Calibration
- Source :
- Cox , E G M , Wiersema , R , Eck , R J , Kaufmann , T , Granholm , A , Vaara , S T , Møller , M H , Van Bussel , B C T , Snieder , H , Pleijhuis , R G , Van Der Horst , I C C & Keus , F 2023 , ' External Validation of Mortality Prediction Models for Critical Illness Reveals Preserved Discrimination but Poor Calibration ' , Critical Care Medicine , vol. 51 , no. 1 , pp. 80-90 .
- Publication Year :
- 2023
-
Abstract
- OBJECTIVES: In a recent scoping review, we identified 43 mortality prediction models for critically ill patients. We aimed to assess the performances of these models through external validation. DESIGN: Multicenter study. SETTING: External validation of models was performed in the Simple Intensive Care Studies-I (SICS-I) and the Finnish Acute Kidney Injury (FINNAKI) study. PATIENTS: The SICS-I study consisted of 1,075 patients, and the FINNAKI study consisted of 2,901 critically ill patients. MEASUREMENTS AND MAIN RESULTS: For each model, we assessed: 1) the original publications for the data needed for model reconstruction, 2) availability of the variables, 3) model performance in two independent cohorts, and 4) the effects of recalibration on model performance. The models were recalibrated using data of the SICS-I and subsequently validated using data of the FINNAKI study. We evaluated overall model performance using various indexes, including the (scaled) Brier score, discrimination (area under the curve of the receiver operating characteristics), calibration (intercepts and slopes), and decision curves. Eleven models (26%) could be externally validated. The Acute Physiology And Chronic Health Evaluation (APACHE) II, APACHE IV, Simplified Acute Physiology Score (SAPS)-Reduced (SAPS-R)‚ and Simplified Mortality Score for the ICU models showed the best scaled Brier scores of 0.11‚ 0.10‚ 0.10‚ and 0.06‚ respectively. SAPS II, APACHE II, and APACHE IV discriminated best; overall discrimination of models ranged from area under the curve of the receiver operating characteristics of 0.63 (0.61–0.66) to 0.83 (0.81–0.85). We observed poor calibration in most models, which improved to at least moderate after recalibration of intercepts and slopes. The decision curve showed a positive net benefit in the 0–60% threshold probability range for APACHE IV and SAPS-R. CONCLUSIONS: In only 11 out of 43 avai<br />OBJECTIVES: In a recent scoping review, we identified 43 mortality prediction models for critically ill patients. We aimed to assess the performances of these models through external validation. DESIGN: Multicenter study. SETTING: External validation of models was performed in the Simple Intensive Care Studies-I (SICS-I) and the Finnish Acute Kidney Injury (FINNAKI) study. PATIENTS: The SICS-I study consisted of 1,075 patients, and the FINNAKI study consisted of 2,901 critically ill patients. MEASUREMENTS AND MAIN RESULTS: For each model, we assessed: 1) the original publications for the data needed for model reconstruction, 2) availability of the variables, 3) model performance in two independent cohorts, and 4) the effects of recalibration on model performance. The models were recalibrated using data of the SICS-I and subsequently validated using data of the FINNAKI study. We evaluated overall model performance using various indexes, including the (scaled) Brier score, discrimination (area under the curve of the receiver operating characteristics), calibration (intercepts and slopes), and decision curves. Eleven models (26%) could be externally validated. The Acute Physiology And Chronic Health Evaluation (APACHE) II, APACHE IV, Simplified Acute Physiology Score (SAPS)-Reduced (SAPS-R)‚ and Simplified Mortality Score for the ICU models showed the best scaled Brier scores of 0.11‚ 0.10‚ 0.10‚ and 0.06‚ respectively. SAPS II, APACHE II, and APACHE IV discriminated best; overall discrimination of models ranged from area under the curve of the receiver operating characteristics of 0.63 (0.61-0.66) to 0.83 (0.81-0.85). We observed poor calibration in most models, which improved to at least moderate after recalibration of intercepts and slopes. The decision curve showed a positive net benefit in the 0-60% threshold probability range for APACHE IV and SAPS-R. CONCLUSIONS: In only 11 out of 43 available mortality prediction models, the performance could be studied usin
Details
- Database :
- OAIster
- Journal :
- Cox , E G M , Wiersema , R , Eck , R J , Kaufmann , T , Granholm , A , Vaara , S T , Møller , M H , Van Bussel , B C T , Snieder , H , Pleijhuis , R G , Van Der Horst , I C C & Keus , F 2023 , ' External Validation of Mortality Prediction Models for Critical Illness Reveals Preserved Discrimination but Poor Calibration ' , Critical Care Medicine , vol. 51 , no. 1 , pp. 80-90 .
- Notes :
- English
- Publication Type :
- Electronic Resource
- Accession number :
- edsoai.on1439547377
- Document Type :
- Electronic Resource