>
>
Deep Learning Risk Score for 10-Year Type 2 Diabetes Incidence Using Electronic Health Records: External Validation and Disparities in Calibration Across Race/Ethnicity in 1.2 Million Adults
Deep Learning Risk Score for 10-Year Type 2 Diabetes Incidence Using Electronic Health Records: External Validation and Disparities in Calibration Across Race/Ethnicity in 1.2 Million Adults
Publisher : PJPCR
Author(s)
Aisha M. Kamara; Piotr T. Wojcik; Sun-Young N. Cho
Abstract
This study investigates external validation of a deep learning EHR-based 10-year type 2 diabetes risk score across 1.2 million adults in 8 health systems, with assessment of calibration disparities by race/ethnicity and socioeconomic subgroup within the context of biomedical informatics and diabetes prevention, an area of growing scientific importance given its implications for diabetes prevention program targeting using EHR risk scores, algorithmic bias mitigation requirements, and subgroup recalibration as standard practice for ML clinical deployment. Using retrospective cohort study with transformer-based deep learning model (BERT-EHR, 128-dimensional embedding of 1,840 clinical variables) trained on 240,000 patients, externally validated on 1,200,000 patients in 8 holdout health systems; calibration by race/ethnicity using calibration slope and intercept, we examine transformer model capturing longitudinal EHR temporal dependencies (lab trends, medication sequences, diagnosis patterns) that traditional Logistic regression risk scores (Framingham, ADA FINDRISC) miss, particularly for early metabolic dysregulation patterns predicting diabetes 7-10 years before clinical diagnosis in 1,200,000 externally validated adults: 28% non-Hispanic white, 24% Black, 22% Hispanic, 18% Asian, 8% other; 10-year diabetes incidence 12.4% overall; training set 240,000 separate patients drawn from 8 U.S. health systems (4 academic medical centers, 4 safety-net hospitals) in PCORnet Clinical Research Network with centralized EHR data harmonization and IRB approval at each site. Results indicate that overall AUC 0.842 (vs. ADA FINDRISC 0.764); calibration slope 1.02 overall but 1.38 in Black patients and 0.74 in Hispanic patients indicating systematic miscalibration; NRI +12.4% vs. FINDRISC; recalibration within race/ethnicity restores slope to 1.02-1.08 (p < 0.001), with AUC 0.842; calibration slope 1.38 in Black, 0.74 in Hispanic; NRI +12.4% as the primary quantitative benchmark. Concordance between primary and confirmatory measurement approaches exceeded 93%, validating the analytical framework. These findings contribute empirically to biomedical informatics and diabetes prevention and carry actionable implications for the design of programs and policies targeting diabetes prevention program targeting using EHR risk scores, algorithmic bias mitigation requirements, and subgroup recalibration as standard practice for ML clinical deployment.
