Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States.
O, I., J, F., M, P.P., K, A., MT, A., IG, S., H, S., FE, A., M, R., & CC, V.E. (2026). Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States.. BMJ open. https://doi.org/10.1136/bmjopen-2025-116133
O I, J F, M PP, K A, MT A, IG S, et al. Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States.. BMJ open. 2026; doi: 10.1136/bmjopen-2025-116133
O I, J F, M PP, et al. Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States.[J]. BMJ open. 2026. DOI: 10.1136/bmjopen-2025-116133.
@article{o2026,
author = {Ibrahim O and Farina J and Pereyra Pietri M and Awad K and Abbas MT and Scalia IG and Sheashaa H and Abdelfattah FE and Razaghi M and Villa Etchegoyen CC},
title = {Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States.},
journal = {BMJ open},
year = {2026},
doi = {10.1136/bmjopen-2025-116133},
note = {PMID: 42595369},
}
TY - JOUR AU - Ibrahim O AU - Farina J AU - Pereyra Pietri M AU - Awad K AU - Abbas MT AU - Scalia IG AU - Sheashaa H AU - Abdelfattah FE AU - Razaghi M AU - Villa Etchegoyen CC TI - Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States. T2 - BMJ open PY - 2026 DO - 10.1136/bmjopen-2025-116133 AN - PMID:42595369 ER -
OBJECTIVE: To compare the diagnostic accuracy of four available automated electronic medical record (EMR) retrieval methods, including a large language model (LLM)-assisted workflow, against manual chart adjudication for identifying cardiovascular events. DESIGN: Retrospective diagnostic accuracy study. SETTING: Three sites within a single US tertiary health system. PARTICIPANTS: Two adult cohorts with previously adjudicated cardiovascular outcomes were included. Cohort 1 included 2258 patients treated with immune checkpoint inhibitors, and Cohort 2 included 1426 patients who underwent transcatheter aortic valve replacement. PRIMARY AND SECONDARY OUTCOME MEASURES: The reference standard was clinician manual chart adjudication. Outcomes included ischaemic stroke or transient ischaemic attack, myocardial infarction (MI), heart failure (HF) exacerbation or hospitalisation and a composite major adverse cardiovascular events (MACE) outcome. Automated retrieval methods included International Classification of Diseases (ICD) codes, primary diagnosis, problem list and a zero-shot LLM workflow. Area under the (receiver operating characteristic) curve (AUC), sensitivity, specificity and net reclassification improvement were assessed. RESULTS: In Cohort 1, the LLM achieved the highest AUC for stroke (0.920; 95% CI 0.881 to 0.958), MI (0.938; 95% CI 0.905 to 0.971) and composite MACE (0.880; 95% CI 0.854 to 0.907), whereas ICD-based retrieval had a higher AUC for HF (0.882; 95% CI 0.845 to 0.918 vs 0.873; 95% CI 0.831 to 0.914). In Cohort 2, the LLM achieved the highest AUC for all evaluated outcomes: stroke (0.915; 95% CI 0.862 to 0.968), MI (0.928; 95% CI 0.839 to 1.000), HF (0.844; 95% CI 0.803 to 0.884) and composite MACE (0.862; 95% CI 0.829 to 0.895). In Cohort 1, differences in AUC between the LLM and ICD methods were not statistically significant across outcomes, whereas in Cohort 2 the LLM showed significantly higher AUC for stroke and composite MACE. CONCLUSION: In this multisite retrospective validation study, the LLM-assisted workflow showed strong but context-dependent performance for identifying cardiovascular events from the EMR. Performance varied by outcome and cohort, and ICD-based retrieval remained competitive for some use cases. These findings support a complementary role for LLM-assisted extraction in retrospective cardiovascular outcomes research.