Yunpeng, J., Chao, F., Jianqiang, Z., Wenting, L., & Bin, L. (2026). Accuracy of ChatGPT and DeepSeek in answering clinical questions from the 2025 Society for Cardiovascular Angiography & Interventions/Heart Rhythm Society left atrial appendage occlusion guidelines.. The Journal of international medical research. https://doi.org/10.1177/03000605261438348
Yunpeng J, Chao F, Jianqiang Z, Wenting L, Bin L. Accuracy of ChatGPT and DeepSeek in answering clinical questions from the 2025 Society for Cardiovascular Angiography & Interventions/Heart Rhythm Society left atrial appendage occlusion guidelines.. The Journal of international medical research. 2026; doi: 10.1177/03000605261438348
Yunpeng J, Chao F, Jianqiang Z, et al. Accuracy of ChatGPT and DeepSeek in answering clinical questions from the 2025 Society for Cardiovascular Angiography & Interventions/Heart Rhythm Society left atrial appendage occlusion guidelines.[J]. The Journal of international medical research. 2026. DOI: 10.1177/03000605261438348.
@article{yunpeng2026,
author = {Jin Yunpeng and Feng Chao and Zhao Jianqiang and Lin Wenting and Li Bin},
title = {Accuracy of ChatGPT and DeepSeek in answering clinical questions from the 2025 Society for Cardiovascular Angiography & Interventions/Heart Rhythm Society left atrial appendage occlusion guidelines.},
journal = {The Journal of international medical research},
year = {2026},
doi = {10.1177/03000605261438348},
note = {PMID: 41975575},
}
TY - JOUR AU - Jin Yunpeng AU - Feng Chao AU - Zhao Jianqiang AU - Lin Wenting AU - Li Bin TI - Accuracy of ChatGPT and DeepSeek in answering clinical questions from the 2025 Society for Cardiovascular Angiography & Interventions/Heart Rhythm Society left atrial appendage occlusion guidelines. T2 - The Journal of international medical research PY - 2026 DO - 10.1177/03000605261438348 AN - PMID:41975575 ER -
ObjectiveTo evaluate the accuracy of ChatGPT and DeepSeek in answering guideline-based clinical questions in cardiology.MethodsIn August 2025, responses generated from four large language models to eight clinical questions based on the 2025 Society for Cardiovascular Angiography & Interventions/Heart Rhythm Society guidelines were evaluated. Three cardiologists independently rated accuracy using a six-point Likert scale: (a) completely incorrect; (b) more incorrect than correct; (c) nearly equally correct and incorrect; (d) more correct than incorrect; (e) nearly all correct; and (f) completely correct. Reproducibility (Fleiss' kappa coefficient, five repeated queries) and inter-rater reliability (intraclass correlation coefficient) were assessed.ResultsThe median (interquartile range) accuracy scores were 5.5 (5, 6) for ChatGPT-5, 6 (5, 6) for ChatGPT-4o, and 5 (4, 6) for both DeepSeek-R1 and DeepSeek-V3, with a significant overall difference (p < 0.001). Pairwise comparisons showed significantly higher accuracy for ChatGPT models than for DeepSeek models (all p < 0.001), whereas no significant differences were observed between ChatGPT-5 and ChatGPT-4o (p = 0.518) or between DeepSeek-R1 and DeepSeek-V3 (p = 0.812). Reproducibility (Fleiss' kappa coefficient) was excellent for ChatGPT-5 (0.803) and good for ChatGPT-4o (0.574), DeepSeek-R1 (0.577), and DeepSeek-V3 (0.618). Overall inter-rater reliability was moderate (intraclass correlation coefficient = 0.463).ConclusionsChatGPT and DeepSeek demonstrated high accuracy and reproducibility but moderate inter-rater reliability, necessitating further validation for educational use.