Please ensure Javascript is enabled for purposes of website accessibility
ORIGINAL ARTICLE
Artificial intelligence in cardiology: evaluating the accuracy of ChatGPT-o1-preview in a medical specialisation exam
 
More details
Hide details
1
Department of Biophysics, Faculty of Medical Sciences in Zabrze, Medical University of Silesia in Katowice, Poland
 
2
Students’ Scientific Association of Computer Analysis and Artificial Intelligence at the Department of Radiology and Nuclear Medicine of the Medical University of Silesia, Katowice, Poland
 
3
Faculty of Automatic Control, Electronics, and Computer Science, Department of Distributed Systems and Informatic Devices, Silesian University of Technology, Gliwice, Poland
 
4
Department of Radiodiagnostics, Interventional Radiology and Nuclear Medicine, Medical University of Silesia, Katowice, Poland
 
These authors had equal contribution to this work
 
 
Submission date: 2025-05-22
 
 
Final revision date: 2025-08-22
 
 
Acceptance date: 2025-08-25
 
 
Online publication date: 2026-09-18
 
 
Corresponding author
Natalia Denisiewicz   

Students' Scientific Association of Computer Analysis and Artificial Intelligence at the Department of Radiology and Nuclear Medicine of the Medical University of Silesia in Katowice
 
 
 
KEYWORDS
TOPICS
ABSTRACT
Introduction:
Recent advancements in artificial intelligence (AI) language models, such as ChatGPT, have sparked interest in their application across various domains, including medicine.

Aim of the research:
This study evaluates the performance of the ChatGPT-o1-preview model on the Polish Specialist Examination (PES) in cardiology.

Material and methods:
A total of 120 single-choice questions from the Spring 2023 PES cardiology exam were analysed. Each question was submitted to the model five times in separate sessions using a standardised prompt. We assessed the accuracy of the responses and the model’s consistency by introducing a “Confidence Index of Language Model” (defined as the ratio of the frequency of the most commonly selected answer to the total number of sessions). Questions were classified according to Bloom’s Taxonomy, and statistical relationships between model performance and question characteristics were analysed.

Results:
ChatGPT-o1-preview provided 97 correct answers (80.83%), well above the passing threshold of 60%. Statistically significant correlations were observed between answer accuracy and difficulty indices (both the Human Difficulty Index and the Confidence Index of Language Model). Questions that were more difficult for human examinees also proved more challenging for the model. No significant differences in model performance were found based on question type (clinical versus other) or content (comprehension and critical thinking versus memory).

Conclusions:
ChatGPT-o1-preview successfully completed a specialised cardiology exam. It is essential to conduct further, more extensive research based on a significantly larger and more diverse set of questions. Only then will it be possible to reliably determine the true potential and limitations of AI in medical applications.
REFERENCES (12)
1.
OpenAI. Introducing ChatGPT. OpenAI (2022). Available at: https://openai.com/index/chatg....
 
2.
Tan S, Xin X, Wu D. ChatGPT in medicine: prospects and challenges: a review article. Int J Surg. 2024; 110: 3701-3706.
 
3.
Liu J, Wang C, Liu S. Utility of ChatGPT in clinical practice. J Med Internet Res. 2023; 25: e48568.
 
4.
Sharma A, Medapalli T, Alexandrou M, Brilakis E, Prasad A. Exploring the role of ChatGPT in cardiology: a systematic review of the current literature. Cureus. 2024; 16: e58936.
 
5.
Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepaño C, Madriaga M, Aggabao R, Diaz-Candido G, Maningo J, Tseng V. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLoS Digital Health. 2023; 2: e0000198.
 
6.
The Lancet Digital Health. ChatGPT: friend or foe? Lancet Digital Health. 2023; 5: e102.
 
7.
Dave T, Athaluri SA, Singh S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intelligence. 2023; 6: 1169595.
 
8.
Anderson LW, Krathwohl DR. A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom’s Taxonomy of Educational Objectives. Longman 2001.
 
9.
Centrum Egzaminów Medycznych, “Specjalizacja – Kardiologia,” CEM (2025). Available at: https://www.cem.edu.pl/aktualn....
 
10.
Huwiler J, Oechslin L, Biaggi P, Tanner FC, Wyss CA. Experimental assessment of the performance of artificial intelligence in solving multiple-choice board exams in cardiology. Swiss Medical Weekly. 2024; 154: 3547.
 
11.
Skalidis I, Cagnina A, Luangphiphat W, Mahendiran T, Muller O, Abbe E, Fournier S. ChatGPT takes on the European Exam in Core Cardiology: an artificial intelligence success story? Eur Heart J Digital Health. 2023; 4: 279-281.
 
12.
Bielówka M, Kufel J, Rojek M, Kaczyńska D, Czogalik Ł, Mitręga A, Bartnikowska W, Kondoł D, Palkij K, Mielcarska S. An investigative analysis – ChatGPT’s capability to excel in the Polish speciality exam in pathology. Pol J Pathol. 2024; 75: 236-240.
 
eISSN:2300-6722
ISSN:1899-1874
Journals System - logo
Scroll to top