Muhammad, Shabbir and Altaf, Bandy and Bushra, Mehboob and Usman, Mahboob and Ambreen, Ambreen and Feras, Almarshad and Adam Siti, Khadijah (2026) Quality of AI vs human-generated single best answer questions: a systematic review and meta-analysis. Медичні перспективи = Medicni perspektivi (Medical perspectives), Т. 31 (2). pp. 135-148. ISSN 2307-0404 (print), 2786-4804 (online)
|
Text
135-148-2-2026.pdf Download (987kB) |
Abstract
Single Best Answer Questions (SBAs) are essential and resource-intensive assessment tools in health professions education. Artificial Intelligence (AI), such as large language models (LLMs), can automate the creation of SBAs; however, evidence comparing the quality of AI-generated and human-created items is still dispersed. The purpose of the study is to compare the psychometric quality, measured by difficulty and discrimination indices, of AIgenerated SBAs with those authored by humans in health professions education. The current study followed PRISMA guidelines. The search was conducted on Scopus, PubMed, and Google Scholar. Studies published through April 25th, 2025, and those that directly compared AI- and human-generated SBAs and reported the mean, standard deviation, and sample size for both difficulty and discrimination indices were included. Two reviewers independently extracted the data. Standardized mean differences (SMDs) were calculated and combined using random-effects models (Jamovi MAJOR module, version 2.6.44-06 March 2025). Heterogeneity and publication bias were assessed. Four studies met the inclusion criteria, providing eight comparison outcomes (4 for difficulty, 4 for discrimination). The combined analysis of both outcomes revealed no statistically significant difference, overall (SMD= -0.084, 95% CI: -0.65 to 0.49, p=0.773); however, the heterogeneity was very high (I²=92.7%). Separate analyses revealed that AI-generated questions were significantly easier than human-generated questions (SMD= +0.541, 95% CI: 0.17 to 0.91, p=0.004; I²=62.3%). Conversely, human-authored questions demonstrated significantly higher discrimination indices than AI-generated questions (SMD= -0.701, 95% CI: -1.33 to -0.08, p=0.028; I²=86.2%). No evidence of publication bias was found. AI-generated items tend to be easier, potentially aiding accessibility, whereas human-authored items currently exhibit superior discriminatory power, which is crucial for robust assessment. High heterogeneity underscores context dependency
| Item Type: | Article |
|---|---|
| Additional Information: | doi: 10.26641/2307-0404.2026.2.365915 |
| Uncontrolled Keywords: | Key words: artificial intelligence, item difficulty, item discrimination, large language models, medical education, meta-analysis, psychometrics, single best answer; штучний інтелект, складність завдань, розрізнення завдань, великі мовні моделі, медична освіта, метааналіз, психометрія, єдина найкраща відповідь |
| Subjects: | Medicine. Education |
| Divisions: | University periodicals > Medical perspectives |
| Depositing User: | Ирина Медведева |
| Date Deposited: | 11 Sep 2026 07:07 |
| Last Modified: | 11 Sep 2026 07:07 |
| URI: | http://repo.dma.dp.ua/id/eprint/10276 |
Actions (login required)
![]() |
View Item |


