Can large language models generate exam questions comparable to humans? A systematic review and meta-analysis study in medical education

Noor H.S. Alani, Jay Amaranathan, Abdulhalim Dandoush, Sally Rye, Paula Toko King, Marshall H. Chin

Open source

DOI
10.1080/0142159x.2026.2691072
Published
2026-06-21
Container
Medical Teacher
Publisher
Informa UK Limited
Open access
unknown

Credibility signals

uncertain Score 64/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.1080/0142159x.2026.2691072,
  title = {Can large language models generate exam questions comparable to humans? A systematic review and meta-analysis study in medical education},
  author = {Noor H.S. Alani and Jay Amaranathan and Abdulhalim Dandoush and Sally Rye and Paula Toko King and Marshall H. Chin},
  year = {2026},
  journal = {Medical Teacher},
  doi = {10.1080/0142159x.2026.2691072},
  url = {https://doi.org/10.1080/0142159x.2026.2691072}
}

RIS

TY  - JOUR
TI  - Can large language models generate exam questions comparable to humans? A systematic review and meta-analysis study in medical education
AU  - Noor H.S. Alani
AU  - Jay Amaranathan
AU  - Abdulhalim Dandoush
AU  - Sally Rye
AU  - Paula Toko King
AU  - Marshall H. Chin
PY  - 2026
JO  - Medical Teacher
DO  - 10.1080/0142159x.2026.2691072
UR  - https://doi.org/10.1080/0142159x.2026.2691072
ER  - 

APA

Alani, N. H., Amaranathan, J., Dandoush, A., Rye, S., King, P. T., & Chin, M. H. (2026). Can large language models generate exam questions comparable to humans? A systematic review and meta-analysis study in medical education. Medical Teacher. https://doi.org/10.1080/0142159x.2026.2691072

Source records