The performance of ChatGPT and other large language models on multiple-choice questions in biomedical disciplines: A meta-analysis.
- DOI
- 10.1002/ase.70262
- Published
- 2026-05-19
- Container
- Anat Sci Educ
- Publisher
- Not recorded
- Open access
- yes
Credibility signals
limited evidence Score 45/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.
Show all credibility signals
- cautionDOI registered: No matching Crossref record was present in this response.
- cautionDOI resolves: No matching Crossref record was present in this response.
- not scoredDirectory of Open Access Journals: No matching DOAJ record was present in this response. No allow-list match; this is not evidence of low credibility.
- not scoredMEDLINE indexed: Not checked or no result supplied; no credibility inference made.
- not scoredOpenAlex core source: Not checked or no result supplied; no credibility inference made.
- not scoredKnown publisher allow-list: Not checked or no result supplied; no credibility inference made.
- not scoredROR affiliation: Not checked or no result supplied; no credibility inference made.
- not scoredRetraction Watch retraction: No retraction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch expression of concern: No expression of concern notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch correction: No correction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch reinstatement: No reinstatement notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- supportingOpen access status: Normalized open-access status: open.
- not scoredPublication license: Not checked or no result supplied; no credibility inference made.
- not scoredPublication version: A publication version was supplied but is not scored.
- cautionMetadata completeness: 5 of 6 scored descriptive metadata groups are present; missing fields increase uncertainty.
Cite this work
BibTeX
@article{allodium:10.1002/ase.70262,
title = {The performance of ChatGPT and other large language models on multiple-choice questions in biomedical disciplines: A meta-analysis.},
author = {Cheverko CM and Mavrych V and Bolgova O and Mohamed FRR and Westrick J and Juarez L and Rush E and Solka KA and Doubleday AF and Byram JN and Becker R and Gomez V and Ganeng BKA and Hoffman LA and Roach VA and Brown KM and DeVaul N and Garnett CN and Herriott HL and Lufler RS and Mussell JC and Balta JY and Pascoe MA and Middleton JW and Duffy S and Stephens GC and Wilson AB.},
year = {2026},
journal = {Anat Sci Educ},
doi = {10.1002/ase.70262},
url = {https://doi.org/10.1002/ase.70262}
}RIS
TY - JOUR TI - The performance of ChatGPT and other large language models on multiple-choice questions in biomedical disciplines: A meta-analysis. AU - Cheverko CM AU - Mavrych V AU - Bolgova O AU - Mohamed FRR AU - Westrick J AU - Juarez L AU - Rush E AU - Solka KA AU - Doubleday AF AU - Byram JN AU - Becker R AU - Gomez V AU - Ganeng BKA AU - Hoffman LA AU - Roach VA AU - Brown KM AU - DeVaul N AU - Garnett CN AU - Herriott HL AU - Lufler RS AU - Mussell JC AU - Balta JY AU - Pascoe MA AU - Middleton JW AU - Duffy S AU - Stephens GC AU - Wilson AB. PY - 2026 JO - Anat Sci Educ DO - 10.1002/ase.70262 UR - https://doi.org/10.1002/ase.70262 ER -
APA
CM, C., V, M., O, B., FRR, M., J, W., L, J., E, R., KA, S., AF, D., JN, B., R, B., V, G., BKA, G., LA, H., VA, R., KM, B., N, D., CN, G., HL, H., RS, L., JC, M., JY, B., MA, P., JW, M., S, D., GC, S., & AB., W. (2026). The performance of ChatGPT and other large language models on multiple-choice questions in biomedical disciplines: A meta-analysis.. Anat Sci Educ. https://doi.org/10.1002/ase.70262
Source records
- europe-pmc · retrieved 2026-09-25T02:27:22.697Z