Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools.
- DOI
- 10.1038/s41431-026-02054-5
- Published
- 2026 Apr
- Container
- European journal of human genetics : EJHG
- Publisher
- Not recorded
- Open access
- yes
Credibility signals
limited evidence Score 45/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.
Show all credibility signals
- cautionDOI registered: No matching Crossref record was present in this response.
- cautionDOI resolves: No matching Crossref record was present in this response.
- not scoredDirectory of Open Access Journals: No matching DOAJ record was present in this response. No allow-list match; this is not evidence of low credibility.
- not scoredMEDLINE indexed: Not checked or no result supplied; no credibility inference made.
- not scoredOpenAlex core source: Not checked or no result supplied; no credibility inference made.
- not scoredKnown publisher allow-list: Not checked or no result supplied; no credibility inference made.
- not scoredROR affiliation: Not checked or no result supplied; no credibility inference made.
- not scoredRetraction Watch retraction: No retraction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch expression of concern: No expression of concern notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch correction: No correction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch reinstatement: No reinstatement notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- supportingOpen access status: Normalized open-access status: open.
- not scoredPublication license: Not checked or no result supplied; no credibility inference made.
- not scoredPublication version: A publication version was supplied but is not scored.
- cautionMetadata completeness: 5 of 6 scored descriptive metadata groups are present; missing fields increase uncertainty.
Cite this work
BibTeX
@article{allodium:10.1038/s41431-026-02054-5,
title = {Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools.},
author = {Reese JT and Chimirri L and Bridges Y and Danis D and Caufield JH and Gargano MA and Kroll C and Schmeder A and Liu F and Wissink K and McMurry JA and Graefe ASL and Niyonkuru E and Korn DR and Casiraghi E and Valentini G and Jacobsen JOB and Haendel M and Smedley D and Mungall CJ and Robinson PN},
year = {2026},
journal = {European journal of human genetics : EJHG},
doi = {10.1038/s41431-026-02054-5},
url = {https://doi.org/10.1038/s41431-026-02054-5}
}RIS
TY - JOUR TI - Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools. AU - Reese JT AU - Chimirri L AU - Bridges Y AU - Danis D AU - Caufield JH AU - Gargano MA AU - Kroll C AU - Schmeder A AU - Liu F AU - Wissink K AU - McMurry JA AU - Graefe ASL AU - Niyonkuru E AU - Korn DR AU - Casiraghi E AU - Valentini G AU - Jacobsen JOB AU - Haendel M AU - Smedley D AU - Mungall CJ AU - Robinson PN PY - 2026 JO - European journal of human genetics : EJHG DO - 10.1038/s41431-026-02054-5 UR - https://doi.org/10.1038/s41431-026-02054-5 ER -
APA
JT, R., L, C., Y, B., D, D., JH, C., MA, G., C, K., A, S., F, L., K, W., JA, M., ASL, G., E, N., DR, K., E, C., G, V., JOB, J., M, H., D, S., CJ, M., & PN, R. (2026). Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools.. European journal of human genetics : EJHG. https://doi.org/10.1038/s41431-026-02054-5
Source records
- pubmed · retrieved 2026-09-26T06:35:59.200Z