Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools.

Reese JT, Chimirri L, Bridges Y, Danis D, Caufield JH, Gargano MA, Kroll C, Schmeder A, Liu F, Wissink K, McMurry JA, Graefe ASL, Niyonkuru E, Korn DR, Casiraghi E, Valentini G, Jacobsen JOB, Haendel M, Smedley D, Mungall CJ, Robinson PN

Open source

DOI
10.1038/s41431-026-02054-5
Published
2026 Apr
Container
European journal of human genetics : EJHG
Publisher
Not recorded
Open access
yes

Credibility signals

limited evidence Score 45/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.1038/s41431-026-02054-5,
  title = {Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools.},
  author = {Reese JT and Chimirri L and Bridges Y and Danis D and Caufield JH and Gargano MA and Kroll C and Schmeder A and Liu F and Wissink K and McMurry JA and Graefe ASL and Niyonkuru E and Korn DR and Casiraghi E and Valentini G and Jacobsen JOB and Haendel M and Smedley D and Mungall CJ and Robinson PN},
  year = {2026},
  journal = {European journal of human genetics : EJHG},
  doi = {10.1038/s41431-026-02054-5},
  url = {https://doi.org/10.1038/s41431-026-02054-5}
}

RIS

TY  - JOUR
TI  - Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools.
AU  - Reese JT
AU  - Chimirri L
AU  - Bridges Y
AU  - Danis D
AU  - Caufield JH
AU  - Gargano MA
AU  - Kroll C
AU  - Schmeder A
AU  - Liu F
AU  - Wissink K
AU  - McMurry JA
AU  - Graefe ASL
AU  - Niyonkuru E
AU  - Korn DR
AU  - Casiraghi E
AU  - Valentini G
AU  - Jacobsen JOB
AU  - Haendel M
AU  - Smedley D
AU  - Mungall CJ
AU  - Robinson PN
PY  - 2026
JO  - European journal of human genetics : EJHG
DO  - 10.1038/s41431-026-02054-5
UR  - https://doi.org/10.1038/s41431-026-02054-5
ER  - 

APA

JT, R., L, C., Y, B., D, D., JH, C., MA, G., C, K., A, S., F, L., K, W., JA, M., ASL, G., E, N., DR, K., E, C., G, V., JOB, J., M, H., D, S., CJ, M., & PN, R. (2026). Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools.. European journal of human genetics : EJHG. https://doi.org/10.1038/s41431-026-02054-5

Source records