Holistic evaluation of large language models for medical tasks with MedHELM.
- DOI
- 10.1038/s41591-025-04151-2
- Published
- 2026 Mar
- Container
- Nature medicine
- Publisher
- Not recorded
- Open access
- yes
Credibility signals
limited evidence Score 45/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.
Show all credibility signals
- cautionDOI registered: No matching Crossref record was present in this response.
- cautionDOI resolves: No matching Crossref record was present in this response.
- not scoredDirectory of Open Access Journals: No matching DOAJ record was present in this response. No allow-list match; this is not evidence of low credibility.
- not scoredMEDLINE indexed: Not checked or no result supplied; no credibility inference made.
- not scoredOpenAlex core source: Not checked or no result supplied; no credibility inference made.
- not scoredKnown publisher allow-list: Not checked or no result supplied; no credibility inference made.
- not scoredROR affiliation: Not checked or no result supplied; no credibility inference made.
- not scoredRetraction Watch retraction: No retraction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch expression of concern: No expression of concern notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch correction: No correction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch reinstatement: No reinstatement notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- supportingOpen access status: Normalized open-access status: open.
- not scoredPublication license: Not checked or no result supplied; no credibility inference made.
- not scoredPublication version: A publication version was supplied but is not scored.
- cautionMetadata completeness: 5 of 6 scored descriptive metadata groups are present; missing fields increase uncertainty.
Cite this work
BibTeX
@article{allodium:10.1038/s41591-025-04151-2,
title = {Holistic evaluation of large language models for medical tasks with MedHELM.},
author = {Bedi S and Cui H and Fuentes M and Unell A and Wornow M and Banda JM and Kotecha N and Keyes T and Mai Y and Oez M and Qiu H and Jain S and Schettini L and Kashyap M and Fries JA and Swaminathan A and Chung P and Haredasht FN and Lopez I and Aali A and Tse G and Nayak A and Vedak S and Jain SS and Patel B and Fayanju O and Shah S and Goh E and Yao DH and Soetikno B and Reis E and Gatidis S and Divi V and Capasso R and Saralkar R and Chiang CC and Jindal J and Pham T and Ghoddusi F and Lin S and Chiou AS and Hong HJ and Roy M and Gensheimer MF and Patel H and Schulman K and Dash D and Char D and Downing L and Grolleau F and Black K and Mieso B and Zahedivash A and Yim WW and Sharma H and Lee T and Kirsch H and Lee J and Ambers N and Lugtu C and Sharma A and Mawji B and Alekseyev A and Zhou V and Kakkar V and Helzer J and Revri A and Bannett Y and Daneshjou R and Chen J and Alsentzer E and Morse K and Ravi N and Aghaeepour N and Kennedy V and Chaudhari A and Wang T and Koyejo S and Lungren MP and Horvitz E and Liang P and Pfeffer MA and Shah NH},
year = {2026},
journal = {Nature medicine},
doi = {10.1038/s41591-025-04151-2},
url = {https://doi.org/10.1038/s41591-025-04151-2}
}RIS
TY - JOUR TI - Holistic evaluation of large language models for medical tasks with MedHELM. AU - Bedi S AU - Cui H AU - Fuentes M AU - Unell A AU - Wornow M AU - Banda JM AU - Kotecha N AU - Keyes T AU - Mai Y AU - Oez M AU - Qiu H AU - Jain S AU - Schettini L AU - Kashyap M AU - Fries JA AU - Swaminathan A AU - Chung P AU - Haredasht FN AU - Lopez I AU - Aali A AU - Tse G AU - Nayak A AU - Vedak S AU - Jain SS AU - Patel B AU - Fayanju O AU - Shah S AU - Goh E AU - Yao DH AU - Soetikno B AU - Reis E AU - Gatidis S AU - Divi V AU - Capasso R AU - Saralkar R AU - Chiang CC AU - Jindal J AU - Pham T AU - Ghoddusi F AU - Lin S AU - Chiou AS AU - Hong HJ AU - Roy M AU - Gensheimer MF AU - Patel H AU - Schulman K AU - Dash D AU - Char D AU - Downing L AU - Grolleau F AU - Black K AU - Mieso B AU - Zahedivash A AU - Yim WW AU - Sharma H AU - Lee T AU - Kirsch H AU - Lee J AU - Ambers N AU - Lugtu C AU - Sharma A AU - Mawji B AU - Alekseyev A AU - Zhou V AU - Kakkar V AU - Helzer J AU - Revri A AU - Bannett Y AU - Daneshjou R AU - Chen J AU - Alsentzer E AU - Morse K AU - Ravi N AU - Aghaeepour N AU - Kennedy V AU - Chaudhari A AU - Wang T AU - Koyejo S AU - Lungren MP AU - Horvitz E AU - Liang P AU - Pfeffer MA AU - Shah NH PY - 2026 JO - Nature medicine DO - 10.1038/s41591-025-04151-2 UR - https://doi.org/10.1038/s41591-025-04151-2 ER -
APA
S, B., H, C., M, F., A, U., M, W., JM, B., N, K., T, K., Y, M., M, O., H, Q., S, J., L, S., M, K., JA, F., A, S., P, C., FN, H., I, L., A, A., G, T., A, N., S, V., SS, J., B, P., O, F., S, S., E, G., DH, Y., B, S., E, R., S, G., V, D., R, C., R, S., CC, C., J, J., T, P., F, G., S, L., AS, C., HJ, H., M, R., MF, G., H, P., K, S., D, D., D, C., L, D., F, G., K, B., B, M., A, Z., WW, Y., H, S., T, L., H, K., J, L., N, A., C, L., A, S., B, M., A, A., V, Z., V, K., J, H., A, R., Y, B., R, D., J, C., E, A., K, M., N, R., N, A., V, K., A, C., T, W., S, K., MP, L., E, H., P, L., MA, P., & NH, S. (2026). Holistic evaluation of large language models for medical tasks with MedHELM.. Nature medicine. https://doi.org/10.1038/s41591-025-04151-2
Source records
- pubmed · retrieved 2026-09-25T11:43:17.575Z