Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases.
- DOI
- 10.1038/s41746-026-02942-6
- Published
- 2026 Jul 2
- Container
- NPJ digital medicine
- Publisher
- Not recorded
- Open access
- unknown
Credibility signals
limited evidence Score 43/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.
Show all credibility signals
- cautionDOI registered: No matching Crossref record was present in this response.
- cautionDOI resolves: No matching Crossref record was present in this response.
- not scoredDirectory of Open Access Journals: No matching DOAJ record was present in this response. No allow-list match; this is not evidence of low credibility.
- not scoredMEDLINE indexed: Not checked or no result supplied; no credibility inference made.
- not scoredOpenAlex core source: Not checked or no result supplied; no credibility inference made.
- not scoredKnown publisher allow-list: Not checked or no result supplied; no credibility inference made.
- not scoredROR affiliation: Not checked or no result supplied; no credibility inference made.
- not scoredRetraction Watch retraction: No retraction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch expression of concern: No expression of concern notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch correction: No correction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch reinstatement: No reinstatement notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredOpen access status: Not checked or no result supplied; no credibility inference made.
- not scoredPublication license: Not checked or no result supplied; no credibility inference made.
- not scoredPublication version: A publication version was supplied but is not scored.
- cautionMetadata completeness: 5 of 6 scored descriptive metadata groups are present; missing fields increase uncertainty.
Cite this work
BibTeX
@article{allodium:10.1038/s41746-026-02942-6,
title = {Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases.},
author = {Shi P and Li J and Yang Z and Cao Y and Ju L and Bao H and Xu H and Fan Y and Ren T and Xiao Y and Zhu J and Chen J and Dong Y and Hu X and Tong Y and Wang Z and Xiang Y and Zhao J and Zhou J and Temelkuran B and Zhou Y and Jin L and Li C and Wang M and Xiong H and Keane PA and Ge J and Wang N},
year = {2026},
journal = {NPJ digital medicine},
doi = {10.1038/s41746-026-02942-6},
url = {https://doi.org/10.1038/s41746-026-02942-6}
}RIS
TY - JOUR TI - Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases. AU - Shi P AU - Li J AU - Yang Z AU - Cao Y AU - Ju L AU - Bao H AU - Xu H AU - Fan Y AU - Ren T AU - Xiao Y AU - Zhu J AU - Chen J AU - Dong Y AU - Hu X AU - Tong Y AU - Wang Z AU - Xiang Y AU - Zhao J AU - Zhou J AU - Temelkuran B AU - Zhou Y AU - Jin L AU - Li C AU - Wang M AU - Xiong H AU - Keane PA AU - Ge J AU - Wang N PY - 2026 JO - NPJ digital medicine DO - 10.1038/s41746-026-02942-6 UR - https://doi.org/10.1038/s41746-026-02942-6 ER -
APA
P, S., J, L., Z, Y., Y, C., L, J., H, B., H, X., Y, F., T, R., Y, X., J, Z., J, C., Y, D., X, H., Y, T., Z, W., Y, X., J, Z., J, Z., B, T., Y, Z., L, J., C, L., M, W., H, X., PA, K., J, G., & N, W. (2026). Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases.. NPJ digital medicine. https://doi.org/10.1038/s41746-026-02942-6
Source records
- pubmed · retrieved 2026-09-26T05:13:04.052Z