Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases.

Shi P, Li J, Yang Z, Cao Y, Ju L, Bao H, Xu H, Fan Y, Ren T, Xiao Y, Zhu J, Chen J, Dong Y, Hu X, Tong Y, Wang Z, Xiang Y, Zhao J, Zhou J, Temelkuran B, Zhou Y, Jin L, Li C, Wang M, Xiong H, Keane PA, Ge J, Wang N

Open source

DOI
10.1038/s41746-026-02942-6
Published
2026 Jul 2
Container
NPJ digital medicine
Publisher
Not recorded
Open access
unknown

Credibility signals

limited evidence Score 43/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.1038/s41746-026-02942-6,
  title = {Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases.},
  author = {Shi P and Li J and Yang Z and Cao Y and Ju L and Bao H and Xu H and Fan Y and Ren T and Xiao Y and Zhu J and Chen J and Dong Y and Hu X and Tong Y and Wang Z and Xiang Y and Zhao J and Zhou J and Temelkuran B and Zhou Y and Jin L and Li C and Wang M and Xiong H and Keane PA and Ge J and Wang N},
  year = {2026},
  journal = {NPJ digital medicine},
  doi = {10.1038/s41746-026-02942-6},
  url = {https://doi.org/10.1038/s41746-026-02942-6}
}

RIS

TY  - JOUR
TI  - Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases.
AU  - Shi P
AU  - Li J
AU  - Yang Z
AU  - Cao Y
AU  - Ju L
AU  - Bao H
AU  - Xu H
AU  - Fan Y
AU  - Ren T
AU  - Xiao Y
AU  - Zhu J
AU  - Chen J
AU  - Dong Y
AU  - Hu X
AU  - Tong Y
AU  - Wang Z
AU  - Xiang Y
AU  - Zhao J
AU  - Zhou J
AU  - Temelkuran B
AU  - Zhou Y
AU  - Jin L
AU  - Li C
AU  - Wang M
AU  - Xiong H
AU  - Keane PA
AU  - Ge J
AU  - Wang N
PY  - 2026
JO  - NPJ digital medicine
DO  - 10.1038/s41746-026-02942-6
UR  - https://doi.org/10.1038/s41746-026-02942-6
ER  - 

APA

P, S., J, L., Z, Y., Y, C., L, J., H, B., H, X., Y, F., T, R., Y, X., J, Z., J, C., Y, D., X, H., Y, T., Z, W., Y, X., J, Z., J, Z., B, T., Y, Z., L, J., C, L., M, W., H, X., PA, K., J, G., & N, W. (2026). Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases.. NPJ digital medicine. https://doi.org/10.1038/s41746-026-02942-6

Source records