A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology (RadM-Bench): Evaluation Development and Validation.

Wu Q, Wu Q, Zhang P, Yi Z, Shen Y, Bai Y, Tan H, Dong P, Xue Z, Roberts N, Wang M

Open source

DOI
10.2196/92183
Published
2026 Aug 7
Container
Journal of medical Internet research
Publisher
Not recorded
Open access
yes

Credibility signals

limited evidence Score 45/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.2196/92183,
  title = {A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology (RadM-Bench): Evaluation Development and Validation.},
  author = {Wu Q and Wu Q and Zhang P and Yi Z and Shen Y and Bai Y and Tan H and Dong P and Xue Z and Roberts N and Wang M},
  year = {2026},
  journal = {Journal of medical Internet research},
  doi = {10.2196/92183},
  url = {https://doi.org/10.2196/92183}
}

RIS

TY  - JOUR
TI  - A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology (RadM-Bench): Evaluation Development and Validation.
AU  - Wu Q
AU  - Wu Q
AU  - Zhang P
AU  - Yi Z
AU  - Shen Y
AU  - Bai Y
AU  - Tan H
AU  - Dong P
AU  - Xue Z
AU  - Roberts N
AU  - Wang M
PY  - 2026
JO  - Journal of medical Internet research
DO  - 10.2196/92183
UR  - https://doi.org/10.2196/92183
ER  - 

APA

Q, W., Q, W., P, Z., Z, Y., Y, S., Y, B., H, T., P, D., Z, X., N, R., & M, W. (2026). A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology (RadM-Bench): Evaluation Development and Validation.. Journal of medical Internet research. https://doi.org/10.2196/92183

Source records