The benchmark illusion: on the curious underestimation of AI and the uncertainty of its creators' gold standards.

Pohlkamp C, Haferlach T

Open source

DOI
10.1016/j.eclinm.2026.104184
Published
2026 Sep
Container
EClinicalMedicine
Publisher
Not recorded
Open access
yes

Credibility signals

limited evidence Score 45/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.1016/j.eclinm.2026.104184,
  title = {The benchmark illusion: on the curious underestimation of AI and the uncertainty of its creators' gold standards.},
  author = {Pohlkamp C and Haferlach T},
  year = {2026},
  journal = {EClinicalMedicine},
  doi = {10.1016/j.eclinm.2026.104184},
  url = {https://doi.org/10.1016/j.eclinm.2026.104184}
}

RIS

TY  - JOUR
TI  - The benchmark illusion: on the curious underestimation of AI and the uncertainty of its creators' gold standards.
AU  - Pohlkamp C
AU  - Haferlach T
PY  - 2026
JO  - EClinicalMedicine
DO  - 10.1016/j.eclinm.2026.104184
UR  - https://doi.org/10.1016/j.eclinm.2026.104184
ER  - 

APA

C, P., & T, H. (2026). The benchmark illusion: on the curious underestimation of AI and the uncertainty of its creators' gold standards.. EClinicalMedicine. https://doi.org/10.1016/j.eclinm.2026.104184

Source records