Reply to Hu et al.: Applying different evaluation standards to humans vs. Large Language Models overestimates AI performance.

Leivada E, Günther F, Dentella V

Open source

DOI
10.1073/pnas.2406752121
Published
2024 Sep 3
Container
Proceedings of the National Academy of Sciences of the United States of America
Publisher
Not recorded
Open access
yes

Credibility signals

limited evidence Score 45/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.1073/pnas.2406752121,
  title = {Reply to Hu et al.: Applying different evaluation standards to humans vs. Large Language Models overestimates AI performance.},
  author = {Leivada E and Günther F and Dentella V},
  year = {2024},
  journal = {Proceedings of the National Academy of Sciences of the United States of America},
  doi = {10.1073/pnas.2406752121},
  url = {https://doi.org/10.1073/pnas.2406752121}
}

RIS

TY  - JOUR
TI  - Reply to Hu et al.: Applying different evaluation standards to humans vs. Large Language Models overestimates AI performance.
AU  - Leivada E
AU  - Günther F
AU  - Dentella V
PY  - 2024
JO  - Proceedings of the National Academy of Sciences of the United States of America
DO  - 10.1073/pnas.2406752121
UR  - https://doi.org/10.1073/pnas.2406752121
ER  - 

APA

E, L., F, G., & V, D. (2024). Reply to Hu et al.: Applying different evaluation standards to humans vs. Large Language Models overestimates AI performance.. Proceedings of the National Academy of Sciences of the United States of America. https://doi.org/10.1073/pnas.2406752121

Source records