Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.

Pujol O, Ferrer R, Coelho A, Oettl FC, Zsidai B, Leal-Blanquet J, Hirschmann MT, Samuelsson K.

Open source

DOI
10.1002/ksa.70544
Published
2026-07-23
Container
Knee Surg Sports Traumatol Arthrosc
Publisher
Not recorded
Open access
no

Credibility signals

limited evidence Score 43/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.

Show all credibility signals

Cite this work

BibTeX

@article{allodium:10.1002/ksa.70544,
  title = {Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.},
  author = {Pujol O and  Ferrer R and  Coelho A and  Oettl FC and  Zsidai B and  Leal-Blanquet J and  Hirschmann MT and  Samuelsson K.},
  year = {2026},
  journal = {Knee Surg Sports Traumatol Arthrosc},
  doi = {10.1002/ksa.70544},
  url = {https://doi.org/10.1002/ksa.70544}
}

RIS

TY  - JOUR
TI  - Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.
AU  - Pujol O
AU  -  Ferrer R
AU  -  Coelho A
AU  -  Oettl FC
AU  -  Zsidai B
AU  -  Leal-Blanquet J
AU  -  Hirschmann MT
AU  -  Samuelsson K.
PY  - 2026
JO  - Knee Surg Sports Traumatol Arthrosc
DO  - 10.1002/ksa.70544
UR  - https://doi.org/10.1002/ksa.70544
ER  - 

APA

O, P., R, F., A, C., FC, O., B, Z., J, L., MT, H., & K., S. (2026). Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.. Knee Surg Sports Traumatol Arthrosc. https://doi.org/10.1002/ksa.70544

Source records