Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.
- DOI
- 10.1002/ksa.70544
- Published
- 2026-07-23
- Container
- Knee Surg Sports Traumatol Arthrosc
- Publisher
- Not recorded
- Open access
- no
Credibility signals
limited evidence Score 43/100 under policy 1.0.0. This is a metadata assessment, not a judgment of the paper's conclusions.
Show all credibility signals
- cautionDOI registered: No matching Crossref record was present in this response.
- cautionDOI resolves: No matching Crossref record was present in this response.
- not scoredDirectory of Open Access Journals: No matching DOAJ record was present in this response. No allow-list match; this is not evidence of low credibility.
- not scoredMEDLINE indexed: Not checked or no result supplied; no credibility inference made.
- not scoredOpenAlex core source: Not checked or no result supplied; no credibility inference made.
- not scoredKnown publisher allow-list: Not checked or no result supplied; no credibility inference made.
- not scoredROR affiliation: Not checked or no result supplied; no credibility inference made.
- not scoredRetraction Watch retraction: No retraction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch expression of concern: No expression of concern notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch correction: No correction notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredRetraction Watch reinstatement: No reinstatement notice matched this DOI in the deployed snapshot. No matching event found; coverage may be incomplete.
- not scoredOpen access status: Not checked or no result supplied; no credibility inference made.
- not scoredPublication license: Not checked or no result supplied; no credibility inference made.
- not scoredPublication version: A publication version was supplied but is not scored.
- cautionMetadata completeness: 5 of 6 scored descriptive metadata groups are present; missing fields increase uncertainty.
Cite this work
BibTeX
@article{allodium:10.1002/ksa.70544,
title = {Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.},
author = {Pujol O and Ferrer R and Coelho A and Oettl FC and Zsidai B and Leal-Blanquet J and Hirschmann MT and Samuelsson K.},
year = {2026},
journal = {Knee Surg Sports Traumatol Arthrosc},
doi = {10.1002/ksa.70544},
url = {https://doi.org/10.1002/ksa.70544}
}RIS
TY - JOUR TI - Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions. AU - Pujol O AU - Ferrer R AU - Coelho A AU - Oettl FC AU - Zsidai B AU - Leal-Blanquet J AU - Hirschmann MT AU - Samuelsson K. PY - 2026 JO - Knee Surg Sports Traumatol Arthrosc DO - 10.1002/ksa.70544 UR - https://doi.org/10.1002/ksa.70544 ER -
APA
O, P., R, F., A, C., FC, O., B, Z., J, L., MT, H., & K., S. (2026). Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.. Knee Surg Sports Traumatol Arthrosc. https://doi.org/10.1002/ksa.70544
Source records
- europe-pmc · retrieved 2026-09-25T13:28:20.866Z