Original Article
Diagnostic and treatment recommendation performance of multimodal large language models in failed or painful total hip arthroplasty: an expert-rated benchmark study
Abstract
Background: The role of multimodal large language models in orthopaedic decision-making remains insufficiently explored, particularly in revision arthroplasty. This exploratory pilot benchmark study compared the diagnostic and treatment recommendation performance of generative pre-trained transformer (GPT)-4o and GPT-5 with that of human evaluators in cases of failed or painful total hip arthroplasty (THA).
Methods: Twenty cases from a tertiary referral center were retrospectively selected and converted into standardized clinical vignettes, each accompanied by one representative radiological image. The cases were purposively selected, illustrative, with clear dominant failure mechanisms. Cases were independently assessed by GPT-4o, GPT-5, an arthroplasty fellow, an arthroplasty-trained fourth-year orthopaedic resident, and a third-year orthopaedic resident. Responses were rated by two senior reviewers on a 3-point ordinal scale for correctness and completeness. Paired comparisons across raters were performed on the ordinal scores.
Results: GPT-4o achieved fully correct and fully complete diagnoses in 60% and 55% of cases, respectively, compared with 80% and 80% for GPT-5. For treatment recommendations, GPT-4o achieved 60% fully correct and 45% fully complete responses, whereas GPT-5 achieved 85% and 70%, respectively. Overall paired comparisons on the ordinal scale showed significant differences across raters for diagnostic correctness (P<0.001), diagnostic completeness (P<0.001), treatment correctness (P=0.004), and treatment completeness (P=0.001). However, no pairwise difference between GPT-5 and GPT-4o remained significant after correction for multiple testing.
Conclusions: In this small, purposively selected set of illustrative cases with relatively clear dominant failure mechanisms, GPT-5 showed consistent numerical improvement over GPT-4o across all evaluated domains, but neither model matched the highest-performing individual human raters. Multimodal large language models may support structured clinical reasoning in revision THA, but specialist oversight remains necessary, and these findings should not be extrapolated to more ambiguous real-world revision scenarios.

