Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports: results from the Artificial intelligence to Revolutionize the patient Care pathway in Hip and knEe aRthroplastY (ARCHERY) Project

Luke Farrow; Mingjun Zhong; Lesley Anderson

doi:10.1302/0301-620X.106B7.BJJ-2024-0136

Current issue

Arthroplasty

Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports

results from the Artificial intelligence to Revolutionize the patient Care pathway in Hip and knEe aRthroplastY (ARCHERY) Project

Luke Farrow
Mingjun Zhong
Lesley Anderson

Download PDF

Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports

Farrow L, Zhong M, Anderson L. Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports. Bone Joint J. 2024;106-B(7):688-695. doi:10.1302/0301-620X.106B7.BJJ-2024-0136

Farrow, Luke, et al. “Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports.” The Bone & Joint Journal, vol. 106-B, no. 7, 2024, pp. 688-695., https://doi.org/10.1302/0301-620X.106B7.BJJ-2024-0136

Farrow, L., Zhong, M., & Anderson, L. (2024). Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports. The Bone & Joint Journal, 106-B(7), 688-695. https://doi.org/10.1302/0301-620X.106B7.BJJ-2024-0136

Farrow, L., Zhong, M. and Anderson, L. (2024) “Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports.” The Bone & Joint Journal, 106-B(7), pp. 688-695. Available at: https://doi.org/10.1302/0301-620X.106B7.BJJ-2024-0136

Farrow, Luke, Mingjun Zhong, and Lesley Anderson. “Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports.” The Bone & Joint Journal 106-B, no. 7 (2024): 688-695. https://doi.org/10.1302/0301-620X.106B7.BJJ-2024-0136

Farrow L, Zhong M, Anderson L. Use of natural language processing techniques to predict patient selection for total hip and knee arthroplasty from radiology reports. Bone Joint J. 2024 Jul 1;106-B(7):688-695. https://doi.org/10.1302/0301-620X.106B7.BJJ-2024-0136

Copy to clipboard

Mendeley

BibTeX

EndNote

RIS

Abstract

Aims

To examine whether natural language processing (NLP) using a clinically based large language model (LLM) could be used to predict patient selection for total hip or total knee arthroplasty (THA/TKA) from routinely available free-text radiology reports.

Methods

Data pre-processing and analyses were conducted according to the Artificial intelligence to Revolutionize the patient Care pathway in Hip and knEe aRthroplastY (ARCHERY) project protocol. This included use of de-identified Scottish regional clinical data of patients referred for consideration of THA/TKA, held in a secure data environment designed for artificial intelligence (AI) inference. Only preoperative radiology reports were included. NLP algorithms were based on the freely available GatorTron model, a LLM trained on over 82 billion words of de-identified clinical text. Two inference tasks were performed: assessment after model-fine tuning (50 Epochs and three cycles of k-fold cross validation), and external validation.

Results

For THA, there were 5,558 patient radiology reports included, of which 4,137 were used for model training and testing, and 1,421 for external validation. Following training, model performance demonstrated average (mean across three folds) accuracy, F1 score, and area under the receiver operating curve (AUROC) values of 0.850 (95% confidence interval (CI) 0.833 to 0.867), 0.813 (95% CI 0.785 to 0.841), and 0.847 (95% CI 0.822 to 0.872), respectively. For TKA, 7,457 patient radiology reports were included, with 3,478 used for model training and testing, and 3,152 for external validation. Performance metrics included accuracy, F1 score, and AUROC values of 0.757 (95% CI 0.702 to 0.811), 0.543 (95% CI 0.479 to 0.607), and 0.717 (95% CI 0.657 to 0.778) respectively. There was a notable deterioration in performance on external validation in both cohorts.

Conclusion

The use of routinely available preoperative radiology reports provides promising potential to help screen suitable candidates for THA, but not for TKA. The external validation results demonstrate the importance of further model testing and training when confronted with new clinical cohorts.

Cite this article: Bone Joint J 2024;106-B(7):688–695.

Correspondence should be sent to Luke Farrow. E-mail: Luke.farrow@abdn.ac.uk

For access options please click here

Figure 1

Some description here