Prediction tools are instruments which are commonly used to estimate the prognosis in oncology and facilitate clinical decision-making in a more personalized manner. Their popularity is shown by the increasing numbers of prediction tools, which have been described in the medical literature. Many of these tools have been shown to be useful in the field of soft-tissue sarcoma of the extremities (eSTS). In this annotation, we aim to provide an overview of the available prediction tools for eSTS, provide an approach for clinicians to evaluate the performance and usefulness of the available tools for their own patients, and discuss their possible applications in the management of patients with an eSTS. Cite this article:
Introduction. The ability to walk over various surfaces such as cobblestones, slopes or stairs is a very patient centric and clinically meaningful mobility outcome. Current wearable sensors only measure step counts or walking speed regardless of such context relevant for assessing gait function. This study aims to improve deep learning (DL) models to classify surfaces of walking by altering and comparing model features and sensor configurations. Method. Using a public dataset, signals from 6 IMUs (Movella DOT) worn on various body locations (trunk, wrist, right/left thigh, right/left shank) of 30 subjects walking on 9 surfaces were analyzed (flat ground, ramps (up/down), stairs (up/down), cobblestones (irregular), grass (soft), banked (left/right)). Two variations of a CNN Bi-directional LSTM model, with different Batch Normalization layer placement (beginning vs end) as well as data reduction to individual sensors (versus combined) were explored and
External validation of machine learning predictive models is achieved through evaluation of
Aims. To examine whether natural language processing (NLP) using a clinically based large language model (LLM) could be used to predict patient selection for total hip or total knee arthroplasty (THA/TKA) from routinely available free-text radiology reports. Methods. Data pre-processing and analyses were conducted according to the Artificial intelligence to Revolutionize the patient Care pathway in Hip and knEe aRthroplastY (ARCHERY) project protocol. This included use of de-identified Scottish regional clinical data of patients referred for consideration of THA/TKA, held in a secure data environment designed for artificial intelligence (AI) inference. Only preoperative radiology reports were included. NLP algorithms were based on the freely available GatorTron model, a LLM trained on over 82 billion words of de-identified clinical text. Two inference tasks were performed: assessment after model-fine tuning (50 Epochs and three cycles of k-fold cross validation), and external validation. Results. For THA, there were 5,558 patient radiology reports included, of which 4,137 were used for model training and testing, and 1,421 for external validation. Following training,
Aims. This study aimed to compare the performance of survival prediction models for bone metastases of the extremities (BM-E) with pathological fractures in an Asian cohort, and investigate patient characteristics associated with survival. Methods. This retrospective cohort study included 469 patients, who underwent surgery for BM-E between January 2009 and March 2022 at a tertiary hospital in South Korea. Postoperative survival was calculated using the PATHFx3.0, SPRING13, OPTIModel, SORG, and IOR
Aim. This study aimed to externally validate promising preoperative PJI prediction models in a recent, multinational European cohort. Method. Three preoperative PJI prediction models (by Tan et al., Del Toro et al., and Bülow et al.) which previously demonstrated high levels of accuracy were selected for validation. A multicenter retrospective observational analysis was performed of patients undergoing total hip arthroplasty (THA) and total knee arthroplasty (TKA) between January 2020 and December 2021 and treated at centers in the Netherlands, Portugal, and Spain. Patient characteristics were compared between our cohort and those used to develop the prediction
Introduction. With advances in artificial intelligence, the use of computer-aided detection and diagnosis in clinical imaging is gaining traction. Typically, very large datasets are required to train machine-learning models, potentially limiting use of this technology when only small datasets are available. This study investigated whether pretraining of fracture detection models on large, existing datasets could improve the performance of the model when locating and classifying wrist fractures in a small X-ray image dataset. This concept is termed “transfer learning”. Method. Firstly, three detection models, namely, the faster region-based convolutional neural network (faster R-CNN), you only look once version eight (YOLOv8), and RetinaNet, were pretrained using the large, freely available dataset, common objects in context (COCO) (330000 images). Secondly, these models were pretrained using an open-source wrist X-ray dataset called “Graz Paediatric Wrist Digital X-rays” (GRAZPEDWRI-DX) on a (1) fracture detection dataset (20327 images) and (2) fracture location and classification dataset (14390 images). An orthopaedic surgeon classified the small available dataset of 776 distal radius X-rays (Arbeidsgmeischaft für Osteosynthesefragen Foundation / Orthopaedic Trauma Association; AO/OTA), on which the models were tested. Result. Detection models without pre-training on the large datasets were the least precise when tested on the small distal radius dataset. The model with the best accuracy to detect and classify wrist fractures was the YOLOv8 model pretrained on the GRAZPEDWRI-DX fracture detection dataset (mean average precision at intersection over union of 50=59.7%). This model showed up to 33.6% improved detection precision compared to the same models with no pre-training. Conclusion. Optimisation of machine-learning models can be challenging when only relatively small datasets are available. The findings of this study support the potential of transfer learning from large datasets to improve
Precision health aims to develop personalised and proactive strategies for predicting, preventing, and treating complex diseases such as osteoarthritis (OA). Due to OA heterogeneity, which makes developing effective treatments challenging, identifying patients at risk for accelerated disease progression is essential for efficient clinical trial design and new treatment target discovery and development. To create a reliable and interpretable precision health tool that predicts rapid knee OA progression over a 2-year period from baseline patient characteristics using an advanced automated machine learning (autoML) framework, “Autoprognosis 2.0”. All available 2-year follow-up periods of 600 patients from the FNIH OA Biomarker Consortium were analysed using “Autoprognosis 2.0” in two separate approaches, with distinct definitions of clinical outcomes: multi-class predictions (categorising disease progression into pain and/or radiographic progression) and binary predictions. Models were developed using a training set of 1352 instances and all available variables (including clinical, X-ray, MRI, and biochemical features), and validated through both stratified 10-fold cross-validation and hold-out validation on a testing set of 339 instances.
Introduction. The objective assessment of shoulder function is important for personalized diagnosis, therapies and evidence-based practice but has been limited by specialized equipment and dedicated movement laboratories. Advances in AI-driven computer vision (CV) using consumer RGB cameras (red-blue-green) and open-source CV models offer the potential for routine clinical use. However, key concepts, evidence, and research gaps have not yet been synthesized to drive clinical translation. This scoping review aims to map related literature. Method. Following the JBI Manual for Evidence Synthesis, a scoping review was conducted on PubMed and Scholar using search terms including “shoulder,” “pose estimation,” “camera”, and others. From 146 initial results, 27 papers focusing on clinical applicability and using consumer cameras were included. Analysis employed a Grounded Theory approach guided iterative refinement. Result. Studies primarily used Microsoft Kinect (infrared-based depth sensing, RGB camera; discontinued) or monocular consumer cameras with open-source CV-models, sometimes supplemented by LiDAR (laser-based depth sensing), wearables or markers. Technical validation studies against gold standards were scarce and too inconsistent for comparison. Larger range of motion (RoM) movements were accurately recorded, but smaller movements, rotations and scapula tracking remained challenging. For instance, one larger validation study comparing shoulder angles during arm raises to a marker-based gold-standard reported Pearson's R = 0.98 and a standard error of 2.4deg. OpenPose and Mediapipe were the most used CV-models. Recent efforts try to improve
To examine whether Natural Language Processing (NLP) using a state-of-the-art clinically based Large Language Model (LLM) could predict patient selection for Total Hip Arthroplasty (THA), across a range of routinely available clinical text sources. Data pre-processing and analyses were conducted according to the Ai to Revolutionise the patient Care pathway in Hip and Knee arthroplasty (ARCHERY) project protocol (. https://www.researchprotocols.org/2022/5/e37092/. ). Three types of deidentified Scottish regional clinical free text data were assessed: Referral letters, radiology reports and clinic letters. NLP algorithms were based on the GatorTron model, a Bidirectional Encoder Representations from Transformers (BERT) based LLM trained on 82 billion words of de-identified clinical text. Three specific inference tasks were performed: assessment of the base GatorTron model, assessment after model-fine tuning, and external validation. There were 3911, 1621 and 1503 patient text documents included from the sources of referral letters, radiology reports and clinic letters respectively. All letter sources displayed significant class imbalance, with only 15.8%, 24.9%, and 5.9% of patients linked to the respective text source documentation having undergone surgery. Untrained
Abstract. Introduction. Precision health aims to develop personalised and proactive strategies for predicting, preventing, and treating complex diseases such as osteoarthritis (OA), a degenerative joint disease affecting over 300 million people worldwide. Due to OA heterogeneity, which makes developing effective treatments challenging, identifying patients at risk for accelerated disease progression is essential for efficient clinical trial design and new treatment target discovery and development. Objectives. This study aims to create a trustworthy and interpretable precision health tool that predicts rapid knee OA progression based on baseline patient characteristics using an advanced automated machine learning (autoML) framework, “Autoprognosis 2.0”. Methods. All available 2-year follow-up periods of 600 patients from the FNIH OA Biomarker Consortium were analysed using “Autoprognosis 2.0” in two separate approaches, with distinct definitions of clinical outcomes: multi-class predictions (categorising patients into non-progressors, pain-only progressors, radiographic-only progressors, and both pain and radiographic progressors) and binary predictions (categorising patients into non-progressors and progressors). Models were developed using a training set of 1352 instances and all available variables (including clinical, X-ray, MRI, and biochemical features), and validated through both stratified 10-fold cross-validation and hold-out validation on a testing set of 339 instances.
Identification of patients at risk of not achieving minimally clinically important differences (MCID) in patient reported outcome measures (PROMs) is important to ensure principled and informed pre-operative decision making. Machine learning techniques may enable the generation of a predictive model for attainment of MCID in hip arthroscopy. Aims: 1) to determine whether machine learning techniques could predict which patients will achieve MCID in the iHOT-12 PROM 6 months after arthroscopic management of femoroacetabular impingement (FAI), 2) to determine which factors contribute to their predictive power. Data from the UK Non-Arthroplasty Hip Registry database was utilised. We identified 1917 patients who had undergone hip arthroscopy for FAI with both baseline and 6 month follow up iHOT-12 and baseline EQ-5D scores. We trained three established machine learning algorithms on our dataset to predict an outcome of iHOT-12 MCID improvement at 6 months given baseline characteristics including demographic factors, disease characteristics and PROMs. Performance was assessed using area under the receiver operating characteristic (AUROC) statistics with 5-fold cross validation. The three machine learning algorithms showed quite different performance. The linear logistic regression model achieved AUROC = 0.59, the deep neural network achieved AUROC = 0.82, while a random forest model had the best predictive performance with AUROC 0.87. Of demographic factors, we found that BMI and age were key predictors for this model. We also found that removing all features except baseline responses to the iHOT-12 questionnaire had little effect on performance for the random forest model (AUROC = 0.85). Disease characteristics had little effect on
Introduction. Instability remains a common complication following total hip arthroplasty (THA) and continues to account for the highest percentage of revisions in numerous registries. Many risk factors have been described, yet a patient-specific risk assessment tool remains elusive. The purpose of this study was to apply a machine learning algorithm to develop a patient-specific risk score capable of dynamic adjustment based on operative decisions. Methods. 22,086 THA performed between 1998–2018 were evaluated. 632 THA sustained a postoperative dislocation (2.9%). Patients were robustly characterized based on non-modifiable factors: demographics, THA indication, spinal disease, spine surgery, neurologic disease, connective tissue disease; and modifiable operative decisions: surgical approach, femoral head size, acetabular liner (standard/elevated/constrained/dual-mobility). Models were built with a binary outcome (event/no event) at 1-year and 5-year postoperatively. Inverse Probability Censoring Weighting accounted for censoring bias. An ensemble algorithm was created that included Generalized Linear Model, Generalized Additive Model, Lasso Penalized Regression, Kernel-Based Support Vector Machines, Random Forest and Optimized Gradient Boosting Machine. Convex combination of weights minimized the negative binomial log-likelihood loss function. Ten-fold cross-validation accounted for the rarity of dislocation events. Results. The 1-year model achieved an area under the curve (AUC)=0.63, sensitivity=70%, specificity=50%, positive predictive value (PPV)=3% and negative predictive value (NPV)=99%. The 5-year model achieved an AUC=0.62, sensitivity=69%, specificity=51%, PPV=7% and NPV=97%. All cohort-level accuracy metrics performed better than chance. The two most influential predictors in the model were surgical approach and acetabular liner. Conclusions. This machine learning algorithm demonstrates high sensitivity and NPV, suggesting screening tool utility. The model is strengthened by a multivariable dataset portending differential dislocation risk. Two modifiable variables (approach and acetabular liner) were the most influential in dislocation risk. Calculator utilization in “app” form could enable individualized risk prognostication. Furthermore, algorithm development through machine learning facilitates perpetual
Femoroacetabular Impingement (FAI) syndrome, characterised by abnormal hip contact causing symptoms and osteoarthritis, is measured using the International Hip Outcome Tool (iHOT). This study uses machine learning to predict patient outcomes post-treatment for FAI, focusing on achieving a minimally clinically important difference (MCID) at 52 weeks. A retrospective analysis of 6133 patients from the NAHR who underwent hip arthroscopic treatment for FAI between November 2013 and March 2022 was conducted. MCID was defined as half a standard deviation (13.61) from the mean change in iHOT score at 12 months. SKLearn Maximum Absolute Scaler and Logistic Regression were applied to predict achieving MCID, using baseline and 6-month follow-up data. The
Aims. The purpose of this study was to develop a convolutional neural network (CNN) for fracture detection, classification, and identification of greater tuberosity displacement ≥ 1 cm, neck-shaft angle (NSA) ≤ 100°, shaft translation, and articular fracture involvement, on plain radiographs. Methods. The CNN was trained and tested on radiographs sourced from 11 hospitals in Australia and externally validated on radiographs from the Netherlands. Each radiograph was paired with corresponding CT scans to serve as the reference standard based on dual independent evaluation by trained researchers and attending orthopaedic surgeons. Presence of a fracture, classification (non- to minimally displaced; two-part, multipart, and glenohumeral dislocation), and four characteristics were determined on 2D and 3D CT scans and subsequently allocated to each series of radiographs. Fracture characteristics included greater tuberosity displacement ≥ 1 cm, NSA ≤ 100°, shaft translation (0% to < 75%, 75% to 95%, > 95%), and the extent of articular involvement (0% to < 15%, 15% to 35%, or > 35%). Results. For detection and classification, the algorithm was trained on 1,709 radiographs (n = 803), tested on 567 radiographs (n = 244), and subsequently externally validated on 535 radiographs (n = 227). For characterization, healthy shoulders and glenohumeral dislocation were excluded. The overall accuracy for fracture detection was 94% (area under the receiver operating characteristic curve (AUC) = 0.98) and for classification 78% (AUC 0.68 to 0.93). Accuracy to detect greater tuberosity fracture displacement ≥ 1 cm was 35.0% (AUC 0.57). The CNN did not recognize NSAs ≤ 100° (AUC 0.42), nor fractures with ≥ 75% shaft translation (AUC 0.51 to 0.53), or with ≥ 15% articular involvement (AUC 0.48 to 0.49). For all objectives, the
Advances in cancer therapy have prolonged patient survival even in the presence of disseminated disease and an increasing number of cancer patients are living with metastatic bone disease (MBD). The proximal femur is the most common long bone involved in MBD and pathologic fractures of the femur are associated with significant morbidity, mortality and loss of quality of life (QoL). Successful prophylactic surgery for an impending fracture of the proximal femur has been shown in multiple cohort studies to result in longer survival, preserved mobility, lower transfusion rates and shorter post-operative hospital stays. However, there is currently no optimal method to predict a pathologic fracture. The most well-known tool is Mirel's criteria, established in 1989 and is limited from guiding clinical practice due to poor specificity and sensitivity. The ideal clinical decision support tool will be of the highest sensitivity and specificity, non-invasive, generalizable to all patients, and not a burden on hospital resources or the patient's time. Our research uses novel machine learning techniques to develop a model to fill this considerable gap in the treatment pathway of MBD of the femur. The goal of our study is to train a convolutional neural network (CNN) to predict fracture risk when metastatic bone disease is present in the proximal femur. Our fracture risk prediction tool was developed by analysis of prospectively collected data of consecutive MBD patients presenting from 2009–2016. Patients with primary bone tumors, pathologic fractures at initial presentation, and hematologic malignancies were excluded. A total of 546 patients comprising 114 pathologic fractures were included. Every patient had at least one Anterior-Posterior X-ray and clinical data including patient demographics, Mirel's criteria, tumor biology, all previous radiation and chemotherapy received, multiple pain and function scores, medications and time to fracture or time to death. We have trained a convolutional neural network (CNN) with AP X-ray images of 546 patients with metastatic bone disease of the proximal femur. The digital X-ray data is converted into a matrix representing the color information at each pixel. Our CNN contains five convolutional layers, a fully connected layers of 512 units and a final output layer. As the information passes through successive levels of the network, higher level features are abstracted from the data. The model converges on two fully connected deep neural network layers that output the risk of fracture. This prediction is compared to the true outcome, and any errors are back-propagated through the network to accordingly adjust the weights between connections, until overall prediction accuracy is optimized. Methods to improve learning included using stochastic gradient descent with a learning rate of 0.01 and a momentum rate of 0.9. We used average classification accuracy and the average F1 score across five test sets to measure
Technology within medicine has great potential to bring about more accessible, efficient, and a higher quality delivery of care. Paediatric supracondylar fractures are the most common elbow fracture in children and at our institution often have high rates of unnecessary long term clinical follow-up, leading to an inefficient use of healthcare and patient resources. This study aims to evaluate patient and clinical factors that significantly predict necessity for further clinical visits following closed reduction and percutaneous pinning. A total of 246 children who underwent closed reduction and percutaneous pinning following supracondylar humerus fractures were prospectively enrolled over a two year period. Patient demographics, perioperative course, goniometric measurements, functional outcome measures, clinical assessment and decision making for further follow up were assessed. Categorical and continuous variables were analyzed and screened for significance via bivariate regression. Significant covariates were used to develop a predictive model through multivariate logistical regression. A probability cut-off was determined on the Receiver Operator Characteristic (ROC) curve using the Youden index to maximize sensitivity and specificity. The regression
Advances in cancer therapy have prolonged cancer patient survival even in the presence of disseminated disease and an increasing number of cancer patients are living with metastatic bone disease (MBD). The proximal femur is the most common long bone involved in MBD and pathologic fractures of the femur are associated with significant morbidity, mortality and loss of quality of life (QoL). Successful prophylactic surgery for an impending fracture of the proximal femur has been shown in multiple cohort studies to result in patients more likely to walk after surgery, longer survival, lower transfusion rates and shorter post-operative hospital stays. However, there is currently no optimal method to predict a pathologic fracture. The most well-known tool is Mirel's criteria, established in 1989 and is limited from guiding clinical practice due to poor specificity and sensitivity. The goal of our study is to train a convolutional neural network (CNN) to predict fracture risk when metastatic bone disease is present in the proximal femur. Our fracture risk prediction tool was developed by analysis of prospectively collected data for MBD patients (2009–2016) in order to determine which features are most commonly associated with fracture. Patients with primary bone tumors, pathologic fractures at initial presentation, and hematologic malignancies were excluded. A total of 1146 patients comprising 224 pathologic fractures were included. Every patient had at least one Anterior-Posterior X-ray. The clinical data includes patient demographics, tumor biology, all previous radiation and chemotherapy received, multiple pain and function scores, medications and time to fracture or time to death. Each of Mirel's criteria has been further subdivided and recorded for each lesion. We have trained a convolutional neural network (CNN) with X-ray images of 1146 patients with metastatic bone disease of the proximal femur. The digital X-ray data is converted into a matrix representing the color information at each pixel. Our CNN contains five convolutional layers, a fully connected layers of 512 units and a final output layer. As the information passes through successive levels of the network, higher level features are abstracted from the data. This model converges on two fully connected deep neural network layers that output the fracture risk. This prediction is compared to the true outcome, and any errors are back-propagated through the network to accordingly adjust the weights between connections. Methods to improve learning included using stochastic gradient descent with a learning rate of 0.01 and a momentum rate of 0.9. We used average classification accuracy and the average F1 score across test sets to measure
This is quite an innovative study that should lead to a multicentre validation trial. We have developed an FDG-PET/MRI texture-based model for the prediction of lung metastases (LM) in newly diagnosed patients with soft-tissue sarcomas (STSs) using retrospective analysis. In this work, we assess the
Background. The advent of value-based conscientiousness and rapid-recovery discharge pathways presents surgeons, hospitals, and payers with the challenge of providing the same total hip arthroplasty episode of care in the safest and most economic fashion for the same fee, despite patient differences. Various predictive analytic techniques have been applied to medical risk models, such as sepsis risk scores, but none have been applied or validated to the elective primary total hip arthroplasty (THA) setting for key payment-based metrics. The objective of this study was to develop and validate a predictive machine learning model using preoperative patient demographics for length of stay (LOS) after primary THA as the first step in identifying a patient-specific payment model (PSPM). Methods. Using 229,945 patients undergoing primary THA for osteoarthritis from an administrative database between 2009– 16, we created a naïve Bayesian model to forecast LOS after primary THA using a 3:2 split in which 60% of the available patient data “built” the algorithm and the remaining 40% of patients were used for “testing.” This process was iterated five times for algorithm refinement, and