Accessibility settings

Published on in Vol 12 (2026)

Preprints (earlier versions) of this paper are available at https://preprints.jmir.org/preprint/91510, first published .
Doctor explaining predicted survival curve and risk drivers to patient on tablet.

Explainable Machine Learning–Based Prediction of Progression-Free Survival in Prostate Cancer: Retrospective Cohort Study

Explainable Machine Learning–Based Prediction of Progression-Free Survival in Prostate Cancer: Retrospective Cohort Study

Original Paper

1Pengiran Anak Puteri Rashidah Sa'adatul Bolkiah Institute of Health Sciences, Universiti Brunei Darussalam, Bandar Seri Begawan, Brunei-Muara District, Brunei Darussalam

2School of Digital Science, Universiti Brunei Darussalam, Bandar Seri Begawan, Brunei Darussalam

3Faculty of Engineering and Computing, Atlantic Technological University, Donegal, Ireland

4The Brunei Cancer Centre, Jerudong Park Medical Centre, Bandar Seri Begawan, Brunei-Muara District, Brunei Darussalam

5Saw Swee Hock School of Public Health, National University of Singapore, Singapore, Singapore

Corresponding Author:

Hein Minn Tun, DipFM, MBBS, MPH

Pengiran Anak Puteri Rashidah Sa'adatul Bolkiah Institute of Health Sciences

Universiti Brunei Darussalam

Jalan Tungku Link

Gadong

Bandar Seri Begawan, Brunei-Muara District, BE1410

Brunei Darussalam

Phone: 673 7428942

Email: 23h8750@ubd.edu.bn


Background: Progression-free survival (PFS) is a critical end point in oncology, yet real-world applications of individualized, explainable machine learning (ML) predictions remain limited.

Objective: This study aimed to develop and validate explainable ML models to predict PFS using retrospective data from a national prostate cancer cohort in Brunei Darussalam.

Methods: We analyzed a retrospective cohort of 212 patients (478 longitudinal observations) treated at the Brunei Cancer Centre (January 2018 to December 2024). Clinical, laboratory, and treatment data were harmonized, with missing values imputed via extremely randomized trees. Longitudinal patterns were captured using a recurrent autoencoder to generate latent representations. We compared 4 modeling approaches: Cox proportional hazards, random survival forest (RSF), gradient boosting survival (GBS), and deep neural network survival models. Performance was evaluated using time-dependent area under the receiver operating characteristic curve (AUC), Harrell C-index, and integrated Brier score (IBS), with Shapley additive explanations (SHAP) used for interpretability.

Results: RSF demonstrated improved discriminative performance and balanced calibration, achieving a C-index of 0.906 and AUCs of 0.941 and 0.917 at 4 and 5 years (IBS=0.0698). In contrast, the traditional CPH model performed poorly (C-index 0.531 and AUC 0.706 at 4 years and 0.833 at 5 years). Deep survival (AUCs of 0.941 at 4 years and 0.917 at 5 years, C-index 0.719, and IBS=0.0887) and GBS (AUCs of 0.765 at 4 years and 0.833 at 5 years, C-index 0.844, and IBS=0.0590) models showed moderate performance. SHAP analysis identified sodium, alanine aminotransferase, mean corpuscular hemoglobin, platelet count, and specific treatment categories as key drivers of increased progression risk.

Conclusions: Tree-based ensemble approaches, particularly RSF integrated with SHAP, offer high accuracy for personalized risk stratification in prostate cancer. These findings highlight the potential of explainable ML to enhance clinical decision-making. However, external validation in a larger multi-institutional, multiomics dataset is required before routine clinical implementation.

JMIR Cancer 2026;12:e91510

doi:10.2196/91510

Keywords



Prostate cancer is the fourth most common cancer in men worldwide, and the incidence of prostate cancer is increasing in Africa, Asia, and Latin America and the Caribbean, raising public health concerns worldwide [1,2]. Global and regional analyses demonstrate that the highest incidence rates are reported in northern Europe, North America, and Australia, largely due to the widespread use of prostate-specific antigen (PSA)–based early detection systems, while mortality is higher in Africa, Asia, and Latin America, indicating the global disparity based on screening and diagnostic capacity [2-5]. In Brunei Darussalam specifically, the national and international cancer registry data demonstrate the increasing prostate cancer deaths in recent decades [6,7]. These epidemiological trends in prostate cancer highlight the need for locally adapted prognostic and management tools that can improve the prostate cancer care continuum in Brunei Darussalam.

Progression-free survival (PFS) and disease-free survival are widely used in prostate cancer clinical trials and observational studies to quantify the duration during which a patient remains free from radiographic or clinical disease progression after treatment [8-10]. PFS often serves as a pragmatic outcome in trials of systemic and targeted therapies, as it can be observed earlier than overall survival and may reflect the treatment effect on tumor control. According to PCWG (Prostate Cancer Working Group) criteria, PSA progression is defined as a confirmed ≥25% and ≥2 ng/mL increase from the PSA nadir (or from baseline after ≥12 weeks if no nadir is achieved), with confirmation to exclude PSA flare, while radiographic or clinical progression, if occurring earlier, takes precedence in determining PFS [11]. Additionally, the understanding of PFS extension for prostate cancer is viewed differently between patients and clinicians. Patients generally require a longer PFS benefit to consider treatment worthwhile and tend to prioritize PSA or radiological responses despite potential side effects, whereas clinicians place greater emphasis on balancing treatment-related toxicity with PFS and may recommend therapy even when the expected PFS gain is modest [12].

Over the past decade, machine learning (ML) methods, including gradient-boosted trees, random survival forests, and deep learning architectures, have been increasingly applied to survival prediction and progression modeling across many cancer types, including lung and prostate cancer [13-17]. Furthermore, multimodal data integration of clinical variables, laboratory tests, imaging and radiomics features, and molecular data to produce individualized risk estimates for progression or death often outperforms traditional statistical models in complex, high-dimensional datasets [17,18]. Despite these advances, most ML and deep learning models for prostate cancer survival prediction rely on cross-sectional or clinical trial datasets with short follow-up periods, limiting their applicability in real-world practice [19-21]. Additionally, explainability methods such as Shapley additive explanations (SHAP) play an important role in the clinical applicability of these models by providing attributions for individual predictions [18,22,23].

However, few studies have explored the sequential analysis of progression free survival with explainable methods to provide locally calibrated, interpretable predictions for progression-free outcomes. Additionally, no previous study in Brunei Darussalam has been conducted to develop interpretable ML survival models for prostate cancer progression. To fill these gaps, our study aimed to develop explainable ML models that predict PFS using retrospective clinical, laboratory, and treatment information from a national cohort in Brunei Darussalam. By developing locally tailored, AI-enabled PFS models, this study contributes to the growing body of evidence supporting the use of ML and deep learning in precision oncology, while also advancing AI-based clinical decision support systems to assist oncologists in prostate cancer management and enhance the continuum of care in Brunei Darussalam.


Study Design and Data Source

This retrospective cohort study included patients with prostate cancer treated at the Brunei Cancer Centre (TBCC), Jerudong Park Medical Centre (JPMC) between January 2018 and December 2024. The study was conducted in accordance with the TRIPOD-AI (Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis adapted for Artificial Intelligence) reporting guidelines to enhance transparency, reproducibility, and clinical applicability (Multimedia Appendix 1). Eligible participants were male patients with a histologically confirmed diagnosis of prostate cancer who received at least 1 treatment modality at TBCC. Patients with secondary metastatic prostate cancer or multiple primary malignancies were excluded following clinical review by oncologists. In total, 212 eligible patients contributed 478 time-indexed observations, capturing longitudinal treatment pathways and disease progression, as illustrated in Figure 1. Demographic, clinical, and treatment data were extracted from multiple sources, including the Brunei Darussalam Healthcare Information and Management System, the JPMC laboratory database, and both electronic and manually documented clinical records.

‎
Figure 1. Overview of cohort data processing and source integration for the study. BruHIMS: Brunei Darussalam Healthcare Information and Management System; JPMC: Jerudong Park Medical Centre; NGS: next-generation sequencing.

Data Structure and Outcome Definition

Data were structured across 2 hierarchical levels: the individual case level and the longitudinal cohort level. At the case level, collected variables encompassed demographic characteristics (including age, ethnicity, and marital status), medical history (eg, comorbidities and family history of malignancy), lifestyle factors (particularly smoking status), anthropometric measures (height, weight, and BMI), tumor-related features (Gleason score, tumor stage, and ECOG performance status), and survival outcomes. At the cohort level, comprehensive treatment information and 23 longitudinal laboratory measures were assembled, including baseline and 1-year posttreatment PSA values, full blood count parameters, liver and renal function indicators, erythrocyte sedimentation rate, calcium, and phosphate levels. Absolute and relative changes in PSA were derived, and tumor staging, Gleason grade groups, and National Comprehensive Cancer Network prognostic risk categories were harmonized to address inconsistencies in coding practices [24]. Except for PSA, all laboratory parameters were baseline measurements obtained before each treatment modification, enabling clinicians to assess laboratory results when determining subsequent therapeutic strategies. Data integrity was ensured through systematic cross-validation across multiple data sources and clinical verification by oncologists before analysis. The primary outcome was PFS.

PFS was calculated on a treatment-episode basis from treatment initiation to the earliest occurrence of PSA-defined progression, which was determined using PCWG-consistent criteria based on treatment-specific baseline PSA, posttreatment PSA nadir, and confirmed PSA rises (≥25% and ≥2.0 ng/mL, with timing and confirmation rules applied) or death from any cause, with patients censored at last follow-up in the absence of an event. A PFS event value of 0 indicates no biochemical progression or survival without progression at the last follow-up (censored), whereas a value of 1 indicates the occurrence of biochemical PSA progression or death from any cause. To conduct a survival analysis, PFS events, which denote the occurrence of progression or death, were used rather than PFS, which indicates the time a patient remains progression-free.

Data Preprocessing and Survival Analysis

Statistical analyses were performed using RStudio (version 4.3.3; R Foundation for Statistical Computing), Python (version 3.10; Python Software Foundation) implemented via Google Colab, and Visual Studio Code (version 1.106.0; Microsoft). Two analytic datasets were created: a case-level dataset consisting of 212 patients and a longitudinal cohort-level dataset comprising 478 time-indexed observations. Data preprocessing procedures included data cleaning, harmonization of tumor staging and Gleason grade groups, elimination of duplicate records, and construction of patient-specific longitudinal sequences. At the case level, continuous variables were compared using the Mann-Whitney U test, while categorical variables were assessed using the Fisher exact test, with a significance threshold set at P<.05. For the cohort-level analysis, mixed-effects linear and logistic regression models were used to evaluate differences between survival groups. The differences between groups, including those stratified by comorbidity status, were assessed using the log-rank test.

Data Splitting and Missing Data Handling

The overall ML model development pipeline is illustrated in Figure 2A. To minimize information leakage, data were split at the individual case level into training (80%) and testing (20%) sets using a random group-splitting strategy [25]. Missing values in both static and longitudinal variables were addressed using extremely randomized trees–based iterative imputation, a nonparametric approach derived from the random forest algorithm that accommodates mixed data types and captures nonlinear relationships [26]. Categorical variables were transformed using one-hot encoding, and continuous variables were standardized using z score normalization to ensure comparable feature scales. All preprocessing steps, including imputation, scaling, and encoder fitting, were performed solely on the training dataset and subsequently applied to the testing dataset.

‎
Figure 2. (A) Overview of the machine learning model development workflow for progression-free survival (PFS). (B) Recurrent autoencoder (RAE) architecture for longitudinal representation. AUC: area under the curve; GRU: gated recurrent unit; PCWG: Prostate Cancer Working Group; PSA: prostate-specific antigen.

Longitudinal Representation Learning Using a Recurrent Autoencoder

A recurrent autoencoder (RAE) was trained on the training cohort to capture temporal patterns in laboratory and treatment histories (Figure 2B) [20]. PFS duration served as the primary temporal input, with all laboratory variables recorded as baseline measurements before each treatment decision. PSA measurements were used exclusively for PFS event determination according to biochemical progression criteria and were not incorporated as predictive features in model training. The PFS event and time were used only as indexing variables, not as features, to avoid learning in the autoencoder, thereby ensuring that the autoencoder learned purely unsupervised representations of longitudinal patient trajectories without exposure to the outcome. The RAE encoder, composed of stacked gated recurrent units (GRUs), transformed each patient’s sequence into a low-dimensional latent embedding, which the decoder then reconstructed to preserve the original temporal patterns. These latent embeddings were concatenated with case-level static features (eg, race and morbidities) to form a multimodal feature vector. To prevent data leakage, the RAE was trained exclusively on the training dataset, and the learned embeddings were fixed before application to the testing set. Hyperparameters included a hidden layer size of 128, an embedding dimension of 32, a learning rate of 0.01, 100 epochs, and a batch size of 32.

Model Development

Traditional Cox Proportional Hazards Model

A Cox proportional hazards model was fitted to the training dataset as a conventional baseline for PFS analysis. The model was adjusted for clinically relevant covariates to quantify the association between predictors and the hazard of disease progression or death, defined as the instantaneous risk of a progression event at a given time, conditional on remaining progression-free up to that point [27,28]. Final covariate selection was guided by model fit, assessed using the log-likelihood ratio test. The proportional hazards assumption was examined using scaled Schoenfeld residuals and corresponding statistical tests to detect systematic departures from time-invariant covariate effects [19-32]. Predicted risk scores were generated for the testing dataset to facilitate direct comparison with ML-based PFS models.

ML Models

Three ML-based survival models were developed using an integrated feature set combining static clinical variables and latent temporal representations derived from the RAE. First, a random survival forest (RSF) model was trained as a nonparametric ensemble method to estimate cumulative hazard functions and rank predictors by importance [33]. Second, a gradient boosting survival (GBS) model (GBT-Survival) was implemented using a Cox partial likelihood loss function, enabling the modeling of nonlinear effects and complex interactions among covariates [34]. Third, a deep neural network–based Cox regression model (DeepSurv) was trained to capture nonlinear relationships between predictors and progression risk by optimizing the Cox partial likelihood [20]. Hyperparameter optimization for all models was performed using grid search cross-validation on the training dataset [35]. For the RSF, the explored parameter grid included the number of estimators (100, 200, 300), maximum depth (5, 7, 10), and minimum samples per split (10, 20, 30). For the GBS (GBT-Survival) model, the grid included learning rates (0.001, 0.01, 0.05, 0.1), the number of estimators (100, 200, 300), and maximum depth (3, 5, 7, 10). Optimal configurations, selected based on the highest concordance index, included RSF (maximum depth=7, minimum samples per split=10, 100 estimators), GBT-Survival (learning rate=0.001, maximum depth=10, 300 estimators), and DeepSurv (hidden layers=[32,35], dropout=0.1, rectified linear unit activation).

Model Evaluation and Explainability

Time-dependent area under the receiver operating characteristic curve (AUC) values could not be reliably estimated at 1, 2, or 3 years for any model because there were too few progression events and insufficient comparable patient pairs within these early periods to assess discrimination accurately. Model performance was assessed using time-dependent AUC values at the 4- and 5-year PFS horizons, together with Harrell concordance index (C-index) to evaluate discriminative ability [36]. Model robustness was examined using bootstrap resampling with 1000 iterations to derive CIs. Calibration was evaluated using the integrated Brier score (IBS) and calibration plots comparing predicted and observed progression-free survival probabilities. To enhance model interpretability, SHAP was applied to the best-performing RSF-Survival model, providing both global and patient-level insights into feature contributions to progression risk. SHAP summary, dependence, and force plots were used to visualize time-dependent predictor effects and identify key drivers of disease progression [19].

Ethical Considerations

Ethics approval for this study was granted by the Medical and Health Research and Ethics Committee, Ministry of Health, Brunei Darussalam (MHREC/MOH/2024/22(1)), on January 25, 2025. All patient data were replaced with pseudonymous research IDs since we received the data from the hospital, and no personally identifiable information was accessed or retained within the analytic or modeling environment.

Software Packages and Analytic Tools

All analyses were performed in Python using established open-source libraries, including pandas and NumPy for data manipulation; scikit-survival for Cox proportional hazards modeling, RSF, GBS analysis, concordance index estimation, and time-dependent AUC; lifelines and pycox for survival modeling (including DeepSurv); PyTorch (torch, torchvision, and torchaudio) for deep learning implementation; fancyimpute for iterative imputation; scikit-learn for data splitting and preprocessing (GroupShuffleSplit and StandardScaler); and SHAP for model interpretability.


Characteristics and Univariable Analysis of Study Participants

The study included 212 patients at the case level, contributing a total of 478 longitudinal observations over 6 years (detailed participant characteristics are presented in Multimedia Appendix 2). PFS events were reported in 45 observations during follow-up. At the case level (n=212), PFS outcomes differed according to tumor-related factors, including prostate cancer stage and Gleason grade group, whereas other comorbid conditions did not differ significantly between PFS outcome groups. Furthermore, at the cohort level (478 observations), significant differences between observations with and without a PFS event were observed in calcium (P<.001), chloride (P<.001), sodium (P<.001), red blood cell count (P<.01), hemoglobin (P<.01), neutrophils (P<.01), eosinophils (P<.01), baseline PSA (P<.001), and posttreatment PSA (P<.001).

Kaplan-Meier Curve for PFS

The PFS analysis with significant variables is presented using the Kaplan-Meier curves in Figure 3. The overall PFS curve demonstrates a steady decline over time, with the probability of remaining progression-free markedly decreasing after approximately 2 years of follow-up. Stratified analyses revealed significant differences across merged clinical stages (log-rank test, P=.0002), with advanced stages showing poorer PFS than earlier stages. Similarly, significant variation in PFS was observed among different prostate cancer treatment groups (log-rank test, P=.0009), indicating heterogeneity in outcomes depending on the therapeutic approach. Furthermore, PFS differed significantly according to place of treatment (log-rank test, P=.0033), with the inpatient treatment group demonstrating consistently lower survival probability over time.

‎
Figure 3. Kaplan-Meier curves for progression-free survival among prostate cancer patients at the Brunei Cancer Centre (TBCC) from January 2018 to December 2024. ADT: androgen deprivation therapy; ARPI: androgen receptor pathway inhibitor; IP: inpatients; OP: outpatients.

Discriminative Performance of PFS Models

The discriminative performance, as measured by the time-dependent AUC at 1, 3, and 5 years (Figure 4), and the Harrell C-index (Table 1), was used to develop and compare 4 survival models. Evaluation of model performance at 4 and 5 years indicated that ML-based models outperformed the traditional Cox proportional hazards models. The Cox model demonstrated limited discriminative ability, with time-dependent AUCs of 0.706 and 0.833 at 4 and 5 years, respectively, along with a C-index of 0.531. In contrast, the tree-based models, RSF and GBT-Survival, showed markedly superior discrimination. The RSF model demonstrated strong discriminative performance at 4 and 5 years (AUCs of 0.941 and 0.917, respectively; C-index 0.906), while the deep survival model (AUCs of 0.941 at 4 years and 0.917 at 5-year; C-index 0.719) and the GBS model (AUCs of 0.765 at 4 years and 0.833 at 5-year, C-index 0.844) demonstrated moderate discriminative performance.

‎
Figure 4. Heatmap of time-dependent area under the receiver operating characteristic curve (AUC) of progression-free survival (PFS) models, showing the relationship between predicted accuracy in fourth and fifth years. Cox: Cox proportional hazards model; DeepSurv: deep neural network–based Cox regression model; GBS: gradient boosting survival; RSF: random survival forest.
Table 1. Discriminative performance (time-dependent area under the receiver operating characteristic curve [AUC], Harrell concordance index [C-Index]) and calibration accuracy (integrated Brier score) of progression-free survival models.
ModelsTime-dependent AUCHarrell C-indexIntegrated Brier score

4 years5 years

Cox proportional hazards0.7060.8330.5310.0527
Random survival forest0.9410.9170.9060.0698
Gradient boosting survival0.7650.8330.8440.0590
Deep neural network survival0.9410.9170.7190.0887

Calibration Performance Using IBS for PFS Model

Calibration performance of the PFS models was assessed using the IBS, where lower values indicate better agreement between predicted and observed outcomes (Table 1). Among the evaluated models, the Cox proportional hazards model demonstrated the best calibration accuracy with the lowest IBS of 0.0527. The GBS model demonstrated comparable calibration performance (IBS=0.0590), followed by the RSF model (IBS=0.0698). In contrast, the deep neural network survival model exhibited the highest IBS (IBS=0.0887), indicating comparatively poorer calibration despite its better discriminative performance.

Although similar time-dependent AUC values were observed between the RSF and deep neural network survival models at specific predefined time points, the limited number of events in the study may reduce the sensitivity of AUC to detect subtle performance differences between models. Therefore, we placed greater emphasis on the C-index and IBS, as these metrics provide a more comprehensive assessment of survival model performance across the entire follow-up period compared with isolated time-point AUC measurements. According to both discriminative and calibration performance, the RSF model demonstrated stable and well-calibrated risk prediction over time.

Model Explainability Using SHAP for RSF Models for PFS

Model explainability for PFS prediction using the RSF was assessed using SHAP, providing both global and individual-level interpretability (Figure 5). At the global level, the SHAP feature importance heatmap (Figure 5A) indicated that biomedical and laboratory parameters, particularly baseline sodium, alanine aminotransferase (ALT), mean corpuscular hemoglobin (MCH) level, weight, and the treatment variable (Rx_C_D_Merged), were the most influential contributors to the model’s PFS predictions, while demographic variables showed comparatively lower impact. These features highlighted the dominant role of laboratory biomarkers in capturing progression risk in patients with prostate cancer. SHAP values reflect model attribution, not a causal effect.

At the individual level, with the sample patient, SHAP bar, force and waterfall plots (Figure 5B) illustrated how specific features contributed to the predicted PFS for a representative patient, demonstrating the direction and magnitude of each feature’s effect on expected survival time. ALT, sodium, MCH, platelet count, and treatment category increased the predicted risk of progression, whereas other features had minimal or neutral influence. These outputs offer clinical interpretability of how individual patient characteristics contribute to PFS predictions, thereby enhancing transparency, trust, and potential clinical utility in the prognosis of prostate cancer.

‎
Figure 5. (A) Heatmap of global Shapley additive explanations (SHAP) feature importance for the random survival forest model. (B) Top 10 features by importance for an individual patient, bar, waterfall and force plots. ALT: alanine transaminase; BPH: benign prostatic hyperplasia; Ca: calcium; Cl: chloride; Hb: Hemoglobin; Hep: hepatitis; Na: sodium; Plat: platelet; PSA: prostate-specific antigen; Rx_C_D_Merged: treatment categories.

Principal Findings

The study findings demonstrated that ML-based models showed improved internal discrimination compared with the traditional Cox proportional hazards model in predicting PFS among patients with prostate cancer, particularly at long-term horizons. While early PFS prediction at 1 to 3 years was not feasible due to insufficient progression events and limited comparable risk sets, robust discrimination emerged at 4 and 5 years. This reflects the delayed progression patterns and sluggish natural history of prostate cancer that are frequently seen in clinical practice [37-39]. Among the evaluated models, RSF showed the most balanced performance, achieving high time-dependent AUC values and stable calibration as measured by the IBS. These findings highlight the importance of aligning evaluation time points with disease biology and reinforce the limitations of early event prediction risk in prostate cancer cohorts [8,12,40].

Additionally, the study provided an explainability output from the RSF model using SHAP, which allows addressing the “black-box” barrier for clinical adoption of AI in oncology [19]. The SHAP-based global and individual explanations revealed that the laboratory biomarkers and treatment-related variables dominated the prediction of PFS. This aligns with emerging evidence that data-driven survival models capture complex, nonlinear associative predictors of risk that are often missed by conventional regression-based approaches [20,41]. Furthermore, the SHAP force and waterfall plots at the individual patient level enable clinicians to understand not only what the model predicts but how a particular prediction is made, which can further improve trust, auditability, and shared decision-making in clinical oncology [19,42,43].

Moreover, the application of a RAE to model the longitudinal cohort structure enabled effective compression of repeated observations into latent representations, providing a novel and efficient approach for managing complex temporal data while preserving clinically relevant information. The performance of RSF and GBS models in this cohort was consistent with earlier reports demonstrating the superiority of tree-based methods for cancer survival prediction [19,44]. Additionally, similar to previous studies on prostate cancer, RSF models outperformed Cox models when handling high-dimensional laboratory and longitudinal data, particularly in heterogeneous real-world cohorts [36,45]. The underperformance of Cox regression in this setting supports the rationale for exploring ML survival methods, particularly RSF, which can capture complex nonlinear interactions without proportional hazards assumptions. Although earlier studies reported strong short-term predictive performance, our findings suggest that PFS prediction in prostate cancer is inherently limited at early points [39,40]. The comparatively weaker calibration of the deep neural network survival model highlighted that higher discrimination does not necessarily translate into reliable absolute risk estimation.

According to the SHAP model, sodium, ALT, and hematological parameters such as MCH and platelet count emerged as influential predictors of PFS. Dysnatremia has been associated with cancer progression and systemic inflammation, while elevated ALT may reflect hepatic stress, metabolic dysfunction, or treatment-related toxicity, all of which have been linked to poorer oncological outcomes [46-48]. Moreover, studies have indicated that the aspartate aminotransferase to ALT ratio can be used as a prognostic biomarker for prostate tumor progression [49,50]. These findings support the growing recognition that routine laboratory biomarkers capture systemic physiological states relevant to cancer progression beyond tumor-centric variables alone [51]. However, the prominence of sodium and ALT in our SHAP analysis does not indicate that these markers are biologically superior to other prognostic factors; rather, within this dataset and modeling framework, they emerged as influential predictive contributors that improved the model’s ability to estimate risk.

This study has several limitations that should be acknowledged. First, the relatively small sample size, while sufficient for internal model development and evaluation, may restrict the generalizability of the results to other populations, particularly in light of regional differences in clinical practice and disease characteristics. This also limits the model’s ability to capture early PFS events. Performance estimates are likely optimistic due to limited event counts. Despite these limitations, our study offers important novel contributions that address key gaps in the literature. No prior study, to our knowledge, has collected the comprehensive range of variables included here or applied a sequential approach to define PFS, whereas most existing studies use a single PFS event per patient, limiting clinical applicability. Second, the retrospective nature of the study may introduce bias due to missing data, unmeasured confounding, and evolving treatment strategies over time. Furthermore, model performance was evaluated using a held-out testing dataset; although cross-validation was applied to reduce overfitting, the absence of external validation remains a limitation.

Third, although the analysis was retrospective and conducted at a single institution, the TBCC functions as the only comprehensive oncology facility in the country and is responsible for the management of nearly all prostate cancer cases nationwide. Nevertheless, the findings highlight the potential utility of ML models incorporating SHAP-based explainability for prognostic risk assessment in prostate cancer. Finally, radiographic progression according to PCWG criteria was not incorporated into the PFS analysis because of the limitations of the imaging data available for this cohort. Future research should include independent external validation and multi-institutional, multiomics datasets with larger and more diverse cohorts to strengthen the generalizability and robustness of these findings.

Conclusions

Our study demonstrates that the ML-based survival models, particularly tree-based ensemble methods, substantially outperform the traditional Cox proportional hazards model in predicting PFS among patients with prostate cancer, especially at longer follow-up horizons. Among the evaluated approaches, the RFS model achieved the most balanced performance, combining strong long-term discrimination with stable calibration, thereby offering a robust framework for prognostic modeling in real-world prostate cancer cohorts. The integration of SHAP-based explainability addressed a key barrier to clinical adoption by providing transparent global and individual-level insights into model predictions highlighting the dominant role of laboratory biomarkers and treatment-related variables in driving PFS risk predictions in this cohort. The application of a RAE to encode longitudinal cohort data further enhanced the model’s ability to capture the complex temporal patterns while reducing dimensionality. These findings support the clinical potential of explainable ML models for personalized risk stratification in prostate cancer, while also emphasizing the need for external validation, larger multi-institutional, multiomics datasets, and prospective evaluation prior to routine clinical implementation.

Acknowledgments

The authors would like to thank the Ministry of Health, Brunei Darussalam, and the Brunei Cancer Centre for granting permission to access the required data for this study. We also acknowledge the support of the Medical and Health Research and Ethics Committee for ethics approval and Universiti Brunei Darussalam for guidance and support. The authors are grateful to all individuals and institutions who contributed to the successful completion of this research. During the preparation of this manuscript, the authors used Grammarly (Superhuman Platform Inc) to assist with language refinement. No generative AI tools were used for data analysis, data interpretation, or the generation of scientific conclusions. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Funding

This work was supported by Universiti Brunei Darussalam through research grant UBD/RSCH/1.6/FICBF/2025/008, awarded to HAR.

Data Availability

The aggregate data and results supporting the findings of this study are included in the published article and its supplementary materials, where applicable. The underlying participant-level dataset is not available because of privacy and confidentiality considerations and to protect the anonymity of study participants. The complete analysis code is publicly available on GitHub [52].

Authors' Contributions

All authors were involved in the conceptualization of this report. HMT was responsible for data curation, investigation, methodology, resources, visualization, and writing of the original draft. HAR was involved in methodology development, supervision, resources, and writing, reviewing, and editing the manuscript. OAM provided supervision and review. LN contributed to supervision, writing, reviewing, and editing the manuscript. MSA and TT contributed to conceptualization, methodology, and resources. All authors approve the content of the manuscript.

Conflicts of Interest

None declared.

Multimedia Appendix 1

Discriminative performance (time-dependent area under the curve, Harrell concordance index) and calibration accuracy (integrated Brier score) of progression-free survival models.

DOCX File , 13 KB

Multimedia Appendix 2

Characteristics of participants stratified by progression-free survival.

DOCX File , 22 KB

  1. Prostate. International Agency of Research on Cancer. URL: https://gco.iarc.who.int/media/globocan/factsheets/cancers/27-prostate-fact-sheet.pdf [accessed 2025-12-17]
  2. Schafer EJ, Laversanne M, Sung H, Soerjomataram I, Briganti A, Dahut W, et al. Recent patterns and trends in global prostate cancer incidence and mortality: an update. Eur Urol. Mar 2025;87(3):302-313. [FREE Full text] [CrossRef] [Medline]
  3. Hassanipour S, Delam H, Arab-Zozani M, Abdzadeh E, Hosseini SA, Nikbakht HA, et al. Survival rate of prostate cancer in Asian countries: a systematic review and meta-analysis. Ann Glob Health. Jan 02, 2020;86(1):2. [FREE Full text] [CrossRef] [Medline]
  4. Zhou X, Wei Q. Global spatiotemporal trends in prostate cancer burden and its socioeconomic disparities: an observational study from 1990 to 2021. J Clin Oncol. Feb 18, 2025;43(5_suppl):325. [FREE Full text] [CrossRef]
  5. Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229-263. [FREE Full text] [CrossRef] [Medline]
  6. Brunei Darussalam Cancer Registry Report 2002-2021. Noncommunicable Disease (NCD) Prevention Unit, Ministry of Health Brunei Darussalam. 2022. URL: https:/​/www.​researchgate.net/​publication/​394256941_Brunei_Darussalam_Cancer_Registry_Report_2002-2021 [accessed 2026-09-25]
  7. Leong E, Ong SK, Si-Ramlee KA, Naing L. Cancer incidence and mortality in Brunei Darussalam, 2011 to 2020. BMC Cancer. May 22, 2023;23(1):466. [FREE Full text] [CrossRef] [Medline]
  8. Halabi S, Vogelzang NJ, Ou SS, Owzar K, Archer L, Small EJ. Progression-free survival as a predictor of overall survival in men with castrate-resistant prostate cancer. J Clin Oncol. Jun 10, 2009;27(17):2766-2771. [FREE Full text] [CrossRef] [Medline]
  9. Halabi S, Roy A, Rydzewska L, Guo S, Godolphin P, Hussain M, et al. Radiographic progression-free survival and clinical progression-free survival as potential surrogates for overall survival in men with metastatic hormone-sensitive prostate cancer. J Clin Oncol. Mar 20, 2024;42(9):1044-1054. [CrossRef] [Medline]
  10. Belin L, Tan A, De Rycke Y, Dechartres A. Progression-free survival as a surrogate for overall survival in oncology trials: a methodological systematic review. Br J Cancer. May 2020;122(11):1707-1714. [FREE Full text] [CrossRef] [Medline]
  11. Scher HI, Morris MJ, Stadler WM, Higano C, Basch E, Fizazi K, et al. Trial design and objectives for castration-resistant prostate cancer: updated recommendations from the Prostate Cancer Clinical Trials Working Group 3. J Clin Oncol. Apr 20, 2016;34(12):1402-1418. [FREE Full text] [CrossRef] [Medline]
  12. Oh EL, Huish W, El-Gamil S, Benson T, Ferguson T. Assessment of patient and clinician perspectives on clinically meaningful extension of progression-free survival in prostate cancer. Eur Urol Open Sci. Nov 12, 2024;70:175-182. [FREE Full text] [CrossRef] [Medline]
  13. Bang S, Ahn YJ, Koo KC. Harnessing machine learning to predict prostate cancer survival: a review. Front Oncol. Jan 10, 2025;14:1502629. [FREE Full text] [CrossRef] [Medline]
  14. Bove S, Arezzo F, Cormio G, Silvestris E, Cafforio A, Comes MC, et al. Explainable machine learning for predicting recurrence-free survival in endometrial carcinosarcoma patients. Front Artif Intell. Dec 06, 2024;7:1388188. [FREE Full text] [CrossRef] [Medline]
  15. Rask Kragh Jørgensen R, Bergström F, Eloranta S, Tang Severinsen M, Bjøro Smeland K, Fosså A, et al. Machine learning-based survival prediction models for progression-free and overall survival in advanced-stage Hodgkin Lymphoma. JCO Clin Cancer Inform. Apr 2024;8:e2300255. [FREE Full text] [CrossRef] [Medline]
  16. Vershinina O, Sushentsev N, Zaikin A, Blyuss O, Barrett T, Ivanchenko M. Improving the potential for predicting prostate cancer progression in patients on active surveillance using explainable artificial intelligence. Cancers (Basel). Nov 07, 2025;17(22):3598. [FREE Full text] [CrossRef] [Medline]
  17. Li Y, Chai X, Yang M, Xiong J, Zeng J, Chen Y, et al. Accurate prediction of disease-free and overall survival in non-small cell lung cancer using patient-level multimodal weakly supervised learning. NPJ Precis Oncol. Jun 19, 2025;9(1):197. [FREE Full text] [CrossRef] [Medline]
  18. Chen Z, Ouyang H, Sun B, Ding J, Zhang Y, Li X. Utilizing explainable machine learning for progression-free survival prediction in high-grade serous ovarian cancer: insights from a prospective cohort study. Int J Surg. May 01, 2025;111(5):3224-3234. [FREE Full text] [CrossRef] [Medline]
  19. Krzyziński M, Spytek M, Baniecki H, Biecek P. SurvSHAP(t): time-dependent explanations of machine learning survival models. Knowl Based Syst. Feb 28, 2023;262:110234. [FREE Full text] [CrossRef]
  20. Katzman JL, Shaham U, Cloninger A, Bates J, Jiang T, Kluger Y. DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med Res Methodol. Feb 26, 2018;18(1):24. [FREE Full text] [CrossRef] [Medline]
  21. Lee C, Yoon J, van der Schaar M. Dynamic-DeepHit: a deep learning approach for dynamic survival analysis with competing risks based on longitudinal data. IEEE Trans Biomed Eng. Jan 2020;67(1):122-133. [FREE Full text] [CrossRef] [Medline]
  22. Sathekge M, Bruchertseifer F, Vorster M, Lawal IO, Knoesen O, Mahapane J, et al. Predictors of overall and disease-free survival in metastatic castration-resistant prostate cancer patients receiving 225 Ac-PSMA-617 radioligand therapy. J Nucl Med. Jan 2020;61(1):62-69. [FREE Full text] [CrossRef] [Medline]
  23. Ponce-Bobadilla AV, Schmitt V, Maier CS, Mensing S, Stodtmann S. Practical guide to SHAP analysis: explaining supervised machine learning model predictions in drug development. Clin Transl Sci. Nov 2024;17(11):e70056. [FREE Full text] [CrossRef] [Medline]
  24. Prostate cancer, version 2. National Comprehensive Cancer Network. 2025. URL: https:/​/www.​nccn.org/​login?ReturnURL=https:/​/www.​nccn.org/​professionals/​physician_gls/​pdf/​prostate.​pdf [accessed 2025-05-07]
  25. GroupShuffleSplit. Scikit-Learn. URL: https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupShuffleSplit.html [accessed 2026-09-23]
  26. Geurts P, Ernst D, Wehenkel L. Extremely randomized trees. Mach Learn. Mar 2, 2006;63:3-42. [FREE Full text] [CrossRef]
  27. Breslow NE. Analysis of survival data under the proportional hazards model. Int Stat Rev. Apr 1975;43(1):45-57. [FREE Full text] [CrossRef]
  28. Kalbfleisch JD, Schaubel DE. Fifty years of the Cox model. Annu Rev Stat Appl. 2023;10:1-23. [FREE Full text] [CrossRef]
  29. In J, Lee DK. Survival analysis: part II - applied clinical data analysis. Korean J Anesthesiol. Oct 2019;72(5):441-457. [FREE Full text] [CrossRef] [Medline]
  30. Schoenfeld D. Partial residuals for the proportional hazards regression model. Biometrika. Apr 1982;69(1):239-241. [FREE Full text] [CrossRef]
  31. Abeysekera WW, Sooriyarachchi MR. Use of Schoenfeld’s global test to test the proportional hazards assumption in the Cox proportional hazards model: an application to a clinical study. J Natn Sci Foundation Sri Lanka. Mar 29, 2009;37(1):41-45. [FREE Full text] [CrossRef]
  32. Grambsch PM, Therneau TM. Proportional hazards tests and diagnostics based on weighted residuals. Biometrika. Sep 1994;81(3):515-526. [FREE Full text] [CrossRef]
  33. Berkowitz M, Altman RM, Loughin TM. Random forests for survival data: which methods work best and under what conditions? Int J Biostat. Apr 24, 2024;20(2):315-345. [FREE Full text] [CrossRef] [Medline]
  34. Chen Y, Jia Z, Mercola D, Xie X. A gradient boosting algorithm for survival analysis via direct optimization of concordance index. Comput Math Methods Med. 2013;2013:873595. [FREE Full text] [CrossRef] [Medline]
  35. Adnan M, Alarood AA, Uddin MI, Ur Rehman I. Utilizing grid search cross-validation with adaptive boosting for augmenting performance of machine learning models. PeerJ Comput Sci. Feb 21, 2022;8:e803. [FREE Full text] [CrossRef] [Medline]
  36. Mogensen UB, Ishwaran H, Gerds TA. Evaluating random forests for survival analysis using prediction error curves. J Stat Softw. Sep 2012;50(11):1-23. [FREE Full text] [CrossRef] [Medline]
  37. Pound CR, Partin AW, Eisenberger MA, Chan DW, Pearson JD, Walsh PC. Natural history of progression after PSA elevation following radical prostatectomy. JAMA. May 05, 1999;281(17):1591-1597. [FREE Full text] [CrossRef] [Medline]
  38. Hamdy FC, Donovan JL, Lane JA, Mason M, Metcalfe C, Holding P, et al. 10-year outcomes after monitoring, surgery, or radiotherapy for localized prostate cancer. N Engl J Med. Oct 13, 2016;375(15):1415-1424. [FREE Full text] [CrossRef] [Medline]
  39. Wilt TJ, Jones KM, Barry MJ, Andriole GL, Culkin D, Wheeler T, et al. Follow-up of prostatectomy versus observation for early prostate cancer. N Engl J Med. Jul 13, 2017;377(2):132-142. [FREE Full text] [CrossRef] [Medline]
  40. Kattan MW, Wheeler TM, Scardino PT. Postoperative nomogram for disease recurrence after radical prostatectomy for prostate cancer. J Clin Oncol. May 1999;17(5):1499-1507. [FREE Full text] [CrossRef] [Medline]
  41. Lundberg SM, Lee SI. A unified approach to interpreting model predictions. In: NIPS'17: Proceedings of the 31st International Conference on Neural Information Processing Systems. New York, NY. Curran Associates Inc; 2017:4768-4777.
  42. Tun HM, Rahman HA, Naing L, Malik OA. Trust in artificial intelligence-based clinical decision support systems among health care workers: systematic review. J Med Internet Res. Jul 29, 2025;27:e69678. [FREE Full text] [CrossRef] [Medline]
  43. Holzinger A, Langs G, Denk H, Zatloukal K, Müller H. Causability and explainability of artificial intelligence in medicine. Wiley Interdiscip Rev Data Min Knowl Discov. 2019;9(4):e1312. [FREE Full text] [CrossRef] [Medline]
  44. Adatorwovor R, Ogunsanya ME, Huang B, Charnigo R, Abraham O. Comparison of machine learning models for colon cancer survival: predictive modeling approach. JMIR Cancer. Nov 26, 2025;11:e72665. [FREE Full text] [CrossRef] [Medline]
  45. Ching T, Himmelstein DS, Beaulieu-Jones BK, Kalinin AA, Do BT, Way GP, et al. Opportunities and obstacles for deep learning in biology and medicine. J R Soc Interface. Apr 2018;15(141):20170387. [FREE Full text] [CrossRef] [Medline]
  46. Stangl-Kremser J, Kramer G, Shariat SF. Association of hyponatremia with survival in patients with castration-resistant prostate cancer: a clinical commentary. Clin Genitourin Cancer. Dec 2019;17(6):e1188-e1192. [FREE Full text] [CrossRef] [Medline]
  47. Djamgoz MB. Hyponatremia and cancer progression: possible association with sodium-transporting proteins. Bioelectricity. Mar 01, 2020;2(1):14-20. [FREE Full text] [CrossRef] [Medline]
  48. Bañez LL, Loftis RM, Freedland SJ, Presti JCJ, Aronson WJ, Amling CL, et al. The influence of hepatic function on prostate cancer outcomes after radical prostatectomy. Prostate Cancer Prostatic Dis. Jun 2010;13(2):173-177. [FREE Full text] [CrossRef] [Medline]
  49. Mitsui Y, Yamabe F, Hori S, Uetani M, Aoki H, Sakurabayashi K, et al. Longitudinal change in castration-resistant prostate cancer biomarker AST/ALT ratio reflects tumor progression. Sci Rep. Sep 15, 2023;13(1):15292. [FREE Full text] [CrossRef] [Medline]
  50. Zhou J, He Z, Ma S, Liu R. AST/ALT ratio as a significant predictor of the incidence risk of prostate cancer. Cancer Med. Aug 2020;9(15):5672-5677. [FREE Full text] [CrossRef] [Medline]
  51. Hanahan D. Hallmarks of cancer: new dimensions. Cancer Discov. Jan 2022;12(1):31-46. [FREE Full text] [CrossRef] [Medline]
  52. Explainable-machine-learning-PSF-. GitHub. URL: https://github.com/heinminntun/Explainable-Machine-Learning-PSF- [accessed 2026-09-23]


‎
ALT: alanine aminotransferase
AUC: area under the receiver operating characteristic curve
C-index: concordance index
GBS: gradient boosting survival
GRU: gated recurrent unit
IBS: integrated Brier score
JPMC: Jerudong Park Medical Centre
MCH: mean corpuscular hemoglobin
ML: machine learning
PCWG: Prostate Cancer Working Group
PFS: progression-free survival
PSA: prostate-specific antigen
RAE: recurrent autoencoder
RSF: random survival forests
SHAP: Shapley additive explanations
TBCC: Brunei Cancer Centre


Edited by M Balcarras; submitted 20.Jan.2026; peer-reviewed by R Ismayilov, C Colak, S Banerjee, J Guo, V Palama, A Arora; comments to author 07.Apr.2026; revised version received 08.Apr.2026; accepted 09.Apr.2026; published 05.Oct.2026.

Copyright

©Hein Minn Tun, Lin Naing, Owais Ahmed Malik, Muhammad Syafiq Abdullah, Thu Ta, Hanif Abdul Rahman. Originally published in JMIR Cancer (https://cancer.jmir.org), 05.Oct.2026.

This is an open-access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Cancer, is properly cited. The complete bibliographic information, a link to the original publication on https://cancer.jmir.org/, as well as this copyright and license information must be included.