Key Takeaways
- Researchers developed a machine-learning model to estimate the risk of hepatocellular carcinoma returning after ablation.
- The retrospective study included 447 patients with previously untreated, solitary hepatocellular carcinoma treated at a single medical center in China.
- The best-performing model used five factors: tumor size, alpha-fetoprotein, sex, glypican-3 expression, and Barcelona Clinic Liver Cancer stage.
- In the held-out test set, the model achieved an area under the curve of 0.790 for overall recurrence. Its highest time-specific AUC was 0.809 for recurrence within the first year after ablation.
- The model has not been externally or prospectively validated and has not been shown to improve surveillance decisions, treatment, survival, or other patient outcomes.
Introduction
Researchers developed a machine-learning model using clinical factors and GPC-3 expression to estimate hepatocellular carcinoma recurrence risk after ablation.
Hepatocellular carcinoma, or HCC, is the most common type of primary liver cancer. Even after an early tumor is treated with ablation, the cancer can return, making continued surveillance an important part of care.
Researchers are increasingly investigating whether machine learning can help identify patients with different levels of recurrence risk and whether those risks change over time.
A peer-reviewed study published online September 7, 2026, in BMC Cancer developed a machine-learning model that combined clinical characteristics with tumor expression of glypican-3, or GPC-3, to estimate recurrence risk after ablation (Li et al., 2026).
The model showed moderate ability to distinguish patients who experienced recurrence from those who did not, with its highest time-specific performance during the first year after treatment.
However, this was a retrospective study conducted at one medical center. The model has not yet been tested prospectively in an independent population, and the study did not determine whether using its predictions improves medical decisions or patient outcomes.
What the Study Examined
Li and colleagues analyzed 447 patients with treatment-naïve, solitary HCC who underwent ablation at the Fifth Medical Center of the Chinese People’s Liberation Army General Hospital in Beijing (Li et al., 2026).
The researchers used an ensemble feature-selection approach combining logistic regression, minimum Redundancy Maximum Relevance, and Mutual Information to identify characteristics that were most useful for predicting recurrence.
Five variables were selected:
- Tumor size
- Alpha-fetoprotein, or AFP
- Sex
- GPC-3 expression
- Barcelona Clinic Liver Cancer, or BCLC, stage
AFP is a blood biomarker commonly used as part of HCC assessment. BCLC staging incorporates features of the cancer, liver function, and a patient’s general condition to classify HCC and help guide management.
The researchers compared multiple predictive algorithms, including logistic regression, random forest, support vector machine, and Extreme Gradient Boosting, commonly called XGBoost.
XGBoost was selected as the best-performing model (Li et al., 2026).
The 447 patients were divided into training, validation, and test sets in a 3:1:1 ratio.
Keeping a test set separate from the data used for model training provides a more meaningful evaluation than assessing performance only in the training data. However, all three datasets came from the same institution.
The test set therefore represents held-out internal testing, not external validation in a separate hospital or patient population.
What Researchers Found
For overall HCC recurrence, the XGBoost model achieved an area under the receiver operating characteristic curve, or AUC, of 0.819 in the training set, 0.813 in the validation set, and 0.790 in the test set (Li et al., 2026).
AUC measures how well a model distinguishes between people who experience an outcome and those who do not. An AUC of 0.5 represents discrimination no better than chance, while an AUC of 1.0 represents perfect discrimination.
An AUC does not, however, show whether a model improves care, whether its risk estimates are well calibrated for individual patients, or whether acting on those predictions produces better outcomes.
The researchers also evaluated recurrence during different periods after ablation.
In the test set, the AUC was:
- 0.809 for recurrence within the first year
- 0.776 for recurrence between one and two years
- 0.781 for recurrence after two years
The highest time-specific AUC was therefore seen for recurrence during the first year after ablation (Li et al., 2026).
Kaplan-Meier analyses also showed separation in recurrence-free survival between patients the model classified as higher risk and those classified as lower risk.
The researchers additionally examined model performance across subgroups based on sex, age, and tumor size. They did not detect statistically significant differences in performance among the examined subgroups, with interaction P values greater than 0.7. However, these analyses do not establish that the model performs equally well across those groups, particularly given the limited sample size (Li et al., 2026).
Overall, the findings show that the model could distinguish different levels of recurrence risk within the population studied. They do not mean it can determine with certainty whether an individual patient’s cancer will return.
Why GPC-3 May Be Important
One distinctive aspect of the model is the inclusion of GPC-3 alongside more conventional clinical characteristics.
GPC-3 is a cell-surface protein studied as a potential HCC biomarker and therapeutic target. In the new model, tumor GPC-3 expression was combined with tumor size, AFP, sex, and BCLC stage rather than being used as a stand-alone predictor (Li et al., 2026).
Including information related to tumor biology may help a prediction model capture information not represented by conventional clinical variables alone.
But GPC-3 should not be interpreted as a universally available or standardized measurement.
A model that depends partly on GPC-3 may be affected by how tumor tissue is obtained, how the protein is measured, and how expression is classified. If those procedures differ between hospitals, the model may not perform in exactly the same way elsewhere.
That is another reason independent validation is necessary before the model can be considered for routine clinical use.
How the Findings Compare With Earlier Research
Machine learning has previously been investigated as a way to predict HCC recurrence following ablation.
A 2022 study examined health records from 1,574 patients with early-stage HCC who underwent microwave ablation at four hospitals (An et al., 2022).
The researchers compared logistic regression, random forest, support vector machine, and XGBoost models for predicting early recurrence.
XGBoost produced the strongest discrimination, with an AUC of 0.75 in the training set, 0.74 in an internal validation set, and 0.76 in an external validation set (An et al., 2022).
That earlier work provides useful context because it included multiple hospitals and tested the model in an external validation population.
The 2026 study takes a different approach by incorporating GPC-3 and assessing recurrence during separate periods after ablation (Li et al., 2026).
Timing may matter because factors associated with recurrence shortly after treatment may not be identical to those associated with recurrence years later.
A 2025 review of recurrence after curative-intent HCC resection or ablation described aggressive features of the original tumor, including tumor size, tumor number, and vascular invasion, as important factors in early recurrence. For later recurrence, characteristics such as viral infection and liver cirrhosis may play a greater role, although definitions of early and late recurrence vary and their underlying mechanisms can overlap (Meng et al., 2025).
The stronger first-year performance seen in the new model is compatible with the possibility that characteristics of the original tumor are particularly informative during the early post-treatment period.
However, the study was designed to develop a prediction model. It cannot establish why recurrence occurs earlier in some patients than in others.
What the Results May Mean
If future studies can reliably identify patients at different levels of recurrence risk, prediction models could potentially help researchers investigate more individualized approaches to post-ablation surveillance.
For example, future clinical studies could test whether risk-stratified monitoring improves the timing or efficiency of recurrence detection compared with established surveillance strategies.
That possibility remains theoretical for this particular model. The study did not test alternative surveillance schedules or determine whether changing the frequency of follow-up according to predicted risk would improve patient outcomes.
The researchers also did not demonstrate that doctors using the model make better decisions or that model-guided care results in earlier detection of recurrent cancer, more effective treatment, longer survival, or better quality of life.
This distinction is especially important when interpreting medical artificial-intelligence research.
A model may achieve respectable statistical performance without providing meaningful clinical benefit. Predictive discrimination is only one step toward establishing that a tool is useful and safe in real-world medicine.
The current findings therefore support additional research rather than changes to routine HCC follow-up.
Patients should continue to follow the surveillance plan recommended by their oncology or hepatology team. This research model should not be used independently to change imaging, laboratory monitoring, or treatment.
Limitations
The study has several limitations that substantially affect how its results should be interpreted.
The most important is its retrospective, single-center design.
All 447 patients came from the same medical center in China. Although investigators divided them into separate training, validation, and test sets, all three groups originated from the same institution (Li et al., 2026).
A held-out test set is useful for estimating performance on data that were not used to train the model, but it does not provide the same evidence as external validation in an independent medical center.
Patient characteristics, underlying liver disease, imaging practices, laboratory methods, tissue testing, ablation techniques, and clinical workflows may differ between hospitals. Any of those differences could affect performance.
The population was also highly selected. The model was developed in treatment-naïve patients with solitary HCC who underwent ablation.
Its results therefore should not automatically be generalized to people with multiple tumors, more advanced HCC, previously treated or recurrent disease, or patients managed primarily with surgical resection, liver transplantation, systemic therapy, or other treatments.
GPC-3 creates an additional portability issue because tissue availability, testing methods, and classification procedures may differ across institutions.
The retrospective design also leaves the study susceptible to selection bias, missing information, and unmeasured differences that statistical modeling cannot necessarily eliminate.
Another important limitation is that prediction does not establish causation. Variables that help an algorithm classify recurrence risk should not automatically be interpreted as biological causes of recurrence.
The subgroup results also require caution. Failure to detect statistically significant differences between subgroups does not prove equal performance, especially when a study was not large enough to provide highly precise estimates within every subgroup.
Most importantly, the study did not evaluate clinical impact.
Researchers have not shown that using the model improves surveillance decisions, recurrence detection, treatment selection, survival, quality of life, or any other patient-centered outcome.
Prospective evaluation and independent multicenter validation are therefore necessary before its clinical usefulness can be established.
Funding, Competing Interests, and Ethics
The researchers reported that the study received no specific grant from any funding agency in the public, commercial, or nonprofit sectors (Li et al., 2026).
The authors declared no competing interests.
The study was approved by the Ethics Committee of the Fifth Medical Center of the Chinese People’s Liberation Army General Hospital. According to the paper, informed consent was waived because the study was retrospective and used anonymized clinical data (Li et al., 2026).
Final Thoughts
The study provides evidence that a relatively compact machine-learning model can distinguish different levels of HCC recurrence risk after ablation within the patient population in which it was developed.
Using tumor size, AFP, sex, GPC-3 expression, and BCLC stage, the XGBoost model achieved a test-set AUC of 0.790 for overall recurrence. Its highest time-specific test-set AUC was 0.809 for recurrence within the first year after ablation (Li et al., 2026).
Those results are encouraging, but they should not be confused with evidence that the model improves medical care.
The next major step is independent, prospective validation across multiple hospitals and patient populations. Researchers would then need to determine whether using its predictions actually improves surveillance decisions or patient outcomes compared with established care.
Until that evidence is available, the model is best viewed as a promising research tool for recurrence-risk assessment rather than an established method for deciding how patients should be monitored after liver cancer ablation.
References
An, C., Yang, H., Yu, X., Han, Z.-Y., Cheng, Z., Liu, F., Dou, J., Li, B., Li, Y., Li, Y., Yu, J., & Liang, P. (2022). A machine learning model based on health records for predicting recurrence after microwave ablation of hepatocellular carcinoma. Journal of Hepatocellular Carcinoma, 9, 671–684. https://doi.org/10.2147/JHC.S358197
Li, H., Li, X., Li, J., Sun, Q., Bian, L., Zhang, Z., Wang, L., & Gao, Y. (2026). Machine learning-based prediction of time-dependent recurrence risk after ablation for hepatocellular carcinoma integrating glypican-3 and clinical features: A retrospective study. BMC Cancer. Advance online publication. https://doi.org/10.1186/s12885-026-16804-7
Meng, F., Wang, J., Zhu, X.-D., Zhang, M., Zhang, X., Cheng, D., Zhang, X., & Liu, L. (2025). Risk factors for recurrence in patients with hepatocellular carcinoma after curative resection or ablation. Journal of Hepatocellular Carcinoma, 12, 2501–2511. https://doi.org/10.2147/JHC.S552316




