This web page was created programmatically, to learn the article in its authentic location you’ll be able to go to the hyperlink bellow:
https://journals.plos.org/plosone/article%3Fid%3D10.1371/journal.pone.0357291
and if you wish to take away this text from our website please contact us
Lung most cancers stays one of many main causes of cancer-related mortality worldwide, the place early detection and dependable danger evaluation are important for enhancing outcomes. This research develops and evaluates a spread of ensemble studying fashions for lung most cancers prediction utilizing life-style and scientific indicators, with an emphasis on each predictive efficiency and interpretability. Five base learners—logistic regression, k-nearest neighbors, naïve Bayes, assist vector machine, and linear discriminant evaluation—have been used to assemble a number of boosting and bagging fashions. Building on these, voting and stacking ensembles have been designed by selectively combining high-performing fashions. All approaches have been evaluated on the unique dataset in addition to on balanced and upsampled variants derived by means of artificial augmentation. Model efficiency was assessed utilizing accuracy, precision, recall, F1-score, Matthews correlation coefficient (MCC), and AUC-ROC. The outcomes present that ensemble approaches persistently outperform particular person fashions, with voting and stacking demonstrating superior efficiency over boosting and bagging strategies. The stacking mannequin achieved the strongest total efficiency throughout all evaluated fashions. On the unique dataset, which gives a extra sensible illustration of sensible deployment circumstances, it attained an accuracy of 93.53%. Performance additional improved on the balanced and upsampled datasets on the upsampled dataset. To improve transparency, SHAP and LIME have been employed to offer international and native interpretability, respectively, enabling identification of key scientific components and patient-specific danger drivers. The evaluation highlights each alignment with identified scientific patterns and dataset-driven variations, supporting knowledgeable interpretation of mannequin outputs. The outcomes counsel that stacking-based ensembles can enhance danger prediction from life-style and scientific indicators whereas sustaining mannequin transparency by means of SHAP and LIME explanations. These findings spotlight the potential of interpretable ensemble studying as a decision-support instrument for lung most cancers danger evaluation.
Citation: Ganie SM, Dutta Pramanik PK, Zhao Z (2026) Lung most cancers danger prediction utilizing interpretable ensemble fashions on life-style and scientific knowledge. PLoS One 21(9):
e0357291.
https://doi.org/10.1371/journal.pone.0357291
Editor: Amgad Muneer, The University of Texas, MD Anderson Cancer Center, UNITED STATES OF AMERICA
Received: September 22, 2025; Accepted: August 13, 2026; Published: September 3, 2026
This is an open entry article, freed from all copyright, and could also be freely reproduced, distributed, transmitted, modified, constructed upon, or in any other case utilized by anybody for any lawful objective. The work is made out there below the Creative Commons CC0 public area dedication.
Data Availability: The dataset used within the experiment is publicly out there at https://www.kaggle.com/datasets/mysarahmadbhat/lung-cancer.
Funding: The creator(s) declared that monetary assist was acquired for this work and/or its publication. ZZ is partially funded by his startup at The University of Texas Health Science Center at Houston, Houston, Texas, USA. This work was additionally supported and funded by the Deanship of Scientific Research, Vice Presidency for Graduate Studies and Scientific Research, King Faisal University, Saudi Arabia (Grant No: KFU264609 to S.M.G.). We declare that the funders had no function in research design, knowledge assortment and evaluation, choice to publish, or preparation of the manuscript.
Competing pursuits: The authors have declared that no competing pursuits exist.
The fashionable life-style has induced many important and life-threatening ailments in humankind all over the world. Among these, most cancers continues to be the main reason behind mortality, answerable for round 10 million deaths in 2020, which equates to at least one in each six fatalities globally [1]. Lung most cancers is likely one of the commonest kinds of most cancers [2]. According to the World Health Organization (WHO), 1.8 million individuals died worldwide in 2020 attributable to lung most cancers. Hence, it has turn out to be a major problem for healthcare professionals and sufferers. It is famous that if the illness is detected at an early stage, an individual’s possibilities of survival may be elevated.
For the early detection of lung most cancers, it’s essential to precisely distinguish between benign (non-cancerous) and malignant (cancerous) cells [3]. Various lung analysis strategies exist that utilise imaging strategies, together with CT (computed tomography), isotope, X-ray, and MRI (magnetic resonance imaging), in addition to molecular screening, amongst others [4]. These image-analysis-based diagnostic strategies nonetheless face the problem of distinguishing benign from malignant cancers, as they share similarities. Importantly, these strategies make use of a reactive method, i.e., they’ll detect a cancerous cell as soon as it has turn out to be such. Instead, a proactive method could be a greater tactic, through which the likelihood of an individual getting most cancers, primarily based on her life-style knowledge, is predicted. Machine studying strategies present a possibility to establish patterns in life-style and scientific knowledge which will assist earlier danger evaluation.
Machine studying has proven important potential for medical professionals and researchers to find hidden prospects in medical, scientific, and demographic knowledge and higher serve human beings [5,6]. In the healthcare sector, machine studying is helpful for analysing ailments with acceptable accuracy, permitting medical doctors to make higher medical choices, which in flip results in improved analysis and prognosis. Machine studying strategies are efficiently utilized to foretell and diagnose deadly ailments, similar to most cancers, coronary heart illness, and mind tumours. However, conventional machine studying approaches could battle to seize more and more complicated and non-linear relationships current in massive heterogeneous healthcare datasets [7].
Traditional machine studying algorithms, similar to logistic regression (LR), naive Bayes (NB), Ok-nearest neighbour (KNN), assist vector machine (SVM), linear discriminant evaluation (LDA), and so forth., have some limitations in most cancers prediction because of the illness’s complexity, dynamicity, and the heterogeneity of knowledge [8]. Some of the numerous limitations are summarized under.
Employing extra subtle machine studying methodologies, similar to deep studying and ensemble strategies, could assist handle these constraints. In current years, ensemble studying paradigms, similar to boosting, bagging, voting, and stacking, have attracted rising consideration from the analysis neighborhood. These approaches have proven promising outcomes, demonstrating improved accuracy within the prediction of ailments [9].
Ensemble studying represents a machine studying paradigm that mixes the predictions of a number of particular person fashions to enhance total accuracy and scale back the chance of overfitting [10]. Several advantages are related to using ensemble studying strategies in most cancers prediction. Some of them are talked about under.
It is crucial to acknowledge that the easy software of ensemble studying doesn’t inherently yield the anticipated benefits. The number of ensemble strategies and foundational fashions needs to be knowledgeable by the particular traits of the issue, dataset properties, and out there computational assets [10]. When utilized appropriately, ensemble studying has the potential to markedly improve the efficiency of machine studying fashions [11].
Although ensemble fashions enhance predictive accuracy, they usually lack interpretability and obscure the prediction course of. Furthermore, whereas explainable synthetic intelligence (XAI) is more and more prevalent in healthcare, its software in ensemble fashions for lung most cancers prediction stays restricted [12–14]. Most analysis has targeted on elucidating particular person fashions with XAI fairly than complicated ensemble programs, thereby constraining the comprehensibility of those fashions in real-world scientific contexts. Despite developments in predictive fashions inside healthcare, there stays a paucity of fashions that combine the accuracy of ensemble studying with the interpretability of XAI for lung most cancers. Given the multitude of interconnected components influencing these circumstances, understanding mannequin predictions is crucial for clinicians to tailor remedies. Traditional ensemble strategies, though efficient, usually perform as “black boxes,” concealing the rationale behind predictions. This opacity diminishes their utility in scientific settings, the place offering clear, patient-specific insights for moral, knowledgeable decision-making is essential.
To handle this problem, this research proposes an XAI-enhanced stacking mannequin that not solely delivers excessive predictive efficiency for lung most cancers danger evaluation but additionally facilitates clear interpretation of mannequin outputs. In the area of scientific decision-making, the place transparency and accountability are important, interpretable AI strategies similar to SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) play a pivotal function. SHAP values present each international and native insights into characteristic significance, facilitating a complete understanding of the components influencing lung most cancers danger predictions. LIME, alternatively, gives localised explanations, enabling the interpretation of particular predictions on an instance-by-instance foundation. By integrating these explainability strategies, we guarantee statistical rigor in efficiency validation, thereby rendering our method extra dependable and explainable for scientific adoption.
The main contributions of this paper are as follows:
The remainder of the paper is organized as follows. Section 2 presents the associated work. Section 3 describes the proposed methodology and gives an outline of the thought-about machine studying and ensemble studying algorithms. Section 4 particulars the dataset. Section 5 presents and analyzes the experimental outcomes. Sections 6 and seven report the statistical significance evaluation and interprets the proposed fashions utilizing XAI, respectively. Section 8 discusses the findings and their scientific implications. Finally, Section 9 concludes the research, outlining its limitations and instructions for future work.
Machine studying and ensemble-based methodologies have demonstrated substantial potential to provide constant, dependable, and legitimate outcomes, thereby discovering widespread software throughout numerous real-world domains. In the context of lung most cancers prediction, important analysis has employed these strategies to reinforce diagnostic accuracy. For instance, Abdullah et al. [15] carried out a comparative evaluation when it comes to accuracies SVM, KNN, and a convolutional neural community (CNN) in classifying lung most cancers in its early phases. In their research, SVM achieved the very best accuracy (95.56%), whereas KNN achieved the bottom (88.40%). Patra [16] utilized machine studying classifiers, together with NB, KNN, and Radial Basis Function (RBF), utilizing the Weka instrument to analyse the lung most cancers knowledge. The dataset consisted of 32 cases with 57 options, collected from the University of California, Irvine (UCI) Machine Learning Repository. Among all of the classifiers, RBF achieved the very best accuracy of 81.25% utilizing 10-fold cross-validation. The creator recommended that characteristic choice strategies and integration with different machine studying fashions might improve the efficacy of the proposed work. Radhika et al. [17] employed machine studying strategies similar to DT, SVM, NB, and LR on two publicly out there datasets (Data World and UCI) for lung most cancers detection. By using the UCI dataset, the LR mannequin attained the very best accuracy fee of 96.90%, whereas SVM outperformed nicely and achieved an accuracy of 99.20% through the use of the world knowledge repository.
Numerous research have employed quite a lot of datasets, algorithms, and methodologies to advance analysis in lung most cancers prediction. An rising variety of researchers have explored the effectiveness of ensemble studying strategies in enhancing diagnostic and prognostic accuracy for lung most cancers. For occasion, Ahmad and Mayyaa [18] designed a instrument for lung most cancers prediction utilizing bagging and randomized node optimization. The authors thought-about numerous components related to lung most cancers, similar to genetic dangers, air air pollution, smoking, mud allergy, and so forth. and signs like coughing blood, fatigue, dry cough, shortness of breath, chest ache, and others. In a associated research, Safiyari and Javidan [19] employed 5 broadly adopted ensemble strategies, together with bagging, dagging, AdaBoost (ADB), multi-boosting, and random subspace utilizing knowledge from the Surveillance, Epidemiology, and End Results (SEER) program. Wang et al. [20] launched an ensemble-based mannequin combining Random Forest (RF) with a self-paced studying bootstrap method to reinforce lung most cancers classification and prognosis. Their methodology employed a sampling technique that progressively built-in knowledge from high- to lower-quality samples. Mamun et al. [21] applied numerous ensemble algorithms, together with Extreme Gradient Boosting (XGB), Light Gradient Boosting Machine (LGBM), bagging, and ADB, evaluated by means of 10-fold cross-validation. Among these, XGB demonstrated superior efficiency, attaining an accuracy of 94.42%.
Ensemble studying methodologies, when paired with XAI, have attracted important consideration in lung most cancers prediction. This curiosity stems from their twin skill to reinforce predictive efficiency whereas offering interpretability—a vital part in scientific decision-making and early analysis. Numerous research have aimed to stability predictive accuracy with transparency. For occasion, RF has been reported to attain 98.9% accuracy throughout numerous datasets, highlighting its potential in medical diagnostics [22]. Another modern framework, DeepXplainer, combines CNNs with XGB and achieves 97.43% accuracy. This mannequin employs SHAP to supply each native and international interpretability, thereby enhancing the transparency of the system’s predictions [23]. An interpretable radiomics-based mannequin was developed utilizing a stacked interpretable sequencing cells (SISC) framework, which not solely demonstrated state-of-the-art outcomes but additionally supplied visualization instruments like response maps to help radiologists in choice assist [24]. Similarly, the ExMed platform, which includes a number of characteristic attribution strategies, has been utilized to artificial healthcare knowledge to foretell lung most cancers sufferers’ life expectancy, additional illustrating the applicability of XAI along side ensemble studying fashions [25]. In one other contribution, Siddhartha et al. [26] used a bagging-based ensemble methodology to categorise the survival standing of lung most cancers sufferers by figuring out and decoding the chance components. They developed a machine studying pipeline (RF) by mitigating the issue of imbalanced knowledge. The central contribution of this research lies in its use of XAI to interpret RF predictions, a important side for fostering belief and reliability in healthcare functions. However, the research lacks a comparative evaluation with different ensemble or deep studying fashions, which limits insights into the relative effectiveness of RF. Zhang et al. [27] aimed to develop a prediction mannequin utilizing machine studying to establish high-risk teams for early identification and prevention of lung most cancers. The efficiency of the fashions, together with LR, NB, RF, and XGB was evaluated utilizing the AUC (space below the curve), Brier loss, log loss, precision, recall, and F1 scores. The SHAP interpreter was used to visualise the fashions, displaying the affect of every predictive issue on lung most cancers danger. Another research, using the SEER 2020 database, in contrast a number of conventional classifiers—LR and DT—with ensemble strategies like XGB, LGBM, and ADB, in addition to deep studying fashions similar to DNNs. LGBM emerged because the main performer, with an accuracy of 87.45% and an AUC of 0.91 [28]. The research highlighted the necessity to combine XAI instruments to reinforce the interpretability of deep studying frameworks, notably for his or her potential use in scientific trials.
In abstract, integrating ensemble strategies—particularly superior strategies like stacking—with interpretability frameworks similar to SHAP holds important promise for enhancing each the efficiency and transparency of lung most cancers prediction fashions. These mixed approaches assist construct clinician belief in AI programs, addressing a long-standing barrier to AI deployment in scientific settings [29,30]. However, the present physique of labor primarily focuses on a restricted set of ensemble strategies, with minimal comparative analysis throughout a broader vary of fashions. Most research additionally confine themselves to traditional efficiency indicators like accuracy and F1-score, overlooking extra nuanced metrics that would higher assess scientific applicability. Furthermore, whereas XAI has been employed in lots of research, its integration with complicated ensemble frameworks stays underexplored. There is proscribed understanding of how mannequin interpretability may be systematically built-in into scientific workflows or how totally different options work together synergistically to have an effect on mannequin outputs. This hole underscores the necessity for future analysis to analyze extra complicated characteristic interactions, broader datasets, and rigorous efficiency analysis frameworks. Additionally, comparative research that consider ensemble strategies alongside conventional fashions might present invaluable insights into their respective benefits, limitations, and real-world applicability in healthcare settings.
This part gives a complete overview of the analysis methodology employed and the ensemble studying strategies utilised within the experiment.
The following phases current the experiment:
Ensemble studying is a robust paradigm in machine studying through which a number of base fashions, also known as weak learners, are mixed to create a extra dependable and correct predictive framework. The core idea underpinning ensemble strategies is that the collective efficiency of a number of reasonably efficient fashions can exceed that of a single, probably extra complicated mannequin.
In ensemble studying, the weak learners are sometimes easy algorithms, similar to DT or LR, which individually exhibit solely marginal predictive functionality above random likelihood. These weak fashions are educated utilizing totally different subsets of the info, using numerous algorithms, and even totally different characteristic units. This variety permits every mannequin to seize totally different features or patterns throughout the knowledge. The ensemble mannequin thus advantages from the aggregated data of its constituent learners, successfully compensating for particular person mannequin weaknesses and enhancing total predictive accuracy [31].
The ultimate prediction of an ensemble mannequin is normally obtained by combining the predictions of every particular person weak learner. This integration of a number of fashions can improve total efficiency, because the ensemble can appropriate the errors and biases of particular person fashions. Ensemble studying has been proven to enhance accuracy, robustness, and generalization capabilities in comparison with utilizing a single mannequin.
In this research, we experimented with a number of ensemble strategies, together with bagging, boosting, voting, and stacking, to find out probably the most appropriate method for lung most cancers prediction utilizing a life-style and scientific dataset. The purpose is to leverage the inherent benefits of ensemble studying to develop a predictive mannequin that achieves excessive accuracy and reliability within the early analysis and danger analysis of lung most cancers.
In this research, we employed 5 generally used machine studying algorithms because the weak learners for the ensemble mannequin: LR [32], KNN [33], NB [34], SVM [35], and LDA [36]. LR is a well-liked linear classification algorithm that fashions the likelihood of a binary end result, whereas KNN is a non-parametric methodology that classifies knowledge factors primarily based on their proximity. NB is a probabilistic classifier that applies Bayes’ theorem to make predictions, and SVM is a robust algorithm that finds the optimum hyperplane to separate lessons. Finally, LDA is a dimensionality discount method that initiatives knowledge right into a lower-dimensional house whereas preserving class separability.
Bagging is an ensemble method through which a number of fashions are educated on totally different subsets of the coaching knowledge, generated through sampling with alternative. This resampling technique permits particular person knowledge factors to seem a number of instances inside a given subset. Each mannequin is educated on a distinct random pattern. During prediction, the outputs of particular person fashions are aggregated through voting or averaging. Random sampling introduces randomness, which helps scale back variance, and aggregation helps scale back bias. In this research, we examined 5 bagging algorithms: DT [37], BDT [38], RF [39], ET [40], and BME [38].
Boosting builds an ensemble sequentially, with every new mannequin trying to appropriate the errors of the earlier fashions. It focuses extra on beforehand misclassified cases fairly than the general error fee. This forces new fashions to focus on beforehand misclassified cases. New fashions are added one after the other, and every is educated on a weighted dataset that assigns larger weights to beforehand misclassified cases. In this research, we examined 5 boosting algorithms: GB [41], XGB [42], CB [43], LGBM [44], and ADB [45].
Multiple fashions are educated independently on the total coaching dataset in voting [46]. During prediction, every mannequin predicts a category or worth for a check occasion. The predictions are mixed utilizing a easy voting mechanism, similar to majority voting, for classification. For regression, the common of predictions from particular person fashions turns into the ultimate prediction. In this research, LR, CB, XGB, RF, and ET are employed as the bottom fashions within the voting course of.
In stacking, a secondary mannequin, also known as a meta-learner, is educated to combine the outputs of a number of main fashions generally known as base learners [46]. Initially, the bottom learners are educated on the unique coaching knowledge, and their predictions on a validation set are collected. These predictions function enter options for the meta-learner, which is then educated to generate the ultimate prediction. At the inference stage, the bottom fashions produce predictions on the check knowledge, that are subsequently handed to the educated meta-learner to find out the final word output. The key goal of stacking is to optimize the mixture of base fashions by studying from their particular person prediction patterns. This layered method permits the ensemble to mannequin intricate relationships amongst base learners, usually leading to enhanced predictive accuracy [47]. In this research, we utilised the identical set of base learners as within the voting ensemble, particularly LR, CB, XGB, RF, and ET.
The lung most cancers dataset utilized on this research was sourced from a publicly accessible repository on Data World (https://www.kaggle.com/datasets/mysarahmadbhat/lung-cancer), contributed by staceyinrobert. This dataset includes 309 data and contains 16 demographic attributes. Among these, the primary 15 variables function predictors (unbiased variables), whereas the ultimate attribute represents the goal variable (dependent variable). An in depth description of the dataset attributes is supplied in Table 1. Before mannequin implementation, the dataset was examined for knowledge high quality points, particularly outliers and lacking values. The interquartile vary methodology was employed to establish potential outliers, and imputation strategies have been used to evaluate lacking knowledge. However, no outliers or lacking values have been detected, indicating a clear and full dataset appropriate for evaluation.
The authentic dataset was comparatively small for coaching ensemble studying fashions. To study the affect of dataset dimension and sophistication distribution on predictive efficiency, two augmentation methods have been adopted. First, the minority class was balanced utilizing the Synthetic Minority Over-sampling Technique (SMOTE). Second, each lessons have been upsampled to create a considerably bigger dataset for investigating mannequin behaviour below elevated pattern availability. Accordingly, three dataset variants have been thought-about:
SMOTE was employed to generate artificial minority-class samples by interpolating between neighbouring minority cases within the characteristic house. The implementation adopted the usual configuration of the imbalanced-learn library utilizing sampling_strategy = ‘auto’ and k_neighbors = 5. Under this setting, every artificial pattern was generated by interpolating between a minority occasion and one in all its 5 nearest minority neighbours recognized utilizing the Euclidean distance metric. The similar SMOTE configuration was used persistently all through all augmentation experiments to make sure methodological consistency.
After producing three datasets, normalization was utilized to every utilizing the min-max scaling method, as outlined in Eq. 1, the place, x denotes the attribute worth, and xmin and xmax symbolize the attribute’s minimal and most values, respectively. This methodology remodeled the values of all attributes to a standardized vary between 0 and 1, thereby making certain uniformity in scale throughout the datasets.
Because the unique dataset was extraordinarily small and extremely imbalanced, the balanced and upsampled datasets have been generated previous to cross-validation. We acknowledge that this method could introduce optimistic bias and doesn’t absolutely get rid of the potential of info leakage. Consequently, the outcomes obtained on the balanced and upsampled datasets needs to be interpreted primarily as comparative analyses of class-balancing methods fairly than definitive estimates of real-world efficiency.
Table 2 exhibits the descriptive statistics of the parameters throughout the three datasets. Here, complete report counts in every dataset, together with the imply, normal deviation, and minimal and most values of every attribute, are famous.
This part presents and discusses the experimental particulars and outcomes achieved utilizing the designed ensemble algorithms for lung most cancers prediction.
A pc system with Intel® CoreTM i9-10900K CPU (3.70 GHz) was used for the experiment. The system’s different {hardware} specs are – RAM: 64 GB (DDR4), HDD: 2 TB, SSD: 500GB (NVMe). The pc was working Windows 11 Pro. Programs have been written in Python on Jupyter Notebook.
Several normal analysis metrics, as detailed in Table 3 have been used to evaluate the designed fashions’ skill to foretell lung most cancers.
As mentioned in Section 3.1, this part presents the main points of the ensemble mannequin constructing processes in two phases.
To get hold of dependable mannequin optimisation whereas lowering the chance of overfitting, stratified 10-fold cross-validation (ok = 10) was employed throughout mannequin improvement, as illustrated in Fig 2. For every dataset variant, the info have been first partitioned into 75% coaching and 25% testing subsets. The stratified 10-fold cross-validation process was then carried out solely on the coaching knowledge for hyperparameter optimisation and mannequin choice. Stratification preserved the category distribution inside every fold, offering steady mannequin optimisation regardless of the comparatively small dimension of the unique dataset.
Within every fold, knowledge preprocessing and have scaling have been fitted solely on the coaching partition and subsequently utilized to the corresponding validation partition. The fashions have been educated on the coaching folds and validated on the corresponding held-out fold. After completion of cross-validation, the optimized mannequin was retrained utilizing the total coaching dataset and subsequently evaluated on the unbiased 25% check set. The efficiency metrics reported on this research correspond to this ultimate test-set analysis.
Because the unique dataset contained solely 309 cases, together with 39 non-cancer instances, the balanced and upsampled datasets have been generated previous to cross-validation to allow significant mannequin coaching and comparative analysis below totally different class-distribution settings. We acknowledge that this design selection could introduce optimistic bias and subsequently the outcomes obtained on the balanced and upsampled datasets needs to be interpreted primarily as comparative analyses fairly than definitive estimates of real-world efficiency.
Hyperparameter optimization performs a pivotal function in influencing the habits of coaching algorithms and considerably impacts mannequin efficiency. To fine-tune the hyperparameters and guarantee optimum efficiency of the proposed ensemble fashions, each grid search and random search strategies have been employed [47]. Among these, grid search yielded extra favorable outcomes. Table 4 presents an in depth abstract of the hyperparameter configurations for every mannequin.
In the primary section, we constructed fifteen fashions for lung most cancers prediction, as follows:
Each of those fashions was applied utilizing all three dataset variants. As described in Section 5.3.1, the info have been partitioned into coaching (75%) and testing (25%) subsets, whereas mannequin choice and hyperparameter optimisation have been carried out utilizing stratified 10-fold cross-validation on the coaching knowledge.
The efficiency of every mannequin on every dataset as per the thought-about analysis metrics, is proven in Fig 3. Among the fifteen experimented fashions, LR, LDA, and BDT achieved the very best prediction accuracy of 90.71% on the precise dataset, whereas SVM (81.17%), KNN (86.1%), and DT (87.09%) are the worst fashions when it comes to accuracy. On the balanced dataset, the highest three fashions by accuracy are RF (94.86%), CB (94.59%), and GB (94.05%). Whereas on the upsampled dataset, DT (97.78%), ET (97.43%), and RF (96.71%) achieved the very best accuracy. Across the three datasets, RF achieved the very best common accuracy at 93.63%, adopted by ET (93.06%) and LGBM (92.98%). Likewise, the efficiency for the opposite metrics may be interpreted from the determine.
In this section, we constructed voting and stacking fashions utilizing the highest 5 fashions (as per their accuracy, as offered in Section 5.3.3) for every dataset, as proven in Table 5. Fig 4 exhibits the efficiency of the voting and stacking fashions utilizing every mixture for every dataset (as proven in Table 5). The outcomes present that there’s not a lot distinction between the voting and stacking fashions. Except for accuracy, the place stacking performs barely higher than voting throughout all datasets, efficiency is combined throughout datasets for different metrics. However, stacking outperformed voting on the upsampled dataset in all metrics, besides precision and AUC.
In the ultimate section, we rigorously chosen 5 fashions to construct the voting and stacking ensembles. After evaluating a number of combos, we finalized LR, CB, XGB, RF, and ET. These fashions symbolize essentially totally different studying paradigms: a linear mannequin (LR), boosting-based ensembles (CB and XGB), and bagging-based ensembles (RF and ET). This cross-family choice was meant to seize numerous choice boundaries and complementary prediction behaviours.
To additional justify this choice, we analyzed the variety among the many constituent learners utilizing pairwise Q-statistics and error-correlation measures. Ensemble studying advantages not solely from sturdy particular person learners but additionally from variety amongst their prediction errors. For two classifiers, the Q-statistic is outlined as
the place
The variety evaluation, summarized in Table 6, exhibits that the chosen learners are neither extremely redundant nor fully unbiased. As anticipated, RF and ET exhibited the very best settlement attributable to their related bagging-based tree architectures, whereas combos involving LR and tree-based ensembles demonstrated considerably higher variety. Most of the remaining pairs exhibited average variety, indicating complementary prediction patterns throughout the chosen learners. The noticed variety means that the chosen learners seize totally different features of the info and subsequently present complementary info to the ensemble.
This mixture was applied on all three datasets. The pipelines of the proposed voting and stacking fashions utilizing the ultimate mixture are illustrated in Fig 5. The performances of the voting and stacking fashions throughout the ten folds for every dataset are proven in Fig 6. The optimized hyperparameter values of the constituent fashions are offered in Table 7.
The characteristic significance utilizing the proposed voting and stacking fashions on precise, balanced, and upsampled datasets are proven in Fig 7, Fig 8, and Fig 9, respectively. From these figures, it may be seen that probably the most important components for correct predictions are AL, AC, and AG, whereas GD, SK, and CP have the bottom contributions.
The studying curves of the voting and stacking fashions for precise, balanced, and upsampled datasets are proven in Fig 10, Fig 11, and Fig 12, respectively. These curves illustrate how mannequin efficiency evolves because the coaching knowledge will increase, facilitating the identification of potential overfitting or underfitting. The coaching and validation curves point out the reliability of the hybrid fashions’ studying habits. The curves for each voting and stacking fashions observe clean, near-linear paths, indicating steady studying with out overfitting or underfitting on each the precise and balanced datasets. This means that the fashions maintained steady studying behaviour regardless of using artificial augmentation.
However, utilizing upsampled knowledge, the voting mannequin exhibited a major variation between the coaching and validation curves. In distinction, the stacking mannequin carried out finest throughout all cases in lowering the loss between the coaching and validation phases.
The confusion matrices for the proposed voting and stacking fashions are proven in Table 8. In normal, stacking performs higher on misclassifications. Fig 13 exhibits the common (over 10 folds) efficiency of the voting and stacking fashions throughout the three datasets. The voting and stacking fashions carried out at par. In normal, stacking was barely higher than voting, however in some instances (e.g., AUC), voting was marginally higher. Fig 14 exhibits the efficiency deviations of the voting and stacking fashions throughout ten folds for every dataset. Voting exhibits much less deviation between the precise and balanced datasets, whereas stacking performs barely higher on the upsampled dataset.
The AUC-ROC curves of the voting and stacking fashions for precise, balanced, and upsampled datasets are proven in Fig 15, Fig 16, and Fig 17, respectively. Overall, each stacking and voting fashions carried out equally nicely throughout all datasets. However, the voting mannequin barely outperformed the stacking mannequin in macro-average AUC-ROC on the upsampled dataset.
This part presents the efficiency of the proposed hybrid voting and stacking fashions and compares them in opposition to different fashions on this research and related state-of-the-art approaches within the literature.
This part compares the voting and stacking fashions from Phase III in opposition to fashions from Phases I and II. To guarantee a targeted comparability, the highest three performing fashions for every metric have been chosen from Phase I, contemplating base learners, boosting, and bagging strategies for every dataset. On the precise dataset, LR, LDA, and BDT confirmed the very best accuracies. On the balanced dataset, RF, CB, and GB carried out finest. For the upsampled dataset, DT, ET, and RF achieved the highest accuracies.
The comparative performances throughout numerous metrics, accuracy, precision, recall, F1-score, MCC, kappa, and AUC, are proven in Fig 18, Fig 19, Fig 20, Fig 21, Fig 22, Fig 23, and Fig 24, respectively. The proposed voting and stacking fashions of Phase III outperformed all different fashions in each analysis metric besides recall. Recall that utilizing stacking on the precise dataset is second-best (subsequent to CB), whereas on the upsampled dataset it achieves a recall similar to the most effective recall of LR and surpasses that of different fashions. Overall, the stacking mannequin outperformed the voting mannequin throughout all metrics and datasets.
Fig 18. Accuracy comparison of the top three base models, voting and stacking using the top five models and voting and stacking using selected models for a) actual dataset, b) balanced dataset, and c) upsampled dataset.
Fig 19. Precision comparison of the top three base models, voting and stacking using the top five models and voting and stacking using selected models for a) actual dataset, b) balanced dataset, and c) upsampled dataset.
Fig 20. Recall comparison of the top three base models, voting and stacking using the top five models and voting and stacking using selected models for a) actual dataset, b) balanced dataset, and c) upsampled dataset.
Fig 21. F1-score comparison of the top three base models, voting and stacking using the top five models and voting and stacking using selected models for a) actual dataset, b) balanced dataset, and c) upsampled dataset.
Fig 22. MCC comparison of the top three base models, voting and stacking using the top five models and voting and stacking using selected models for a) actual dataset, b) balanced dataset, and c) upsampled dataset.
Fig 23. Kappa comparison of the top three base models, voting and stacking using the top five models and voting and stacking using selected models for a) actual dataset, b) balanced dataset, and c) upsampled dataset.
Fig 24. AUC of the top three base models, voting and stacking using top five models and voting and stacking using selected models for a) actual dataset, b) balanced dataset, and c) upsampled dataset.
Fig 25 presents a comparability of the composite efficiency scores (CCS) for the highest three base fashions, in addition to voting and stacking ensembles utilizing each the highest 5 and chosen fashions throughout the three dataset variants. The CCS is outlined as the common of a number of analysis metrics (accuracy, precision, recall, F1-score, MCC, Kappa, and AUC) and is used right here as a unified measure for comparative mannequin evaluation. The outcomes present that ensemble fashions typically outperform particular person fashions, with stacking persistently attaining larger CCS values than voting. LightGBM is a notable exception, outperforming the voting ensemble (high 5) on each the precise and balanced datasets. The highest CCS is achieved by the stacking ensemble on the upsampled dataset, with scores of 83.09 and 85.35 for the highest 5 and chosen fashions, respectively. A transparent pattern is noticed throughout datasets: CCS values improve from the unique to the balanced and additional to the upsampled dataset. This displays improved mannequin efficiency below lowered class imbalance and elevated pattern dimension. Tree-based fashions similar to LightGBM, AdaBoost, and XGBoost persistently carry out nicely, doubtless attributable to their skill to seize nonlinear characteristic interactions and deal with class imbalance extra successfully.
Because the balanced and upsampled datasets have been generated through artificial augmentation, a further experiment was carried out to evaluate whether or not the stage at which SMOTE was utilized affected mannequin efficiency. Two settings have been thought-about: (i) artificial samples generated earlier than knowledge splitting and (ii) artificial samples generated solely from the coaching knowledge after splitting. In the latter setting, the info have been first partitioned into coaching and validation subsets, after which SMOTE was fitted and utilized solely to the coaching subset utilizing the identical configuration adopted in the principle experiments. The corresponding validation subset retained its authentic class distribution and contained no artificial observations all through mannequin coaching and analysis, thereby stopping info leakage through the resampling course of. The corresponding outcomes are reported in Table 9.
For the balanced dataset, making use of SMOTE after splitting produced barely larger efficiency throughout most analysis metrics. The selected-5 stacking mannequin, as an illustration, improved from 93.78% to 95.67% accuracy, whereas the F1-score elevated from 93.98% to 95.87%. Similar enhancements have been noticed for the voting ensembles.
The similar comparability was carried out on the upsampled dataset. Although absolutely the values assorted, the general behaviour of the fashions remained unchanged. Stacking persistently outperformed voting, and the selected-5 stacking mannequin continued to attain the most effective outcomes, recording 99.94% accuracy, 99.47% recall, and 99.74% AUC below the post-splitting SMOTE protocol.
An fascinating commentary is that the stricter resampling protocol didn’t scale back efficiency. In a number of instances, the outcomes have been marginally higher. More importantly, the relative rating of the ensemble fashions remained steady throughout each settings. This means that the benefit of the stacking framework isn’t depending on a selected SMOTE implementation technique. At the identical time, the outcomes needs to be interpreted throughout the broader limitations of artificial augmentation and the comparatively small dimension of the unique dataset.
The effectiveness of the proposed ensemble mannequin was validated by means of comparative analysis in opposition to related analysis papers, utilizing a number of efficiency metrics detailed in Table 10. We chosen the stacking mannequin for the upsampled dataset, given its superior efficiency over the voting mannequin in our research. The comparative evaluation contains research that used ensemble studying strategies to foretell lung most cancers from demographic knowledge.
The leads to Table 10 show the advance our stacking method achieves over present strategies. This development stems from the strategic number of fashions and meta-learners, rigorous cross-validation, and complete hyperparameter optimisation.
To rigorously assess efficiency variations among the many evaluated fashions, a non-parametric framework was employed utilizing the Friedman aligned-rank check, adopted by Holm-adjusted put up hoc comparisons. The outcomes point out clear and statistically important variations throughout fashions.
As proven in Table 11, the Friedman check yields extremely important outcomes for each the voting and stacking with chosen base fashions ensembles (p ≈ 0.000), resulting in rejection of the null speculation. This confirms that the noticed efficiency variations are statistically significant fairly than attributable to random variation.
The rating leads to Table 12 additional substantiate this discovering. Both voting and stacking with chosen base fashions obtain the very best rank scores (105.86 and 107.86, respectively), clearly outperforming all baseline fashions. Among the person learners, GB and RF persistently rank because the next-best-performing fashions. In distinction, less complicated fashions similar to SVM, KNN, and Naive Bayes occupy the bottom ranks, indicating comparatively weaker efficiency.
The put up hoc evaluation gives detailed pairwise comparisons. As reported in Table 13, the voting (with chosen fashions) mannequin exhibits statistically important enhancements over a number of baseline strategies, together with SVM, KNN, NB, BME, AdaBoost, CatBoost, LDA, and BDT. However, no statistically important variations are noticed compared with stronger fashions similar to GB, RF, LGBM, LR, DT, and ET. In distinction, the STS mannequin demonstrates broader statistical superiority; it achieves statistically important enhancements over a wider set of fashions, together with XGBoost along with weaker learners. Similar to voting (with chosen fashions), its variations from top-performing fashions similar to GB and RF usually are not statistically important, indicating comparable efficiency on the highest degree.
Finally, the Wilcoxon check leads to Table 14 verify that the proposed ensemble approaches yield statistically important enhancements on the dataset degree (p ≈ 0.0425), additional supporting the robustness of the noticed efficiency positive aspects.
XAI encompasses methodologies that improve the transparency and comprehensibility of AI fashions, thereby rendering their choices accessible to human specialists [48]. This is especially important in domains similar to healthcare, the place belief and accountability are paramount [49]. XAI helps each international and native explanations, that are important for making certain the reliability and utility of fashions in healthcare contexts. Global explanations elucidate the first components influencing illness outcomes, thereby aiding clinicians and researchers, whereas native explanations make clear how these components have an effect on particular person sufferers, successfully bridging the hole between analysis and scientific apply.
To increase the interpretability of the proposed stacking and voting fashions for predicting lung most cancers on a worldwide scale, we employed the SHAP methodology [50]. Grounded in Shapley values from cooperative recreation concept, SHAP gives a constant framework for assessing every characteristic’s impression on particular person predictions, making it a broadly adopted instrument for elucidating complicated machine studying fashions [51]. Additionally, we utilized LIME for native or instant characteristic interpretation [52], providing fast, case-specific insights into mannequin predictions, which is especially useful for real-time decision-making.
Global explanations present complete AI mannequin habits insights throughout affected person populations by figuring out key options like age, genetic markers, and lab outcomes influencing predictions. This evaluation ensures alignment of medical data, validates mannequin choices, and identifies inconsistencies that want refinement. Global explanations additionally assist establish biases, guarantee equity throughout demographics, and assist moral compliance with requirements such because the General Data Protection Regulation (GDPR), the Health Insurance Portability and Accountability Act (HIPAA), and Food and Drug Administration (FDA) rules. This transparency fosters belief in AI-driven medical choices. In our evaluation, we used SHAP characteristic significance values to rank options by their affect on predictions, no matter whether or not the impression is constructive or destructive.
The international analyses of stacking and voting fashions throughout three datasets are depicted by means of Beeswarm plots, ranked by imply absolute SHAP values in Fig 26, Fig 27, and Fig 28, for precise, balanced, and upsampled datasets. Features are ranked from high to backside by their impression on lung most cancers prediction. The x-axis shows SHAP values, representing every characteristic’s impression on the mannequin’s output. A constructive SHAP worth signifies contribution to predicting lung most cancers, whereas a destructive worth suggests the prediction of lung most cancers absence. Red dots symbolize excessive characteristic values, and blue dots symbolize low values, displaying how characteristic magnitude impacts prediction. Red dots point out that larger characteristic values correlate with a better likelihood of lung most cancers. Blue dots point out a decrease likelihood of an affiliation with lung most cancers, as proven by destructive SHAP values. For instance, if a affected person’s characteristic F1 has a SHAP worth of −0.2, it decreases the expected likelihood of lung most cancers by 0.2. Conversely, if F1’s SHAP worth is + 0.3, the mannequin predicts a 0.3 improve, suggesting lung most cancers.
In the stacking mannequin, the options FT, CD, and AC exhibit the very best imply absolute SHAP values, indicating their dominant function in predictions. The broad dispersion of SHAP values for FT—spanning constructive and destructive impacts—suggests a context-dependent affect, the place larger values (pink dots) are strongly related to elevated most cancers danger, whereas decrease values (blue dots) lower the prediction likelihood. Features AL and PP present important however unidirectional results, reinforcing their identified associations with most cancers prognosis. Features WZ, AT, YF, and SK exhibit combined impacts (dots on either side of zero), suggesting context-dependent results. Features SB and GD cluster close to zero, implying negligible contributions, reflecting both organic irrelevance or inadequate dataset illustration. The voting mannequin’s plot exhibits related high options (AC, AL, FT) and shares the identical minimal-impact options (AG, SP, CP, GD) with stacking.
Similarly, utilizing the balanced dataset, options similar to AL, AC, FT, CD, and PP stay high contributors for each fashions. In distinction, options similar to GD, SB, CP, and AG proceed to have minimal impression on each fashions. Using the upsampled dataset, AL, AC, FT, CD, SD, and PP are probably the most influential predictors in each stack and voting. Features similar to GD, SB, CP, and AG as soon as once more present very low or near-zero SHAP values, indicating minimal affect on predictions throughout each fashions.
Local explanations make clear particular person mannequin predictions, that are essential in healthcare for making patient-specific choices. By figuring out key components, similar to biomarkers or medical historical past, these explanations improve transparency, facilitate therapy planning, and foster clinician-patient belief. They additionally reveal errors by highlighting options that result in incorrect outcomes, enabling mannequin refinement. Additionally, they permit professionals to evaluate whether or not predictions align with established analysis, thereby enhancing transparency in AI diagnostics. This research utilises LIME plots to interpret stacking and voting fashions in lung most cancers prediction. LIME generates localized explanations by approximating mannequin habits round cases, environment friendly for real-time use. Compared to SHAP, which gives exact however computationally intensive explanations of Shapley values, LIME affords faster insights for exploratory evaluation. Using LIME and SHAP collectively balances interpretability, depth, and computational effectivity in scientific decision-making.
The LIME outputs of voting fashions are proven in Fig 29–30, and Fig 31 for precise, balanced, and upsampled datasets. The LIME outputs of stacking fashions are proven in Fig 32–33 and Fig 34 for precise, balanced, and upsampled datasets. They present insights into characteristic contributions and prediction chances for particular cases. The visualisation illustrates the affect of options on mannequin choices. The left plot shows the prediction chances for lung most cancers. The center part presents characteristic contributions, whereas the suitable facet lists options with precise values. Orange bars point out options rising lung most cancers danger, whereas blue bars symbolize protecting components.
The voting mannequin predicts 86% probability of most cancers (class 1) and 14% probability of no most cancers (class 0) primarily based on the precise dataset (Fig 29). AC (+0.35) contributes most to lung most cancers prediction, adopted by AT (+0.20), YF (+0.13), SD +(0.09), SK (+0.08), CP (+0.06), and AG (+0.01). AL (−0.24) and FT (−0.23) exert biggest affect on predicting no lung most cancers. PP (−0.17), CO (−0.15), and CD (−0.14) have a destructive impression on lung most cancers prediction. In the balanced dataset (Fig 30), AL (−0.16), CD (−0.16), and FT (+0.13) are probably the most important options indicating the absence of lung most cancers. AC (+0.14) is the strongest predictor for lung most cancers, adopted by SD (+0.9), YF (+0.05), AT (+0.04), and CP (+0.03). In the upsampled dataset (Fig 31), AL (−0.18) has probably the most important destructive impression, whereas AC (+0.14) has probably the most important constructive impression. AG (−0.00), SB (+0.01), and SK (−0.01) minimally affect prediction.
For the stacking mannequin (Fig 32), CD (−0.27) primarily signifies the affected person doesn’t have lung most cancers. FT (−0.22), AL (−0.19), PP (−0.18), CO (−0.16), and WZ (−0.14) are important destructive determinants. AC (+0.20) and SD (+0.18) are the strongest constructive determinants for lung most cancers. Other constructive determinants embody AT (+0.12), YF (+0.12), SK (+0.10), and CP (+0.08). An identical sample seems within the balanced dataset (Fig 33). In the upsampled dataset (Fig 34), AL (−0.23) has probably the most important destructive impression, whereas SD (−0.21) has probably the most important constructive impression. AT (−0.01), SB (+0.01), CP (+0.02), SK (+0.02), and CO (0.02) have a minimal affect on the prediction.
The improvement of interpretable ensemble fashions for lung most cancers prediction utilizing life-style and scientific indicators carries essential scientific implications, notably for early detection and danger stratification. The experimental outcomes point out that ensemble approaches—particularly stacking and voting—persistently outperform particular person base learners. While efficiency improved considerably on the balanced and upsampled datasets, these outcomes needs to be interpreted with warning, as artificial knowledge augmentation could not absolutely seize the complexity and variability of real-world scientific populations. In this context, the stacking mannequin’s efficiency on the unique dataset (93.53% accuracy) gives a extra sensible estimate of its sensible utility. The near-perfect accuracy noticed on the upsampled dataset demonstrates the mannequin’s skill to take advantage of further coaching info however shouldn’t be seen as consultant of anticipated scientific efficiency.
Although the stacking mannequin achieves the very best total efficiency throughout a number of analysis metrics, its recall is marginally decrease than that of some particular person base fashions. This highlights an inherent trade-off in classification, the place optimizing for total efficiency could result in slight reductions in sensitivity. In the context of lung most cancers prediction, recall is especially important as a result of false negatives can delay analysis and therapy. Therefore, whereas the stacking mannequin gives balanced and strong efficiency, mannequin choice in scientific settings could require prioritizing recall relying on the particular software.
To additional study mannequin behaviour, studying curve evaluation was carried out throughout totally different dataset configurations. On each the unique and SMOTE-balanced datasets, the voting and stacking fashions exhibit carefully aligned coaching and validation curves, indicating steady studying dynamics and good generalization. However, a divergence emerges within the absolutely upsampled dataset: the voting mannequin exhibits a transparent hole between coaching and validation efficiency, suggesting susceptibility to overfitting with artificial knowledge. In distinction, the stacking mannequin maintains shut alignment between the curves, demonstrating higher robustness and resistance to overfitting below the identical circumstances.
These observations are additional supported by stratified cross-validation, which helps mitigate variance in efficiency estimates. Notably, the stacking mannequin achieves persistently sturdy outcomes throughout all dataset variants, indicating that its efficiency isn’t solely depending on artificial augmentation however displays a extra steady and generalizable studying functionality.
The statistical outcomes point out that the stacking mannequin achieves probably the most constant and highest total efficiency, with important enhancements over a number of baseline strategies. However, its efficiency is statistically similar to sturdy tree-based fashions similar to GB and RF. From a scientific perspective, this distinction is essential. The statistical superiority of stacking over weaker fashions helps its use to enhance reliability, notably by lowering variability throughout affected person instances. This is efficacious in screening or triage settings, the place constant identification of high-risk people is important.
At the identical time, the dearth of statistically important variations between stacking and top-performing tree-based fashions means that comparable predictive efficiency may be achieved with much less complicated fashions. In sensible phrases, which means in settings the place interpretability, ease of deployment, or computational effectivity are priorities, fashions similar to RF or GB could function efficient alternate options. Conversely, in eventualities the place robustness throughout heterogeneous knowledge is crucial—similar to multi-center screening applications or populations with various scientific profiles—the stacking mannequin affords a bonus by integrating numerous studying patterns. This can assist keep steady efficiency when knowledge traits shift.
Through complete SHAP and LIME analyses, we recognized a number of clinically related predictors that each verify and problem the established understanding of lung most cancers danger components. A deep dive into the characteristic significance utilizing SHAP reveals an enchanting alignment with, and some shocking deviations from, established scientific data. The constant identification of fatigue (FT) and alcohol consumption (AC) as high predictors strongly helps the mannequin’s scientific validity. Fatigue is a well-documented constitutional symptom of malignancy, usually linked to systemic irritation and tumor metabolic exercise [53]. However, fatigue is a nonspecific and prevalent symptom in lots of persistent circumstances, suggesting that its predictive energy is strongest when mixed with different indicators, similar to smoking historical past or age [54]. Similarly, persistent alcohol use is a identified, dose-dependent danger issue that may promote carcinogenesis [55,56]. The mannequin’s skill to seize the dose-response relationship, through which larger values push SHAP values constructive (elevated danger), mirrors scientific observations and reinforces confidence within the mannequin’s studying course of. However, epidemiological proof on alcohol as an unbiased danger issue for lung most cancers is combined, with some research suggesting a synergistic impact with smoking attributable to enhanced carcinogen publicity (e.g., acetaldehyde) [57,58]. While smoking stays probably the most established danger issue, our mannequin attributes substantial predictive weight to alcohol consumption. This discovering warrants additional investigation—whether or not the mannequin captures true organic danger or confounding behavioural patterns (e.g., heavy drinkers being extra more likely to smoke). Prospective cohort research might make clear this relationship earlier than scientific integration. Furthermore, the mannequin’s emphasis on swallowing issue (SD) is of excessive scientific significance. Dysphagia is a major “red flag” symptom, usually indicating native tumor development the place the mass could also be compressing the esophagus, suggesting a extra superior stage [59].
However, the evaluation additionally highlights key factors for dialogue, notably relating to options that both battle with or add nuance to established danger hierarchies. One notable commentary is the comparatively low significance attributed to smoking (SK), regardless of its well-established function as the first danger issue for lung most cancers [60]. As proven within the characteristic significance analyses offered in Section 5.3.5, this sample was persistently noticed throughout each voting and stacking fashions and throughout all three dataset variants.
To additional examine this obvious discrepancy, we carried out a multicollinearity evaluation utilizing the Variance Inflation Factor (VIF). For a predictor
the place
The outcomes of VIF evaluation are offered in Table 15. Smoking (VIF = 11.16) and several other smoking-associated predictors, together with yellow fingers (VIF = 19.08), persistent cough (VIF = 17.47), shortness of breath (VIF = 18.78), and alcohol consumption (VIF = 17.79), exhibited substantial multicollinearity. This means that the predictive sign related to smoking is distributed throughout correlated behavioural and symptom-related variables. Consequently, attribution strategies similar to SHAP could allocate significance throughout these associated predictors fairly than concentrating it solely on smoking. Therefore, the comparatively decrease significance assigned to smoking shouldn’t be interpreted as contradicting established epidemiological proof, however fairly as a consequence of multicollinearity throughout the dataset.
An identical sample is noticed for different clinically established components. The comparatively low significance assigned to age (AG) and chest ache (CP) can also be counterintuitive, on condition that superior age is a significant non-modifiable danger issue for most cancers [61], and chest ache is a generally reported scientific symptom in lung most cancers sufferers [62]. These observations point out that the mannequin could also be capturing extra instant, discriminative indicators—similar to fatigue, persistent illness, and swallowing issue—over broader or much less particular indicators. Additionally, dataset-specific traits, similar to class distribution and illustration bias (e.g., aged sufferers with non-malignant respiratory circumstances), could additional affect these patterns.
Importantly, these observations spotlight a basic distinction between predictive modeling and epidemiological causality. SHAP-based characteristic significance displays the contribution of options to the mannequin’s predictions primarily based on the out there knowledge distribution, fairly than their true causal function in illness improvement. As a end result, options which are strongly correlated with the result within the dataset—or that seize downstream results—could obtain larger significance than main danger components. Therefore, deviations from established scientific hierarchies needs to be interpreted as insights into the mannequin’s discovered representations and dataset traits, fairly than contradictions of medical data.
Another intriguing high-impact characteristic is allergy (AL). The relationship between allergy symptoms and lung most cancers is complicated and never absolutely understood. Some analysis suggests {that a} hyper-vigilant immune system in allergic people would possibly supply a protecting impact in opposition to sure cancers, together with lung most cancers [63], though the connection varies relying on the particular allergy and most cancers kind [64]. The mannequin’s constant recognition of this characteristic suggests it has recognized a powerful predictive sample within the knowledge, warranting additional scientific investigation.
Other clinically related options, similar to persistent illness (CD), show notable predictive contributions, with biologically believable mechanisms involving inflammatory pathways. It is smart, as pre-existing circumstances, notably respiratory ones, are identified to extend lung most cancers danger [65]. The minimal affect of gender (GD) contradicts identified epidemiological patterns [66,67], probably reflecting altering smoking demographics or dataset limitations. These findings emphasise the significance of contextual interpretation when making use of mannequin outcomes to scientific decision-making. Future iterations might enhance these options by integrating extra detailed genetic or biomarker knowledge to extend precision.
The scientific implications of those findings might be important. The fashions would possibly enhance danger stratification by figuring out high-risk people past present smoking-based standards, probably broadening eligibility for low-dose CT screening. Including signs like fatigue and coughing could permit earlier detection than conventional strategies, whereas highlighting modifiable danger components, similar to alcohol consumption, might information preventive counselling.
The native, patient-specific explanations supplied by LIME additional improve the mannequin’s scientific utility. By displaying how a mixture of things, like excessive alcohol consumption and fatigue, contributes to a person’s high-risk rating, the mannequin gives actionable insights. This can empower physicians throughout affected person consultations, enabling them to have extra knowledgeable discussions about particular life-style modifications and the concrete components that elevate the affected person’s private danger. Ultimately, whereas the stacking mannequin demonstrates highly effective predictive capabilities, its true scientific promise is unlocked by its interpretability, which gives the essential bridge between statistical predictions and significant, actionable scientific insights.
Overall, from a scientific workflow perspective, the interpretability provided by SHAP and LIME can play a sensible function in decision-making. As illustrated in Fig 35, affected person demographic info, life-style components, and scientific signs may be processed by the proposed stacking mannequin to generate each a lung most cancers danger rating and SHAP/LIME-based explanations highlighting probably the most influential predictors. Global explanations can help clinicians in figuring out high-risk profiles primarily based on combos of life-style and scientific components, probably supporting early screening choices similar to prioritising sufferers for additional diagnostic analysis (e.g., low-dose CT scans). At the person degree, LIME-based explanations present patient-specific insights, serving to clinicians perceive why a selected affected person is flagged as excessive danger and guiding focused follow-up investigations. The ensuing danger rating and explanatory components can then be reviewed alongside typical scientific findings to assist risk-stratified choices, together with routine monitoring, further scientific evaluation, or specialist referral. Furthermore, the identification of modifiable contributors, similar to alcohol consumption, could assist personalised counselling and preventive interventions. Importantly, these explanations additionally permit clinicians to evaluate whether or not mannequin predictions are clinically believable, thereby reinforcing belief and enabling knowledgeable use of AI-assisted suggestions.
Fig 35. Conceptual decision-support workflow illustrating the potential integration of the proposed explainable lung cancer prediction framework into future clinical practice following appropriate validation.
Overall, from a scientific workflow perspective, the interpretability provided by SHAP and LIME can play a sensible function in decision-making. Fig 35 illustrates a conceptual decision-support workflow meant to show how the proposed framework might be built-in right into a future scientific setting following applicable validation. Patient demographic info, life-style components, and scientific signs may be processed by the proposed stacking mannequin to generate each a lung most cancers danger rating and SHAP/LIME-based explanations highlighting probably the most influential predictors. Global explanations can help clinicians in figuring out high-risk profiles primarily based on combos of life-style and scientific components, probably informing choices relating to additional diagnostic analysis (e.g., consideration of low-dose CT screening). At the person degree, LIME-based explanations present patient-specific insights, serving to clinicians perceive why a selected affected person is flagged as excessive danger and guiding focused follow-up investigations. The ensuing danger rating and explanatory components can then be reviewed alongside typical scientific findings to assist danger stratification, together with routine monitoring, further scientific evaluation, or specialist referral. Furthermore, the identification of modifiable contributors, similar to alcohol consumption, could assist personalised counselling and preventive interventions. Importantly, these explanations additionally permit clinicians to evaluate whether or not mannequin predictions are clinically believable, thereby reinforcing belief and enabling knowledgeable use of AI-assisted suggestions. The proposed workflow is meant solely as a conceptual illustration of how explainable AI might assist scientific decision-making and shouldn’t be interpreted as a advice for scientific deployment or referral choices primarily based on the current research alone.
However, implementation challenges warrant cautious consideration, together with the subjective nature of key predictors similar to fatigue and anxiousness, which can necessitate using standardised evaluation instruments to make sure dependable scientific measurement. Although the surprising deviations don’t invalidate the prediction, they spotlight an essential implication: a clinician utilizing this instrument should interpret the output completely. The mannequin signifies the presence of a “symptom,” and the clinician should hyperlink it to the underlying “cause.” This emphasises that the instrument is meant to assist, not substitute, scientific judgment.
From a deployment perspective, there’s an inherent trade-off between predictive efficiency and computational complexity. The proposed stacking ensemble achieves the strongest predictive efficiency however requires higher computational assets than particular person base learners attributable to its multi-stage structure and meta-learning course of. While this extra complexity is unlikely to pose a major problem in most scientific environments, the place danger prediction is carried out offline and doesn’t require real-time response, it could turn out to be an essential consideration in large-scale inhabitants screening programs or resource-constrained healthcare settings. Therefore, the selection of mannequin ought to stability predictive accuracy with out there computational assets and deployment necessities.
Improving characteristic measurement with extra exact knowledge on key variables, similar to smoking (pack-years) and alcohol consumption (normal drinks), might improve predictive accuracy. Combining this with imaging outcomes and molecular markers could result in the event of multimodal prediction instruments, whereas causal evaluation can make clear whether or not the recognized associations replicate real organic relationships. These developments will likely be important for translating mannequin outputs into actionable scientific insights.
Lung most cancers stays one of many deadliest ailments worldwide. Early prediction and detection are important for well timed medical intervention and enhancing affected person outcomes. Machine studying strategies may be utilized for this objective. This research adopted ensemble studying to foretell lung most cancers. Four main ensemble approaches (boosting, bagging, voting, and stacking) have been designed and evaluated for lung most cancers prediction utilizing a life-style dataset. The outcomes confirmed that ensemble fashions outperformed particular person base learners, with the stacking ensemble attaining the strongest total efficiency. On the unique dataset, the stacking mannequin attained an accuracy of 93.53%, demonstrating its potential for lung most cancers danger prediction utilizing life-style and scientific indicators. Although considerably larger efficiency was noticed on the balanced and upsampled datasets; nonetheless, these outcomes needs to be interpreted within the context of artificial knowledge augmentation. Overall, the findings show that ensemble studying can successfully leverage complementary predictive patterns to enhance lung most cancers danger evaluation and assist early detection efforts. When decoding fashions utilizing XAI, in line with the SHAP methodology, options similar to allergy, alcohol consumption, fatigue, persistent illness, and peer strain for smoking and alcohol most importantly affect lung most cancers danger prediction, whereas gender, age, and shortness of breath have the least affect. The findings assist the event of focused interventions and danger administration methods to enhance affected person outcomes in lung most cancers analysis and therapy.
This research has a number of limitations that warrant cautious consideration. A main limitation is the reliance on a single publicly out there Kaggle dataset comprising 309 samples, which inherently restricts the generalizability of the findings. The mannequin’s efficiency could subsequently replicate traits which are particular to this dataset and its underlying demographic and scientific distribution fairly than broadly consultant patterns. The comparatively small pattern dimension additional will increase the chance of overfitting and studying dataset-specific relationships. Additionally, the absence of exterior validation limits the robustness of the outcomes, because the proposed fashions weren’t evaluated on unbiased datasets collected from totally different establishments or affected person populations. Consequently, the findings needs to be interpreted as preliminary, and the proposed framework shouldn’t be used to information scientific choices or low-dose CT referral till potential and exterior validation has been accomplished.
Furthermore, the dataset could not seize the total spectrum of life-style and scientific danger components related to lung most cancers. Real-world affected person populations are extremely heterogeneous, influenced by comorbidities, environmental exposures, genetic predispositions, and different components that aren’t absolutely represented within the present knowledge, probably limiting the mannequin’s applicability in sensible settings. The use of artificial knowledge augmentation introduces further limitations. The technique of upsampling each minority and majority lessons to considerably enlarge the dataset is unconventional and should distort real-world class distributions, probably resulting in overly optimistic efficiency estimates. Accordingly, the outcomes obtained on the balanced and, particularly, the upsampled datasets needs to be interpreted as comparative analyses below artificial knowledge augmentation fairly than as estimates of anticipated real-world scientific efficiency. In addition, making use of resampling previous to cross-validation could introduce bias in efficiency analysis. Specifically, utilizing SMOTE earlier than cross-validation may cause artificial samples derived from the identical authentic cases to seem throughout coaching and validation folds, introducing info leakage and probably inflating efficiency estimates.
To construct upon this work and handle these limitations, a number of enhancements may be explored. The most important subsequent step is to guage the ensemble fashions on bigger, extra numerous, and externally validated datasets. This will contain sourcing knowledge from a number of establishments to evaluate the generalizability and robustness of our method throughout totally different populations and scientific settings. Such validation must also embody a broader vary of life-style and scientific traits to create a extra complete and correct predictive mannequin. Further validation experiments might discover the mannequin’s utility in multi-cancer research. Assessing the efficiency of the proposed ensemble framework throughout different most cancers varieties would assist confirm its generalizability past lung most cancers and supply insights into its broader applicability in oncology. Methodologically, integrating our ensemble strategies with deep studying and superior characteristic engineering could additional increase predictive efficiency. Investigating novel hybrid architectures that synergistically mix these approaches represents a promising route for enhancing predictive accuracy. Ultimately, the scientific worth of this analysis lies in its sensible software. Developing a user-friendly choice assist instrument that includes ensemble fashions might facilitate the sensible software of lung most cancers danger evaluation and early detection in scientific settings. Translating these fashions right into a sensible instrument with a physician-friendly interface is a important step for profitable scientific deployment, facilitating real-time danger evaluation and supporting early detection methods for lung most cancers.
Additionally, implementing bias mitigation methods—similar to analysing underrepresented populations—will likely be essential to enhancing mannequin equity and making certain equitable healthcare outcomes. While SHAP elucidated characteristic significance, the research didn’t completely discover the interpretability of ensemble fashions. Understanding the decision-making mechanisms of those fashions is crucial for his or her implementation in healthcare, the place belief and transparency are paramount. Future analysis might examine mannequin interpretability by means of numerous frameworks to make clear relationships between enter knowledge and predictions. Incorporating multi-modal knowledge, similar to imaging knowledge with scientific indicators, might improve the predictive functionality and scientific relevance of fashions.
Future enhancements contain broadening the scope to include further base learners and hybrid fashions that combine ensemble strategies with deep studying to reinforce predictive accuracy. Examining demographic variety inside databases will likely be essential to making sure equitable healthcare outcomes. Integrating a number of knowledge sources can assist mitigate biases and enhance the mannequin’s skill to establish people in danger. The inclusion of further danger components, similar to genetic predispositions and environmental influences, in ensemble fashions warrants additional investigation. Expanding the set of assessed traits might improve their predictive energy and relevance to personalised therapy.
To progress from analysis to scientific use, potential validation throughout a number of centres is essential. Including causal inference strategies, similar to counterfactual evaluation, can assist decide whether or not important components, similar to alcohol consumption, are causally associated or merely correlated. Additionally, combining imaging and biomarker knowledge might enhance predictive accuracy, resulting in a multi-modal AI diagnostic system. Lastly, creating an interpretable decision-support instrument utilizing ensemble fashions might improve real-world usefulness in healthcare. This know-how would help healthcare professionals in assessing danger and early detection of lung most cancers, finally enhancing affected person outcomes and healthcare companies.
The authors thank the UTHealth Cancer Genomics Core for offering technical assist. The authors additionally gratefully acknowledge the assist supplied by the Deanship of Scientific Research, Vice Presidency for Graduate Studies and Scientific Research, King Faisal University, Saudi Arabia.
This web page was created programmatically, to learn the article in its authentic location you’ll be able to go to the hyperlink bellow:
https://journals.plos.org/plosone/article%3Fid%3D10.1371/journal.pone.0357291
and if you wish to take away this text from our website please contact us
This web page was created programmatically, to learn the article in its unique location you…
This web page was created programmatically, to learn the article in its authentic location you…
This web page was created programmatically, to learn the article in its authentic location you…
This web page was created programmatically, to learn the article in its authentic location you…
This web page was created programmatically, to learn the article in its authentic location you…
This web page was created programmatically, to learn the article in its authentic location you'll…