Abstract:
The challenges posed by lung cancer are immense. It is one of the most serious diseases, health-wise, due to various factors affecting patient outcomes, such as demographic, environmental, behavioral, treatment-related, and clinical factors. With the increasing availability of multidimensional patient data, advanced machine learning technologies can be used to predict prognostic factors and identify determinants of survival. The primary goal of this study was to analyze the existing lung cancer dataset, evaluate several machine learning algorithms for predicting patient outcomes, and identify the most significant predictors. In this study, “Survived” served as the main outcome variable, while “Survival_months” was excluded from the training and testing of models to avoid data leakage. Models including Logistic regression, Random Forest, ExtraTrees, Gradient Boosting, and XGBoost were developed on the basis of a stratified data split with the ratio of 80/20. The performance evaluation of the models was conducted before calculating accuracy, precision, recall, F1-score, as well as the area under the ROC curve. Permutation importance enabled finding out the most significant predictors. The study covered 2,000 patient records and 41 factors. No missing data was found. Overall, the survival rate was 37.8%. Survival was not the same across stages. In Stage I, it was 69.8%. In Stage IV, it dropped to 5.5%.The Gradient Boosting model performed best. It reached about 70% accuracy. Its ROC-AUC was 0.786.The cancer stage was confirmed as the most significant predictor, with the highest importance value compared with other variables. The dataset demonstrates that machine learning methods can be applied to lung cancer survival predictions and that cancer stage is the most prominent variable.
Keywords: Lung Cancer; Machine Learning; Survival Prediction; Global Health; Prognostic Modeling
