SACM - Australia
Permanent URI for this collectionhttps://drepo.sdl.edu.sa/handle/20.500.14154/9648
Browse
19 results
Search Results
Item Restricted An Optimised Hybrid Framework for Predicting Learner Dropout in Massive Open Online Courses Using Machine Learning and Behavioural Analytics(Saudi Digital Library, 2026) Alghamdi, Saad Hamdan A; Ben, SohThis research addresses the persistent challenge of high dropout rates in Massive Open Online Courses (MOOCs) by developing and evaluating an enhanced predictive model capable of identifying learners at risk of disengagement. The study introduces the Integrated Stacked Ensemble Learning Dropout Prediction (ISELDP) model, which combines machine learning and meta-heuristic optimisation to improve predictive performance and interpretability. Employing an explanatory, sequential quantitative methodology, the research was conducted in two phases: (1) developing and testing the proposed model using a secondary in-session dataset, and (2) validating the model with primary behavioural and demographic data collected from 466 MOOC participants through a structured survey. The proposed ISELDP integrates multiple base learners (random forest, AdaBoost, gradient boosting, and XGBoost) with a multilayer perceptron (MLP) meta-learner, supported by an optimised feature selection framework based on a hybrid genetic algorithm–correlation feature selection (GA–CFS) technique. The evaluation across balanced and unbalanced datasets demonstrates that the enhanced ISELDP significantly outperforms benchmark models, achieving an average accuracy of 91% and an F1-score of 88%. The hybrid optimisation process effectively reduces feature redundancy, mitigates overfitting, and enhances the model’s generalisability across heterogeneous data. Complementary statistical analyses of the survey data identified motivation and engagement as the most influential predictors of MOOC retention, mediated by self-management and social influence factors, and indirectly affected by course design. These findings emphasise the centrality of motivational and behavioural dimensions in sustaining learner persistence, offering a broader understanding of dropout triggers beyond in-session activity metrics. The research contributes theoretically by defining the specifications of effective dropout prediction models that integrate deep learning and meta-heuristic optimisation. Methodologically, it advances predictive learning analytics through a validated hybrid framework that balances accuracy, interpretability, and adaptability in dynamic learning environments. Practically, it provides a data-driven foundation for designing proactive retention strategies and personalised interventions in MOOCs. Despite limitations related to data scope, computational complexity, and temporal coverage, the findings establish a replicable and extensible model for future cross-platform and longitudinal studies in predictive educational analytics. Compared to state-of-the-art approaches, this study proposes a hybrid GA–CFS feature selection method embedded in a stacked ensemble framework. The proposed approach yields a more compact feature set while enhancing prediction accuracy and robustness compared to conventional single-method techniques.9 0Item Restricted Quantifying the Human Impact on Vegetation Health in Rawdat Al Khafs, Saudi Arabia, Using Multi-Source Geospatial Data and Machine Learning(Saudi Digital Library, 2026) Alajmi, Abdulhadi; Dewan, AshrafArid and hyper-arid environments are vulnerable to land degradation and have become focal targets for large-scale afforestation initiatives combating desertification. The One Million Tree Initiative at Rawdat Al Khafs, a hyper-arid rawdah ~100 km northeast of Riyadh receiving < 50 mm annual rainfall, operates under the Saudi Green Initiative and Vision 2030. Evaluating afforestation effectiveness in such settings remains methodologically challenging, as existing frameworks rarely separate climatedriven from human-induced vegetation change in hyper-arid conditions. This study addresses these gaps through three inter-connected objectives: (i) establishing a climate-driven vegetation baseline using a Random Forest (RF) regression model trained on pre-plantation Sentinel-2 imagery and multi-source climate data (20182021); (ii) detecting and quantifying residual vegetation trends (2018–2025) that exceed this baseline and may represent a non-climatic signal conditionally attributable to afforestation; and (iii) characterising the spatial fingerprint of the detected anomaly through phase-based proximity analysis and distance-decay modelling. The RF–RESTREND framework combines 10 m Sentinel-2 imagery, CHIRPS precipitation, and ERA5-Land temperature data with Mann–Kendall trend testing, Sen's slope, Difference-in-Differences (DiD) correction, and SHAP analysis, applied to both NDVI and EVI. A persistent positive vegetation anomaly was observed within the plantation boundary from mid-2022 onwards, consistent across both indices and not fully explained by the climate counterfactual. Within the study area, 77.6% of NDVI pixels and 81.7% of EVI pixels exhibited statistically significant greening (Mann–Kendall Z > +1.96), with mean Z-scores of +3.97 and +4.45, respectively. The EVI distance-decay model (R² = 0.53; decay length = 65 m) indicated a localised boundary effect, while SHAP analysis identified spatial position and long-term climatic trends as dominant predictors, consistent with micro-topographic moistureconcentration mechanisms in rawdah ecosystems. Attribution is framed as conditional, reflecting the short pre-plantation training window and the widening climate variability of the post-plantation period. These findings provide an early satellitebased, climate-corrected assessment of an ecological response associated with the One Million Tree Initiative. The RF–RESTREND framework is reproducible and scalable to comparable arid restoration contexts, supporting Vision 2030 monitoring.17 0Item Restricted Evaluating Hybrid AI Approaches in email Spam Detection: A Literature Review(Saudi Digital Library, 2026) Alshamrani, Husam A; Sabar, NasserSpam is still one of the most serious cybersecurity issues of the day, and serves as a method of delivering phishing and malware. In this thesis, two complementary parts are integrated: First, a PRISMA-based systematic review of 55 studies that were published from 2023 to 2025 and analyze the classical, AI-based and ensemble spam detection techniques; and second, a controlled proof-of-concept experiment on the Enron-Spam corpus, where three distinct classical classifiers (Naïve Bayes with Bag-of-Words, Logistic Regression with TF-IDF, Linear SVM with TF-IDF) are compared against a stacking ensemble of the three classifiers with a Logistic Regression meta-classifier. All four configurations were very close together and obtained an accuracy of between 0.989 and 0.993, with the stacking ensemble having the highest F1-score (0.9923) and the lowest false-positive rate (0.0088); but there were no significant differences between the four configurations when a paired statistical test was not used. The thesis contribution lies in the synthesis of the various perspectives of the 2023–2025 hybrid and ensemble approaches and an internally consistent baseline comparison on a standard corpus. The study is restricted to English-language email and to lexical features; further extensions of the study using deep-learning, multilingual, and adversarial techniques are suggested.10 0Item Restricted Harmonic Impact Assessment and Prediction in a Power System with Nonlinear Loads(Saudi Digital Library, 2025) Almogargesh, Yousef Yaqoob Y; Swamidoss, SathiakumarThis thesis presents a simulation-based and machine learning framework for predicting power quality (PQ) indices in residential power systems with nonlinear loads. Four MATLAB/Simulink models representing nonlinear loads, induction motor loads, thyristor converters, and photovoltaic grid-connected systems were developed to generate voltage and current waveforms under different operating conditions. Time-domain and frequency-domain features were extracted from the simulated signals and used to train artificial neural network (ANN) models for predicting six key PQ indices: Total Harmonic Distortion (THD), Crest Factor, Form Factor, Power Factor, RMS Voltage, and Peak Voltage. Both shallow and deep neural network architectures were implemented and evaluated using statistical performance metrics including Mean Squared Error (MSE), Mean Absolute Error (MAE), and Pearson correlation coefficient. The results demonstrate that the proposed approach accurately predicts future power quality conditions while reducing the need for repeated analytical calculations. This framework provides an efficient tool for power quality monitoring and supports the development of intelligent and proactive power system management strategies.21 0Item Restricted Artificial Intelligence (AI)-based Multi-criteria Shipping Industry Provider Selection(Saudi Digital Library, 2026) Khan, Ibraheem Abdulhafiz Q; Hussain, Farookh KhadeerThis thesis highlights that automating the selection of maritime shipping service providers is pivotal to supply-chain performance. By replacing fragmented and subjective practices with transparent analytics, automation reduces cost and time, improves reliability, and ensures decisions are reproducible at scale. To achieve this, the thesis introduces an intelligent multi-criteria search engine (MC-SE) that integrates artificial intelligence (AI) and multi-criteria decision-making (MCDM) to support both shippers and freight companies in identifying reliable, cost-effective providers. The objectives are to (i) develop an AI-based predictive classifier for offshore shipping decisions; (ii) systematically map provider criteria to the service quality framework (SERVQUAL); (iii) propose an AI-assisted approach for criteria weighting; (iv) conduct a SERVQUAL survey for provider-side assessment; and (v) validate the framework through an Australian case study. Methodologically, criteria were extracted from provider websites and benchmark datasets, then clustered into decision attributes using semantic similarity techniques. These clusters were aligned with SERVQUAL dimensions to ensure construct validity. AI-based weighting and supervised learning were applied within an MCDM pipeline to calculate attribute importance, integrate cost as a complementary decision factor, and rank providers objectively. This dual use of structured datasets and unstructured textual content ensures that the framework adapts to both traditional logistics data and dynamic, web-based information sources. Provider-side service quality is structured via SERVQUAL, while cost is modelled as a complementary decision attribute within the overall MC-SE multi-criteria framework. Validation demonstrates strong agreement between the proposed MC-SE weighting and the SERVQUAL survey (mean absolute error (MAE), MAE = 0.014), with dimension-level differences typically within 2–3%. The optimisation classifier, based on a voting ensemble, achieves 82.3% accuracy on held-out test data. These findings show that data-driven weighting, combined with supervised learning, can robustly support provider selection in practice. Overall, this thesis develops a novel, AI-driven framework to support the automated selection of maritime shipping service providers, bridging gaps between academic models and industry practice. Future research will refine the MC-SE framework, evaluate its portability across diverse contexts, and extend the evaluation to incorporate customer-experience evidence that complements provider-side quality and explicit cost trade-offs. Importantly, providers’ clusters are consistently aligned with canonical SERVQUAL dimensions to preserve theoretical and empirical coherence.29 0Item Restricted Integrating Educational Data Mining and Artificial Intelligence to Enhance ICT User Satisfaction and Administrative Efficiency in Saudi Educational Institutions(Saudi Digital Library, 2026) Almaghrabi, Hamad; Soh, BenThe integration of Information and Communication Technology (ICT) in educational administration offers transformative opportunities to enhance efficiency and user satisfaction, but also presents significant challenges. Despite the potential of ICT systems to stream- line processes and support data-driven decision-making, their implementation is often hindered by fragmented infrastructures, inconsistent adoption, and limited alignment with user needs. This thesis addresses these challenges through the design and evaluation of the AI-integrated IiCE framework, developed to strengthen ICT adoption and administrative performance in educational institutions. Educational administrative environments are inherently complex, characterised by mul- tidimensional data, dynamic workflows, and overlapping responsibilities that often expose systemic inefficiencies. The proposed IiCE framework leverages predictive analytics and user-centred design principles to generate actionable insights for optimising ICT utilisa- tion. Its key objectives include identifying the determinants of user satisfaction, enhancing decision-making processes, and fostering an organisational culture that supports technolo- gical innovation and acceptance. Employing a mixed-methods research approach, this study investigates current ICT ad- option practices in Saudi educational institutions. Quantitative and qualitative analyses, incorporating stakeholder perceptions and institutional data, were conducted to uncover adoption barriers and performance gaps. Machine learning (ML) models were applied to predict user satisfaction trends, while SHAP (Shapley Additive Explanations) techniques provided interpretability by highlighting the most influential factors affecting adoption. The framework also integrates adaptive training modules, modular deployment strategies, and continuous feedback mechanisms to ensure sustainability and contextual adaptability. Grounded in Saudi Arabia’s Vision 2030 for digital transformation, the evaluation of the IiCE framework demonstrates its ability to enhance administrative workflows, optim- ise resource allocation, and strengthen stakeholder engagement. Expert validation con- firms its effectiveness in mitigating inefficiencies, promoting collaboration, and supporting evidence-based management practices. This research contributes to the fields of educational administration and ICT innova- tion by presenting an adaptable, AI-driven framework that bridges the gap between tech- nological potential and practical implementation. The findings underscore the value of advanced AI techniques in managing ICT complexity, driving user satisfaction, and im- proving institutional efficiency. Future work may extend this framework through real-time analytics, greater model interpretability, and cross-domain applications for broader educational impact22 0Item Restricted Mortality and Prolonged ICU Stay Analysis in the MIMIC-III Database: A Dual Analytical(Saudi Digital Library, 2025) Aldawsari, Fayez; Concha, Oscar; Welberry, HeidiAbstract Background: Understanding mortality patterns across ICU types and identifying patients at risk for prolonged stays is crucial for resource allocation and clinical decision-making. Objectives: This study examined (1) differences in in-hospital mortality across ICU types after risk adjustment, and (2) developed a predictive model for prolonged ICU stays using early clinical data. Methods: Using MIMIC-III database, we analyzed 61,533 ICU admissions. Propensity score matching with logistic regression compared mortality across five ICU types against MICU. XGBoost classification predicted ICU stays ≥7 days using first 24-hour clinical features. Results: Overall mortality was 4.6%, with 16.1% prolonged stays. After propensity adjustment, CSRU demonstrated significantly lower mortality versus MICU (OR: 0.22, 95% CI: 0.16-0.29), while SICU, CCU, and TSICU showed no significant differences. The XGBoost model achieved excellent discrimination (AUC-ROC: 0.862, sensitivity: 0.264, specificity: 0.978). Glasgow Coma Scale mean score was the most important predictor, followed by vasopressor use and mechanical ventilation. Conclusions: Substantial mortality differences exist across ICU types after risk adjustment, with cardiac surgery patients showing superior outcomes. Early clinical data accurately identifies patients at risk for prolonged stays, enabling proactive resource planning.25 0Item Restricted Time-to-Death Analysis and Prolonged ICU Stay Prediction Using MIMIC-III Data(Saudi Digital Library, 2025) Alanazi, Abdullah; Welberry, HeidiThis study investigated two critical questions in intensive care using the MIMIC-III database. First, we examined whether time-to-death differs across ICU types using survival analysis methods. Second, we developed a machine learning model to predict prolonged ICU stays (≥7 days) from early clinical features, addressing the class imbalance inherent in this outcome. Our survival analysis of 24,754 adult ICU admissions revealed significant mortality differences between ICU types, with SICU and TSICU patients showing approximately 50% lower hazard of death compared to MICU patients (adjusted HR 0.51, 95% CI: 0.44-0.60 and 0.51, 95% CI: 0.42-0.62, respectively). For prolonged stay prediction, our Random Forest model with balanced training achieved strong discrimination (AUC-ROC 0.84) and balanced accuracy (76.6%), outperforming traditional logistic regression. The most important predictive features were Glasgow Coma Scale measures, mechanical ventilation, and vasopressor use—all indicators of illness severity. These findings suggest that ICU type substantially influences mortality risk and that early prediction of prolonged stays is feasible using routinely collected clinical data.30 0Item Restricted AI-Based Approaches for Respiratory Disease Detection Using Audio Signals and Imaging Data(Saudi Digital Library, 2025) Shati, Asmaa; Hassan, Ghulam Mubashar; Datta, AmitavaRespiratory diseases (RDs) remain major global health concerns, typically diagnosed through imaging and auscultation, with cough sounds also offering diagnostic cues. These methods, however, are often subjective and depend on expert interpretation. Advances in machine learning (ML) enable automated RD diagnosis, yet challenges such as limited data, high computational costs, and accessibility gaps persist, underscoring the need for innovative approaches. This thesis proposes a series of novel approaches for automated RD detection, utilizing either cough audio or CXR as input modalities, selected for their availability and affordability. These approaches integrate advanced techniques for segmentation, feature extraction, and subsequent classification, offering practical and cost-effective diagnostic solutions. Extensive evaluation on multiple open-source datasets demonstrates the effectiveness of the proposed approaches across diverse diagnostic contexts.30 0Item Restricted Predicting Client Default Payments Using Machine Learning in Production Environment(Saudi Digital Library, 2025) Alanazi, Reem; LavendiniThis project investigates the application of machine learning techniques to predict client default payments in a credit card setting. Using a dataset of 30,000 Taiwanese clients, the study addresses the challenges of class imbalance, predictive accuracy, and fairness in credit risk assessment. An XGBoost model was developed and enhanced through feature engineering, resampling techniques (SMOTE/ADASYN), and class weighting to improve recall for defaulters while maintaining overall accuracy. Interpretability was achieved using SHAP values, providing transparency into model decisions. To mitigate demographic disparities, particularly across education levels, a fairness-constrained Random Forest was integrated into a two-stage cascade framework, reducing false positives while preserving high recall. The final cascade model achieved 84% accuracy, with 93% recall for non-defaulters and 53% recall for defaulters, significantly outperforming baseline benchmarks. Fairness audits revealed that education-based disparities could be reduced with minimal performance trade-offs, while age-based fairness was largely maintained. The project demonstrates a practical, interpretable, and ethically aware pipeline for credit default prediction, with deployment considerations and directions for future research in cost-sensitive learning, advanced fairness constraints, and real-time monitoring35 0
