Saudi Cultural Missions Theses & Dissertations
Permanent URI for this communityhttps://drepo.sdl.edu.sa/handle/20.500.14154/10
Browse
43 results
Search Results
Item Restricted Intelligent Fault Detection for Belt Conveyor Idlers Using Machine Learning(Saudi Digital Library, 2026) Alharbi, Fahad; Luo, Suhuai; Zhang, Hongyu; Chen, Zhiyong; Wheeler, CraigConveyor belt systems are essential components of modern mining, logistics, and manufacturing operations. However, faults in idlers, which are the rollers that support and guide the belt, can reduce system reliability, cause unplanned downtime, increase maintenance costs, and create safety risks. Conventional inspection methods, such as visual assessment and manual listening, are widely used to identify abnormal idler behaviour, but they are labour intensive, subjective, and difficult to apply across large conveyor systems. Acoustic monitoring offers a promising contactless alternative because developing faults often produce changes in sound before complete failure occurs. However, these fault related acoustic signatures can be subtle and may be masked by environmental and operational noise, making them difficult to interpret using traditional signal processing methods alone. Machine learning (ML) has therefore been increasingly applied to acoustic signals to automate feature extraction, fault detection, and classification. Nevertheless, existing ML approaches still face several important challenges, including the limited availability of labelled fault data, inadequate generalisation across recording conditions and sensing platforms, and insufficient modelling of the spatial and temporal characteristics of acoustic signals. This thesis addresses these challenges by developing intelligent fault detection (IFD) models that automatically analyse acoustic signals from belt conveyor idlers. The research follows a progressive methodology comprising supervised transfer learning, semi-supervised anomaly detection, task-specific spatial–temporal deep learning, and cross-domain Convolutional Neural Network (CNN)–Transformer adaptation. To support these investigations, acoustic datasets were collected using handheld microphones and drone-mounted recorders under multiple idler operating and fault conditions. The first study investigated supervised fault classification using the original handheld acoustic dataset, which contained 255 four-second samples representing Normal, Stage 1, Stage 2, and Stage 3 operating conditions. Embeddings were extracted using YAMNet, a pre-trained audio neural network, and processed using Bidirectional Long Short-Term Memory (BiLSTM) and Bidirectional Gated Recurrent Unit (BiGRU) networks to model contextual relationships in both forward and backward temporal directions. Attention mechanisms and an Extreme Gradient Boosting (XGBoost) classifier were also evaluated. The best-performing YAMNet–BiLSTM configuration achieved an accuracy of 90.59% and an F1-score of 90.57% for four-class fault-stage classification. Although this study established a strong supervised baseline, it required labelled examples from every operating condition and fault stage. This dependence limits practical application because faulty-idler recordings are relatively rare, costly to collect, and difficult to label in industrial environments. To address this limitation, the second study developed CASSAD (Chroma-Augmented Semi-Supervised Anomaly Detection), which was trained exclusively on normal operating samples during training. Following the first study, additional handheld recordings were collected, increasing the dataset from 255 to 468 acoustic samples. This larger dataset enabled CASSAD to be evaluated using a broader set of normal and abnormal operating recordings. CASSAD combines Chroma-STFT, Chroma-CQT, and Chroma-CENS representations with filtering, statistical aggregation, and a one-class support vector machine (OC-SVM). On the expanded 468-sample handheld dataset, the best-performing CASSAD configuration achieved an accuracy of 90.59%, a positive-class F1-score of 92.42%, and an Area Under the Receiver Operating Characteristic Curve of 96.29%. To enable a direct comparison with the YAMNet-based models, CASSAD was also evaluated on the original 255-sample dataset after the three fault stages had been combined into a single Abnormal class. In this binary evaluation, CASSAD achieved an accuracy of 93.00% and an F1-score of 93.25%, compared with an accuracy of 92.18% and an F1-score of 93.00% for the strongest YAMNet-based configuration. These results demonstrate that competitive anomaly-detection performance can be achieved without labelled abnormal samples during training. However, CASSAD provides only binary Normal–Abnormal decisions and uses temporally aggregated features, limiting its ability to distinguish fault severity and model changes in acoustic behaviour over time. To overcome these limitations, the third study developed TD-CLNet, a Time-Distributed CNN–Long Short-Term Memory (LSTM) architecture designed to perform multi-stage fault classification while learning spatial and temporal representations from the acoustic data. The expanded 468-sample handheld dataset was re-segmented into one-second samples and converted into log-Mel-spectrogram frames. TD-CLNet applies a shared CNN feature extractor to each frame and then uses an LSTM to model the resulting temporal feature sequence. This design enables the model to distinguish among Normal, Stage 1, Stage 2, and Stage 3 conditions while preserving temporal information that was reduced through the statistical aggregation used in CASSAD. Under four-fold cross-validation, TD-CLNet achieved a mean accuracy, precision, recall, and weighted F1-score of 92.1%. It outperformed the evaluated conventional CNN–LSTM configurations and provided a small improvement over the strongest YAMNet model re-evaluated on the same expanded dataset. Nevertheless, the sequential processing used by LSTM networks limits parallel computation and provides less direct access to broader global relationships within the acoustic feature sequence. To address these limitations and investigate cross-domain generalisation, the fourth study developed hybrid CNN–Transformer models using four pre-trained CNN backbones: ResNet-18, DenseNet-121, EfficientNet-B0, and ShuffleNet-V2. Two acoustic feature representations were investigated: Mel-spectrograms, which represent the distribution of signal energy across perceptually scaled frequency bands over time, and Mel-Frequency Cepstral Coefficients (MFCCs), which provide a compact representation of the short-term spectral envelope. The CNN backbones extracted local spectral representations, while Transformer encoders modelled broader contextual relationships within the feature sequences. In the first phase, the models were trained and evaluated using 0.5-second segments derived from the handheld source-domain recordings. The ResNet-18–Transformer models achieved a cross-fold mean accuracy of 96.9%, with a 95% confidence interval of 96.4%–97.6%, while a ResNet-18–Transformer ensemble increased the handheld-domain accuracy to 98.0%. In the second phase, the models trained on the handheld source-domain dataset were adapted to the drone-acquired target-domain dataset. The drone recordings represented a more challenging sensing environment because they were affected by rotor noise, changing recording distances, varying microphone positions, and environmental interference. Multiple fine-tuning strategies were evaluated, included full fine-tuning (Full-FT), freezing the CNN backbone (CNN-Frozen), freezing the Transformer encoder (TR-Frozen), and 𝐿2-SP regularisation. Mean–Covariance Alignment (MCA) was compared with a cross-entropy (CE) baseline and several established domain-adaptation methods. MFCCs produced the strongest CE baseline in several drone-domain experiments, achieving an accuracy of 90.5% under 𝐿2-SP. However, MFCC performance generally decreased when MCA was applied. In contrast, MCA improved the Mel-spectrogram Full-FT configuration, increasing accuracy from 84.7% to 86.0% and the F1-score from 83.4% to 85.6%. These findings demonstrate that domain-adaptation performance depends on the interaction among the acoustic representation, fine-tuning strategy, model architecture, and alignment objective; no single adaptation method was uniformly optimal across all evaluated configurations. In summary, this thesis establishes a connected research pathway for acoustic fault detection in belt conveyor idlers. It provides handheld and drone-acquired acoustic datasets, establishes supervised transfer-learning baselines, introduces CASSAD for anomaly detection without labelled abnormal training samples, develops TD-CLNet for task-specific spatial–temporal fault-stage classification, and proposes a two-phase CNN–Transformer framework for handheld-to-drone domain adaptation. Collectively, these contributions advance acoustic idler monitoring towards more accurate, data-efficient, and adaptable fault detection while identifying the need for further validation across additional industrial sites, conveyor configurations, sensing platforms, and operating conditions before large-scale real-world deployment.23 0Item Restricted Low-Carbon Sustainable Construction in KSA(Saudi Digital Library, 2025) ALBALAWI, MOHAMMED AWAD; ElHalwany, Mohamed MThis research focuses on Saudi Arabia's energy-intensive sectors and presents a data-driven method for accurately forecasting CO₂ emissions from industrial facilities. By combining temporal variables, production metrics, energy consumption data, and operational efficiency indicators, the proposed model employs gradient boosting regression (XGBoost) to capture complex, non-linear connections between input features and emissions output. Compared to traditional statistical methods, this machine learning-based technique offers enhanced projected accuracy and responsiveness, making it suitable for real-time monitoring and decision-making. Besides emissions projections, the model incorporates an adaptive control component that facilitates near-real-time operational adjustments. These adjustments are enhanced by a reinforcement learning loop, which provides continuous optimization and ongoing learning from prior decisions and outcomes. By employing scenario simulation, stakeholders can assess financial implications and develop environmentally friendly policies. Experiments demonstrate that the proposed framework outperforms traditional approaches in prediction accuracy and operational flexibility, providing an effective strategy for sustainable industrial operations in Saudi Arabia16 0Item Restricted Happiness by Proxy: Real-Time Social Diagnostics from Online Search, Sentiment, population and Risk Behaviour(Saudi Digital Library, 2026) Alomair, Maryam Anwar; Boy, fredericTraditional indicators of population well-being, including GDP and survey-based subjective well being measures, face well-known limitations: GDP excludes the psychological, social, and cultural dimensions of human flourishing, while surveys are temporally lagged, expensive, and vulnerable to response biases. This thesis explores whether naturally occurring digital behavioural data — specifically Google Trends search-volume series — can provide complementary, real-time, population-level proxy indicators of subjective well-being that address some of these limitations. The thesis is situated within the subjective well-being (SWB) tradition, targeting its life-satisfaction and affect components (see 1.4.0). "Proxy" is used in a specific sense (1.4.0(b)): aggregate behavioural indicator, predictive feature, and correlational population-level signal — explicitly not substitute measure for individual well-being. All inferences are at the population or sub-population level. Five empirical studies are presented. Study 1 builds an aggregated index from emotion-related Google Trends keywords for the UK (2020–2023) and shows that it covaries with the ONS weekly life-satisfaction series (R² = 0.6535 in random-split test; chronological-split robustness check reported in 1.3.4). Study 2 compares interpretable machine-learning models for life-satisfaction prediction in a real-time monitoring framework intended as a proof of concept for inclusive digital governance, with portability to other settings flagged as a design intention requiring external validation. Study 3 uses LSTM networks to identify three distinct diurnal patterns in aggregate search behaviour, which we describe as "digital chronotypes" — a property of the population-level signal, not of individual users. Study 4 applies the BEAST change-point algorithm to Arabic language Google Trends data for Saudi Arabia (2013–2022), as an exploratory investigation of how aggregate digital signals respond to cultural transitions including Ramadan and Vision 2030 milestones. Study 5 uses the Culture-Based Development (CBD) framework with a stepwise econometric pipeline (Oaxaca-Blinder, OLS-FE, 3SLS, Heckman, hierarchical) to investigate the association between country-of-origin cultural milieu and observed driving behaviour in a multi national sample of drivers in Saudi Arabia, identifying an empirical pattern (the "Paradox of National Behavioural Patterns") whose substantive and statistical interpretations are discussed. The thesis offers partial empirical support for the use of digital behavioural traces as complementary population-level well-being proxies, indicative evidence that some PERMA dimensions leave detectable traces in aggregate digital behaviour, and proof-of-concept demonstrations of how such proxies might be applied in policy and risk-assessment contexts. It does not establish a unified, universally validated measurement system, and the contributions in each empirical chapter are framed accordingly.14 0Item Restricted The Impact Of Electronic Health Records And Machine Learning On Antibiotic Prescribing Practices In Haematology Patients With Bacteraemia.(Saudi Digital Library, 2025) Tashkandi, Smaher; Scott, McLachlanBackground: Bloodstream infections (BSI) are a leading cause of morbidity and mortality among haematology patients, particularly those immunocompromised due to chemotherapy-induced neutropenia. Delays in diagnosis and reliance on broad-spectrum antibiotics contribute to antimicrobial resistance (AMR), now recognised as a critical global health threat. Electronic Health Records (EHR) combined with Machine Learning (ML) offer potential to support early infection recognition and optimise antibiotic prescribing, yet their adoption in high-risk haematology populations remains underexplored. Aim: This dissertation explores how HER and ML-driven tools influence antibiotic prescribing practices for bacteraemia in haematology patients, with specific attention to their potential roles in antimicrobial stewardship, early detection, and safe clinical decision-making. Methods: A case study approach was adopted, drawing on Yin’s structured framework to examine evidence at the intersection of digital health and haematology care. A systematic search of PubMed, MEDLINE, and CINAHL identified 143 records, of which 11 primary studies met inclusion criteria. These studies were critically appraised using Joanna Briggs Institute (JBI) checklists. Thematic synthesis was employed to integrate findings across quantitative and qualitative evidence, with particular attention to methodological quality, relevance to practice, and stakeholder perspectives. Findings: Four themes were identified. First, early infection prediction: ML models consistently outperformed conventional scores (e.g., qSOFA, NEWS2), offering potential for earlier alerts. Second, outcome prognostication: algorithms improved risk stratification for mortality and antimicrobial resistance. Third, support for antibiotic selection and de-escalation: personalised antibiograms and interpretable resistance forecasts showed promise for AMS. Fourth, model and data design: most studies lacked haematology-specific variables such as duration of neutropenia, central venous access, or MDR colonisation, limiting generalisability. Across themes, facilitators and barriers shaped implementation: clinician trust, usability, workflow alignment, governance, and the need for continuous recalibration. Evidence quality overall was moderate, constrained by retrospective single-centre designs and scarce prospective validation. Conclusion: ML–EHR systems hold promise for enhancing infection recognition and optimising antibiotic use in haematology, but their translation into practice is hindered by methodological weaknesses and implementation barriers. Future progress requires prospective, haematology-specific trials, incorporation of relevant risk factors, and development of governance frameworks. Embedding nurses, pharmacists, and wider stakeholders within stewardship structures will be essential to ensure safe, accountable, and effective deployment.4 0Item Restricted A Comparative Analysis of ARIMAX and LSTM for Predicting TSLA Stock Prices Using S&P 500 and VIX Data(Saudi Digital Library, 2025) Alageli, Abdulmajeed; Nishanth, SastryForecasting stock prices is a vital and complex task within the realm of financial research, a challenge that is further exacerbated by rising volatility so typical of modern markets. While statistical models have been employed for a long time, machine-learning- and deep-learning-based methods are formidable alternative methods to tackle financial timeseries data, which are characterised by their complicated and non-linear nature, hence creating significant challenges to analysis methods. The dissertation investigates a long-standing controversy between classical and new forecasting methods by centring on a central question: forecasting the notoriously volatile stock, Tesla Inc. (TSLA), during a time of extreme market volatility from 2020 to 2025. The motivation behind this study comes from a desire to find whether or not a classical forecasting algorithm, aided by further market data, can come even close to matching a cutting-edge deep-learning-based forecasting algorithm during such exceptional times. To understand the research question, a comparative quantitative study was performed with a Cross-Industry Standard Process for Data Mining (CRISP-DM) approach. Two models were proposed and compared: a statistical baseline, called AutoRegressive Integrated Moving Average with eXogenous variables, or ARIMAX, with a Long Short-Term Memory neural network, or LSTM, which represents the deep learning paradigm. The models were trained by using a dataset with historical daily TSLA prices, together with additional data from a popular benchmark, such as the SP 500, and a well-known volatility index, known as the CBOE Volatility Index or VIX, as well as a complete set of engineered technical indicators, with the goal to better understand market trends, volatility, and momentum. The models’ predictive accuracy was rigorously assessed by utilising an unseen test dataset, with strict adherence to popular regression measures. This study’s results were conclusive. The Long Short-Term Memory algorithm essentially dominated the Autoregressive Integrated Moving Average with Exogenous Variables algorithm across all evaluation criteria, suggesting its better capability to recognise complex patterns inherent within the varying nature of stock data. The LSTM attained a Mean Absolute Percentage Error (MAPE) of just 5.02%, compared to the 27.76% MAPE achieved by the ARIMAX algorithm. Similarly, the Root Mean Squared Error (RMSE) generated by the LSTM was over five times less compared to that by the ARIMAX algorithm ($20.34 compared to $107.83). The major contribution of the current work is to offer significant empirical evidence supporting that within environments with high asset volatility, the forecasting accuracy of deep learning models significantly exceeds that of traditional linear statistical models. One of the major implications derived from our work addresses financial practitioners and quantitative analysts, specifying that their overreliance on less complex models with regard to such assets proves to be insufficient, leading to significant forecasting errors. Thus, the study points to a necessity to construct and utilise sophisticated methods from a deep learning perspective to achieve additional and applicable understanding within the dynamic realm of financial markets.Forecasting stock prices is a vital and complex task within the realm of financial research, a challenge that is further exacerbated by rising volatility so typical of modern markets. While statistical models have been employed for a long time, machine-learning- and deep-learning-based methods are formidable alternative methods to tackle financial timeseries data, which are characterised by their complicated and non-linear nature, hence creating significant challenges to analysis methods. The dissertation investigates a long-standing controversy between classical and new forecasting methods by centring on a central question: forecasting the notoriously volatile stock, Tesla Inc. (TSLA), during a time of extreme market volatility from 2020 to 2025. The motivation behind this study comes from a desire to find whether or not a classical forecasting algorithm, aided by further market data, can come even close to matching a cutting-edge deep-learning-based forecasting algorithm during such exceptional times. To understand the research question, a comparative quantitative study was performed with a Cross-Industry Standard Process for Data Mining (CRISP-DM) approach. Two models were proposed and compared: a statistical baseline, called AutoRegressive Integrated Moving Average with eXogenous variables, or ARIMAX, with a Long Short-Term Memory neural network, or LSTM, which represents the deep learning paradigm. The models were trained by using a dataset with historical daily TSLA prices, together with additional data from a popular benchmark, such as the SP 500, and a well-known volatility index, known as the CBOE Volatility Index or VIX, as well as a complete set of engineered technical indicators, with the goal to better understand market trends, volatility, and momentum. The models’ predictive accuracy was rigorously assessed by utilising an unseen test dataset, with strict adherence to popular regression measures. This study’s results were conclusive. The Long Short-Term Memory algorithm essentially dominated the Autoregressive Integrated Moving Average with Exogenous Variables algorithm across all evaluation criteria, suggesting its better capability to recognise complex patterns inherent within the varying nature of stock data. The LSTM attained a Mean Absolute Percentage Error (MAPE) of just 5.02%, compared to the 27.76% MAPE achieved by the ARIMAX algorithm. Similarly, the Root Mean Squared Error (RMSE) generated by the LSTM was over five times less compared to that by the ARIMAX algorithm ($20.34 compared to $107.83). The major contribution of the current work is to offer significant empirical evidence supporting that within environments with high asset volatility, the forecasting accuracy of deep learning models significantly exceeds that of traditional linear statistical models. One of the major implications derived from our work addresses financial practitioners and quantitative analysts, specifying that their overreliance on less complex models with regard to such assets proves to be insufficient, leading to significant forecasting errors. Thus, the study points to a necessity to construct and utilise sophisticated methods from a deep learning perspective to achieve additional and applicable understanding within the dynamic realm of financial markets.5 0Item Restricted Image classification using machine learning for digital forensic investigations(Saudi Digital Library, 2025) Abu Sallamah, Yousef; Akinbi, AlexThe exponential increase in the production of manipulated digital media, along with the sheer abundance of digital forensic evidence creates a serious problem to the task of a law enforcement agency and practitioners of digital forensics. Available image classification forensics tools that do exist are largely commercial and narrow in scope, thus limiting the viability of wide usage in a forensics application. The following dissertation will fill these gaps by proposing an open source and machine learning driven image classification system to classify images related to forensics, with particular sensitivity to weapons and narcotics, which are common categories of cybercrime. The ResNet-50 convolutional neural network model was trained using transfer learning based on more than 47,000 labelled images and optimised to obtain the highest accuracy and robustness rates. The system is designed to include the use of explainable artificial intelligence (XAI) methods (such as Grad-CAM visualisations) to provide improved transparency, interpretability, and possible admissibility in courts of law. Statistical performance assessment, test of robustness under the degradations to occur in practice and usability testing by practitioners were carried out. The benchmark comparison with the state-of-the-arts studies reveals the proposed solution to outperform the traditional support vector machine and Bag-of-Visual-Words approaches in addition to solving the problem of scalability and explainability. The novelty of the presented dissertation lies in the fact that multi-label forensic classification, open-source applicability, and XAI have already been integrated into one framework, and a transparent, reproducible, and practical contribution to the digital forensics field has been introduced. The limitations in terms of representativeness of the databases and resistance to adversarial attacks are discussed, and the future research of federated learning, adversarial robustness, and combination to forensic triage systems are suggested.17 0Item Restricted Computational Approaches for Drug Repositioning and Target Discovery in Alzheimer’s Disease(King Abdullah University of Science and Technology (KAUST), 2024) Alamro, Hind; Gao, XinAlzheimer’s Disease (AD) presents significant challenges to global healthcare systems due to its complex and progressive nature. Despite extensive research, the underlying mechanisms of AD lack clarity, and current treatments only alleviate symptoms without halting disease progression. Consequently, there is an urgent need for computational approaches that can accelerate research efforts and aid in the development of more effective treatments for AD. In this thesis, we address these critical challenges by developing computational and AI-based methods to improve the early detection of AD, identify novel biomarkers, and explore new therapeutic strategies through drug repositioning. To begin with, we focus on identifying key biomarkers associated with AD using gene expression datasets and then expand it to the identification of biomarkers through exploring the association between AD and its comorbidity, resulting in the discovery of new hub genes and miRNAs. Next, we examine the potential for drug repositioning by mining biomedical literature to uncover associations between drugs, targets, and diseases. This task was fulfilled by developing a systematic pipeline to extract valuable information from a curated collection of AD-related literature. The resulting data is subsequently used to construct a disease-specific knowledge graph, which is employed for drug repositioning using advanced graph-based techniques. Overall, this thesis contributes to AD research by employing computational methods, multi-data integration, and literature mining to provide new insights and therapeutic strategies. This work identifies key participants in AD progression and presents a pathway to accelerate the discovery of treatments through computational approaches.10 0Item Restricted Intelligent Data-Driven Models for Accurate Multi-factors Prediction of Carbon Credit Prices(Saudi Digital Library, 2025) Alshatri, Najlaa saad; Ghannam, SafaaThis thesis addresses the challenge of accurately predicting carbon credit prices, which are non-linear, non-stationary, and influenced by multiple correlated external factors such as energy prices, environmental indicators, and economic conditions. Accurate pricing is vital for transparency and effectiveness in carbon markets. A systematic literature review identified research gaps, leading to the development of a Carbon Credit Multi-Factor Prediction (CCMFP) model integrating factor identification and optimized prediction algorithms. The proposed Carbon Credit Multi-Factor Identification (CCMFI) model combines random forest regression with explainable AI to identify the most influential factors among 22 external variables. Feature reduction and extraction techniques, independent component analysis (ICA), nonlinear ICA (NLICA), and principal component analysis (PCA), were then applied, with extracted components used as inputs to SVR and MLP models. Using daily Australian Carbon Credit Units (ACCUs) prices as a case study, experiments evaluated the impact of different factor sets on prediction accuracy. The models achieved an R2 of over 97%, with optimal performance from factors including environmental technology patents, CO2 emissions, renewable energy adoption, global carbon allowances, coal and crude oil prices. These findings enhance market confidence, reduce financial risks, and support global climate change mitigation through effective carbon credit utilization.53 0Item Restricted Long-Term Dependency Margin Maximization Model (LTDM3): Dealing with Concept Drift in Personalized Learning Systems(Saudi Digital Library, 2023) Allogmany, Bander; Josyula, DarsanaAdvances in data analytics and intelligent technologies are enabling smart learning environments that promote personalized learning. Personalized learning systems where learners engage with information in a manner tailored to their unique needs, goals, and abilities have garnered significant academic research attention. If students can achieve their objectives faster than with traditional learning methods, it would increase their motivation and reduce their likelihood of dropping out. It can also offer educators a better understanding of each student’s learning process, enabling them to teach more effectively. Artificial intelligence (AI) plays a vital role in the development of personalized learning systems. Rapid advancements in AI technologies enable tracking and modifying of each student’s learning environment. Machine learning algorithms facilitate the determination of students’ learning styles, abilities, and progress throughout the learning process. One of the major challenges to effective personalization is the resistance of machine learning models to adapt to non-stationary data streams. Machine learning models for personalized learning systems are susceptible to the concept drift phenomenon, a deterioration of the model’s performance over time due to changes in data distribution. These arise due to factors affecting learning ability, including changes in family structure, parental involvement, peer relationships, learner behavior, personal interests, environmental influences such as nutrition and sleep, and so on. For successful personalization, it is critical that underlying predictive and classification models be able to adapt successfully to data changes that contribute to the drift phenomenon. This research proposes a method to address concept drifts in personalized learning systems that involve training using sequential features extracted automatically, noting when concept drifts are causing model deterioration, and automatically adjusting the trained model to improve model performance in the presence of drift. Unlike large language models (LLMs), which usually lack inherent capabilities for indicating that a concept drift has occurred, the approach presented in this dissertation can detect and point out instances of concept drift. Detecting concept drift is important for initiating specific interventions, whereas large language models tend to obscure or overlook such changes in the data. The proposed approach aims to enhance the accuracy and effectiveness of predictive models, ensuring personalized learning systems deliver pertinent and useful recommendations even when student preferences change. While conducting experiments using a real-world dataset related to students’ interactions with educational systems, the proposed model shows impressive results in managing concept drifts. Moreover, the proposed model shows resilience against two major types of drift: incremental and sudden drifts. This indicates that by using the proposed approach, we can ensure that the predictive models maintain their effectiveness in the presence of different types of drift.5 0Item Restricted Pseudo-Labeling for Deep Learning-Based Side-Channel Disassembly Using Contextual Layer and Feature Engineering(Saudi Digital Library, 2025) Alabdulwahab, Saleh Sami S; Son, YunsikEmbedded devices face critical cyber-attacks due to their lightweight design and the sensitive data they handle. Integrating cloud and embedded systems increases the need for security measures against threats. Among these threats are deep learning-based side-channel disassembly attacks, which can expose sensitive information or steal software intellectual properties. Conducting a security test to evaluate the systems against these threats is essential. However, the main challenges include a comprehensive and refined dataset for training deep learning-based side-channel attacks and the lack of public datasets; labeling and profiling such attacks are costly and time-consuming. Additionally, accurately disassembling a single instruction is difficult due to the multiple classes representing each instruction and the obfuscation caused by dummy instructions. This study aimed to create an advanced side-channel evaluation methodology that performs three main deep-learning tasks: profiling using context-aware pseudo-labeling techniques at an instruction level, a disassembly model enhanced with moving log-transformed temporal interaction features, and a sequence labeling model for the detection of dummy instructions using natural language processing techniques. Utilizing gated recurrent units, the proposed pseudo-labeling model achieved 0.996 R2 in estimating the power trace for the assembly instructions. The proposed features improved the disassembly model's accuracy to 0.993, outperforming the related works. Additionally, the detection of dummy instructions using a long short-term memory model reached an accuracy of 0.979. This study provides valuable insights and methodology for measuring the software robustness against side-channel attacks.18 0
