SACM - United Kingdom
Permanent URI for this collectionhttps://drepo.sdl.edu.sa/handle/20.500.14154/9667
Browse
77 results
Search Results
Item Restricted Engineering and Classifying Indirect Prompt Injections: A Socio-Technical Approach to AI Agent Security(Saudi Digital Library, 2026) Aldurayhim, Abdulaziz; Abudayeh, MohammadAs autonomous AI agents become more prevalent in enterprise settings, they are now susceptible to "Indirect Prompt Injection (IPI)" attacks, which embed adversarial payloads in external data sources to manipulate the agent's contextual execution and force unauthorised actions. This dissertation holds that IPI attacks constitute a socio-technical phenomenon because they exploit semantic vulnerabilities in agent architectures and psychological manipulation patterns in human communication. To address this threat, the study proposes a three-dimensional taxonomy based on payload type, injection vector, and social engineering trigger. The initial corpus comprised 1,119 instances; following rigorous preprocessing and deduplication, a refined dataset of 858 validated samples was established for modelling. A separate 1,032-sample verified dataset pack was used only as an independent reference for dataset verification, not as the definitive modelling dataset. A DeBERTa-v3-base classifier is fine-tuned to classify social engineering triggers, using RoBERTa-base as a baseline. On the OSINT subset, DeBERTa's macro-F1 score was 0.983, and the accuracy was 91.7%. The study also presents a probabilistic Confused Deputy threat model and an ACE defence architecture consisting of architectural controls, runtime detection, privilege minimisation, human-in-the-loop oversight, and practitioner education. The findings conclusively demonstrate that integrating advanced NLP detection mechanisms with socio-technical security frameworks significantly enhances enterprise resilience against IPI threats.12 0Item Restricted Using surveys and on-bus sensors to understand and predict bus passenger satisfaction(Saudi Digital Library, 2026) Ghulman, Mayyar; Enoch, Marcus; Kyriakopoulos, Kostas; Beechey, MattBackground: Rapid urbanisation and population growth have increased the demand for mobility in cities, leading to congestion and environmental issues such as air and noise pollution. Public transportation, particularly buses, is vital for accessibility, economic participation, and reducing social isolation, especially in areas with limited rail services. Despite these benefits, bus usage in many developed countries has stagnated at best and, in some cases, declined, while car travel is on the rise, underscoring the need to enhance bus services, make them more attractive, and increase patronage. A key challenge is the reliance on passenger perception surveys to assess bus service quality. Objective measures, such as temperature and crowding, are often collected in a fragmented manner, creating a gap in the evidence available to operators and policymakers. Without integrating perceived service quality (PSQ) with delivered service quality (DSQ), identifying practical service improvements to boost passenger satisfaction becomes obscure. Purpose: This thesis aims to investigate the impact of perceived service quality and actual delivered service quality—measured through real-time bus performance indicators— on bus passenger satisfaction and loyalty. Additionally, it evaluates predictive models that leverage real-time operational data from onboard bus sensors to predict passenger satisfaction. Methods: A quantitative research design was adopted that combined passenger perceptions with measured bus performance. Data were collected from two bus services in the UK (East Midlands): Service 90 (Nottingham–Newark) and Novus Direct (Leicester–New Lubbesthorpe). Perceived service quality was captured through an in-person, web-based customer satisfaction survey, which yielded 208 valid responses, while delivered service quality was measured using an onboard sensor system that recorded five operational indicators: punctuality, busload, speed, vibration, and temperature. The data analysis followed three stages: (1) principal component analysis reduced the perceived service quality indicators to latent dimensions; (2) the hypotheses were tested using PLS-SEM, together with multi-group analysis (MGA) to compare passenger groups; and (3) supervised machine learning was used to predict satisfaction from the operational indicators, with performance evaluated using standard classification metrics. Findings: Perceived service quality dimensions, namely pre-trip experience, on-bus environment, driver-related factors, and safety, were found to significantly contribute to predicting satisfaction, with the onboard environment being the most significant factor. Bus crowding and comfort: driving style, steering, and driver behaviour were among the most important attributes, with the highest loadings. There were no significant differences in these predictions between captive and choice riders, or between work and non-work trip groups, based on MGA. However, separate analyses of each group’s sample produced very few differences. On the other hand, in terms of delivered service quality, only actual bus load and actual speed were found to have a significant, albeit small, negative impact on satisfaction. The Random Forest classifier achieved the strongest predictive performance among the supervised classification models, with temperature, bus load, and punctuality identified as the most important predictors of passenger satisfaction. Contribution: This thesis presents a comprehensive investigation that significantly contributes to the field of public transport satisfaction through data collection and analysis. It is one of the few studies to utilise both subjective survey-based data and objective sensor-derived bus operational data to assess and predict passenger satisfaction with bus travel. Methodologically, the thesis demonstrates a practical approach for capturing, synchronising, and merging high-frequency onboard sensor data with passenger trip-level survey responses. While the integration of subjective and objective measures is established in other domains, this thesis is among the few to apply this approach to bus passenger satisfaction, using a purpose-built onboard sensor system to capture service quality at the individual-trip level and link it directly to passengers’ survey responses. This dual approach generates insights into both satisfaction and loyalty, as well as actionable operational predictions. Conclusion: This study provides bus operators and transport authorities with evidence-based methods for real-time service monitoring and targeted interventions to improve passenger satisfaction and loyalty. This research contributes to the existing body of knowledge, not only regarding buses but also for the public transport sector as a whole.5 0Item Restricted Algorithms for Growth Pattern Based Analysis of Lung Adenocarcinoma Histology(Saudi Digital Library, 2026) AlRubaian, Arwa; Rajpoot, Nasir; Raza, ShanLung Adenocarcinoma (LUAD) is a leading cause of cancer deaths and shows considerable variation in patient outcomes due to its complex histological, immune, and molecular diversity. Traditional computational pathology methods use global, pattern-agnostic representations that do not capture the structured organization of tumor tissue. This thesis develops new, pattern-aware computational approaches to address this gap. First, we introduce CellOMaps, a biologically informed image representation that encodes the cellular organization of LUAD tissue. By focusing on spatial structure rather than raw visual appearance, CellOMaps enables robust, scalable characterization of LUAD histological growth patterns. Using this representation, we propose a framework for dense, patch-level growth pattern classification across Whole Slide Images (WSIs), offering a reproducible and spatially resolved alternative to conventional predominant-pattern assessment and demonstrating improved robustness across institutions. Next, we investigate immune heterogeneity in LUAD by analyzing the spatial distribution of Tumor-Infiltrating Lymphocytes (TILs). We introduce GPS-TILs, a growth pattern–specific digital biomarker that quantifies the presence, density, abundance, and spatial dispersion of immune infiltration within distinct morphological patterns. This pattern-aware immune profiling improves prognostic stratification compared to global TILs assessment and manual grading, highlighting the importance of incorporating histological context into immune analysis. Finally, we address molecular inference from histopathology by predicting Tumor Mutational Burden (TMB) from WSIs. We introduce GRIL-GNN, a graph-based learning framework that aggregates information across spatially organized tissue regions while preserving detailed morphological features. This approach uses a novel unified learning objective that combines supervised learning with self-supervised regularization, enabling robust representation learning under weak supervision. By restricting analysis to localized tissue regions with shared histological patterns, this work examines the distribution of the TMB predictive signal within different morphological patterns. Overall, this work demonstrates that integrating spatial and pattern-aware analysis improves the interpretation of morphological, immune, and molecular signals in LUAD. The approaches proposed here provide a framework for more precise, clinically relevant computational pathology, with potential applications to other cancers characterized by complex tissue structure.14 0Item Restricted Saudi Dialect Sentiment Analysis: A Hybrid Approach and Evaluation on Real-World Consumer Reviews(Saudi Digital Library, 2026) Alsemaree, Ohud; Singh Gill, SukhpalThis thesis presents a hybrid sentiment analysis framework for the Saudi dialect, addressing the limited availability of linguistic resources and annotated datasets for Arabic sentiment analysis. The proposed approach combines lexicon-based methods with machine learning and deep learning techniques to improve sentiment classification performance. A Saudi dialect sentiment lexicon (LSAnArTe) and a large-scale annotated dataset were developed to support the research. Experimental results demonstrate that the hybrid framework outperforms baseline models, achieving high classification accuracy and providing valuable resources for future research in Arabic Natural Language Processing and sentiment analysis.13 0Item Restricted Artificial Intelligence through Machine Learning techniques to enhance the application of 3D body scanning in apparel shape and sizing(Saudi Digital Library, 2026) Alhassawi, Ruqey Ali; Simeon, Gill; Steve, Hayes; Kristina, BrubacherSignificant challenges persist in realising the full potential of technology related to accurate and inclusive body dimension variation and garment sizing and fit. Traditional methods often fail to capture the complexity of human body morphology, highlighting the value of more detailed approaches to analysing body dimension variation. This doctoral research aims to support the visual analysis of anthropometric population data through the integration of artificial intelligence (AI) and machine learning (ML) techniques, addressing limitations in traditional anthropometric methods used for apparel sizing and body–to–pattern mapping. A mixed–methods approach was employed across five interconnected phases, leveraging 3D body scanning (3DBS) technology to analyse and compare real–world body dimensions, classical garment sizing classifications and garment patterns. The research involved: (1) a comprehensive analysis of 3DBS data to establish body dimension diversity, (2) a critical reassessment of the traditional 8–head figure ratio, (3) clustering algorithms (Hierarchical, self–organizing map (SOM), k–means) to classify body types, (4) application of support vector machine (SVM) and principal component analysis–SVM (PCA–SVM) models for accurate size prediction, and (5) enhanced regression analysis to develop a data–driven approach for garment pattern adjustment. A dataset of 677 female participants from a range of ethnic backgrounds was utilised. Significant dimensional variations within conventional size groups were identified, revealing limitations in traditional measurement-based sizing systems within the study sample. Key findings demonstrate frequent deviations from the classical 8–head figure proportion model, emphasising the need for a more comprehensive approach. Clustering algorithms successfully delineated distinct morphological categories, while SVM modelling exposed trade–offs between predictive accuracy and computational complexity. Regression analysis established quantitative relationships between body measurements and pattern block parameters, offering a means of examining how body dimension variation relates to patternmaking practice. This research makes several theoretical, methodological and practical contributions. Theoretically, it provides data-based evidence of body proportion variability within standard size categories, challenges the classical 8-head figure proportion model using measured data, and identifies distinct body shape clusters within the study sample. Methodologically, it applies an integrated analytical framework – combining 3D body scan data, statistical analysis, ML clustering and classification, and regression analysis – to examine body dimension variation and body-to-pattern relationships. Practically, it provides how data-driven analysis of anthropometric variation may inform patternmaking considerations, subject to further applied investigation. This research examines the integration of 3D body scanning and computational techniques within anthropometric analysis. The use of data-based derived visual tools provides a means of representing and exploring body variation within the study sample. The findings highlight the potential relevance of data-driven approaches to sizing and may inform further investigation into how body diversity is represented within garment sizing systems.17 0Item Restricted Machine Learning for Radiotherapy Treatment of Prostate Cancer(Saudi Digital Library, 2026) Alqarni, Maram; Teresa, Guerrero Urbano; Andrew, KingExternal beam radiotherapy (EBRT) and brachytherapy (BT) are both forms of radiation treatment used for prostate cancer to destroy cancer cells. EBRT applies the radiation externally while BT involves placing radioactive seeds inside the prostate. At Guy’s Cancer Centre, both treatment modalities are performed depending on various factors. Each of the treatment modalities involves different imaging modalities used for treatment planning, delivery and follow-up. However, both have some overlapped clinical tasks such as defining the clinical target volume (CTV) and organs at risk (OARs) from imaging data. The work described in this thesis aims to perform research to promote clinical translation of machine learning (ML) techniques to streamline workflows in EBRT and BT. The first piece of work in this thesis focuses on an ML-based segmentation model for prostate MRI. One of the main challenges affecting clinical adoption of ML in MRI segmentation is the domain shift problem. The findings of this piece of work reveal for the first time the significant impact on model performance of using different acquisition/annotation protocols, even if using the same scanner vendor/field strength. It is shown that training an ML model with data that covers the important sources of domain shift can produce a robust model with good generalisability performance. The next piece of work investigates the possibility of race bias in ML-based prostate MRI segmentation. Through experiments on a controlled dataset of White and Black patients, it is shown that the model performance gap between Black and White subjects is dependent on the level of (im)balance between Black and White subjects in the training data. Again, it is shown that training using demographically balanced data can produce a fair and robust model. The conclusion from both of these pieces of work is that model performance can be robust if the training data is sufficiently diverse, both in terms of image characteristics and patient demographics. Building upon these analyses, the thesis next investigates the clinical utility of a diagnostic prostate MRI model trained on diverse data and externally validates it on in-house clinical data. The evaluation of this model encompasses not only standard quantitative metrics but also measurement of inter-observer variability in manual segmentation and assessments of performance on downstream clinical tasks. Next, the thesis investigates the clinical utility of multi-organ ML-based segmentation models. Here, two models are investigated: one for planning MRI called the “FIMRAa-P” model and another radiotherapy CT model called the “PelvisMA-CT” model. Both models are extensively evaluated quantitatively and qualitatively by five observers. The agreement between the quantitative metrics and the qualitative clinical metrics is also investigated for each clinical structure, revealing generally poor agreement between the two. It is also shown that this agreement is dependent on the structure being segmented and the profession of the clinicians who perform the evaluations. One of the main clinical translation outcomes of this thesis is the deployment of PelvisMA-CT by the Clinical Scientific Computing (CSC) group at GSTFT, and its integration into a contouring application called GSTTAutoSeg. This model is currently being used clinically at Guy’s Cancer Centre and the thesis presents the results of a monitoring and enhancement study based on its ongoing clinical use. Overall, the thesis presents a number of key contributions, all aimed at promoting clinical translation of ML in EBRT and BT. It is hoped that the work performed will accelerate the benefits of ML in radiotherapy treatment planning and delivery and ensure that all patients benefit from the introduction of the thoroughly evaluated new technology.8 0Item Restricted Predicting Carbon Credit Prices Using Advanced Machine Learning Techniques(Saudi Digital Library, 2026) Rayan, Najdi; Wang, HaiAccurate forecasting of carbon credit prices supports risk management, investment decisions, and policy assessment in the context of climate action. EU ETS carbon prices exhibit volatility, non-linearity, and non-stationarity, which reduces the effectiveness of traditional forecasting models. This dissertation proposes and evaluates a three-stage hybrid machine learning model for one-day-ahead forecasting of EU Emissions Trading System (EU ETS) carbon prices. The architecture follows a divide-and-conquer strategy. First, Wavelet Packet Decomposition (WPD) decomposes the carbon price signal into multiple frequency components. Second, a Gated Recurrent Unit (GRU) network models temporal dependencies and forecasts the trend component. Third, an Extreme Gradient Boosting (XGBoost) model predicts and corrects the GRU residual errors using wavelet-derived detail components as input features. The model was trained and tested on a dataset covering January 2018 to December 2024. The dataset includes EU ETS carbon prices, Brent crude oil prices, and electricity prices, while the forecasting model is univariate and uses the carbon price series only. On an unseen test set of 510 days, the model achieved a Mean Absolute Percentage Error (MAPE) of 1.66%, a Root Mean Squared Error (RMSE) of 4.86 EUR/ton, and a Mean Absolute Error (MAE) of 4.41 EUR/ton. The results indicate that combining signal decomposition, deep learning, and gradient boosting provides stable forecasting performance for EU ETS carbon prices under realistic evaluation conditions.13 0Item Restricted Analysing Large-Scale Attacks in IoT Environments using ML/DL(Saudi Digital Library, 2025) Bokhari, Mohammed Ibrahim K; Neetesh, SexenaThe fine-grained classification of malicious network traffic presents a significant and persistent challenge in cybersecurity, primarily due to the extreme class imbalance inherent in real-world network data. Conventional machine learning approaches, which apply a single, unitary model to the problem, have demonstrated limited success, often failing to effectively identify rare but critical minority attack classes. This dissertation argues that the conventional model paradigm is fundamentally flawed for this problem space and proposes a hierarchical, multi-stage classification framework as a more robust alternative. This research presents a comprehensive, multi-faceted investigation into this problem, using the 34-class CICIoT2023 dataset as a benchmark. The study was conducted across four distinct experimental paths, comparing two ensemble methods (XGBoost and Random Forest) and two class-handling strategies (a "Grouped" approach that manually merges similar classes and an "Ungrouped" approach that tackles all 34 classes directly). Within this structure, we designed and implemented a 4-tier hierarchical framework that employs a "divide and conquer" strategy, using an initial classifier to handle majority traffic and a class-level routing mechanism to delegate ambiguous samples to specialised recovery tiers. An adaptive resampling strategy was deployed within these tiers, concentrating aggressive SMOTE only where required. The empirical results provide a holistic validation of the proposed architecture. The optimal configuration—an Ungrouped, XGBoost-led hierarchical framework—achieved a final accuracy of 0.9228 and Macro-F1 score of 0.7948, a substantial improvement over all other experimental paths and conventional baselines. More significantly, this approach demonstrated a more than 800% increase in the F1-score for some of the under-represented minority classes. The analysis also revealed a key architectural principle: classifier performance is role-dependent, with different ensemble methods excelling in different roles within the hierarchy, highlighting the importance of managing the bias-variance trade-off at a systemic level. Finally, this work provides a rigorous, data-centric analysis that distinguishes between model limitations and the inherent limitations of the dataset, identifying a "dataset-induced ceiling" on performance for 5 of the 34 classes. The primary contribution of this dissertation is, therefore, a methodologically robust and architecturally novel framework, validated through a comprehensive, multi-path experimental design. The principles of hierarchical decomposition and adaptive resource allocation are domain-agnostic and offer a promising direction for future research into extreme imbalance problems.19 0Item Restricted AI-Powered Multimodel Detection System for Cybersecurity Attacks: Design, Implementation, and Evaluation(Saudi Digital Library, 2025) Alhazmi, Marwan; Nguyen, HoangAs cyber threats have become increasingly complex, so too has the need for advanced detection methods to be able to analyze different types of data. Historically, traditional intrusion detection systems (IDS), have relied on analyzing one form of data, either a statistical analysis of network traffic or an alert log written in text format. These limitations restrict the capability of IDSs to detect the many complexities associated with modern attacks. Therefore, this dissertation proposes an AI powered, multimodel detection system that utilizes a combination of both structured network data, and unstructured alert text, to improve the performance of intrusion detection systems. The methodologies include preprocessing and feature extraction on the CICIDS2017 dataset, machine learning algorithms for the analysis of structured data and Natural Language Processing (NLP) algorithms for the analysis of text data. The multimodel fusion method used late fusion where the predictions from each modality are combined to produce a single prediction. In addition, several classification algorithms were trained and tested including Random Forest, Logistic Regression, and Text Classification. Results showed that the multimodel system significantly outperformed the single-modality systems based on the evaluation metrics of Accuracy, Precision, Recall, and F1-Score. Furthermore, the multimodel fusion strategy enhanced the context of the detection by reducing false positive detections; this addresses a major challenge that is commonly experienced by researchers in the field of Intrusion Detection Systems (IDS). Therefore, this dissertation provides a practical, scalable, multimodel AI-based framework for detecting cybersecurity threats and demonstrates the effectiveness of using a combination of structured and unstructured data sources, along with providing direction for further advancements in Intelligent Intrusion Detection Systems.29 0Item Restricted Evaluating Static, Contextual, and End-to-End Embedding Techniques for Malware Detection on Dynamic API Call Data(Saudi Digital Library, 2026) Basfar, Mohammed Raed; Joey, LamThe rate of malware development continues to challenge cybersecurity, with traditional signature- and heuristic-based techniques overwhelmed by polymorphic and zero-day attacks. Natural language processing (NLP) offers a promising direction by modeling dynamic API call sequences as semantic linguistic data, enabling sophisticated embedding and sequence-learning methods to be used for malware detection. This dissertation contrasts and analyzes three typical embedding methods static, contextual, and end-to-end task-learned representations—under a shared experimental framework. Specifically, it employs Word2Vec embeddings with a Convolutional Neural Network (CNN), contextual BERT embeddings with a CNN, and a Bidirectional Long Short-Term Memory (BiLSTM) network with a trainable embedding layer and weighted loss function to address class imbalance. The experiments were conducted on a dynamic API call dataset of around 44,000 malware and 1,000 benign samples, summarized by the first 100 API calls executed under sandboxed conditions. Results indicate that the Word2Vec + CNN pipeline had the highest overall accuracy and malware detection precision but the lowest benign recall. The BERT + CNN model provided more balanced class performance, but at the expense of added computational overhead. The BiLSTM had the highest benign recall, as it was able to easily distinguish from non-malicious activity, but the lowest precision and hugely added resource use. The findings point out the competing trade-offs among detection accuracy, benign recall, and processing efficiency, highlighting the issue of aligning model selection with actual security contexts' resource constraints and priorities. The study contributes by reporting a comparative systematic review of the embedding approaches for malware detection and offering informative insights into performance vs. efficiency trade-offs. Apart from its scientific significance, it proves the larger potential of NLP-based approaches to supporting malware detection systems and to informing the design of responsive, resource-aware cybersecurity systems.20 0
