Engineering and Classifying Indirect Prompt Injections: A Socio-Technical Approach to AI Agent Security

dc.contributor.advisorAbudayeh, Mohammad
dc.contributor.authorAldurayhim, Abdulaziz
dc.date.accessioned2026-08-11T13:14:17Z
dc.date.issued2026
dc.description.abstractAs autonomous AI agents become more prevalent in enterprise settings, they are now susceptible to "Indirect Prompt Injection (IPI)" attacks, which embed adversarial payloads in external data sources to manipulate the agent's contextual execution and force unauthorised actions. This dissertation holds that IPI attacks constitute a socio-technical phenomenon because they exploit semantic vulnerabilities in agent architectures and psychological manipulation patterns in human communication. To address this threat, the study proposes a three-dimensional taxonomy based on payload type, injection vector, and social engineering trigger. The initial corpus comprised 1,119 instances; following rigorous preprocessing and deduplication, a refined dataset of 858 validated samples was established for modelling. A separate 1,032-sample verified dataset pack was used only as an independent reference for dataset verification, not as the definitive modelling dataset. A DeBERTa-v3-base classifier is fine-tuned to classify social engineering triggers, using RoBERTa-base as a baseline. On the OSINT subset, DeBERTa's macro-F1 score was 0.983, and the accuracy was 91.7%. The study also presents a probabilistic Confused Deputy threat model and an ACE defence architecture consisting of architectural controls, runtime detection, privilege minimisation, human-in-the-loop oversight, and practitioner education. The findings conclusively demonstrate that integrating advanced NLP detection mechanisms with socio-technical security frameworks significantly enhances enterprise resilience against IPI threats.
dc.format.extent56
dc.identifier.citationAldurayhim, A. S. (2026). Engineering and Classifying Indirect Prompt Injections: A Socio-Technical Approach to AI Agent Security (Master's thesis, De Montfort University).
dc.identifier.urihttps://hdl.handle.net/20.500.14154/79888
dc.language.isoen
dc.publisherSaudi Digital Library
dc.subjectIndirect Prompt Injection
dc.subjectAI Agent Security
dc.subjectSocio-Technical Systems
dc.subjectDeBERTa
dc.subjectTransformer Models
dc.subjectMachine Learning
dc.subjectCyber Security
dc.subjectSocial Engineering
dc.subjectAI
dc.subjectArtificial Intelligence
dc.subjectSecurity of AI
dc.subjectAI for Security
dc.titleEngineering and Classifying Indirect Prompt Injections: A Socio-Technical Approach to AI Agent Security
dc.typeThesis
sdl.degree.departmentSchool of Computer Science and Informatics
sdl.degree.disciplineCyber Security
sdl.degree.grantorDe Montfort University
sdl.degree.nameMaster of Cyber Security

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
SACM-Dissertation.pdf
Size:
1.67 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
1.61 KB
Format:
Item-specific license agreed to upon submission
Description:

Copyright owned by the Saudi Digital Library (SDL) © 2026