Engineering and Classifying Indirect Prompt Injections: A Socio-Technical Approach to AI Agent Security

No Thumbnail Available

Date

2026

Journal Title

Journal ISSN

Volume Title

Publisher

Saudi Digital Library

Abstract

As autonomous AI agents become more prevalent in enterprise settings, they are now susceptible to "Indirect Prompt Injection (IPI)" attacks, which embed adversarial payloads in external data sources to manipulate the agent's contextual execution and force unauthorised actions. This dissertation holds that IPI attacks constitute a socio-technical phenomenon because they exploit semantic vulnerabilities in agent architectures and psychological manipulation patterns in human communication. To address this threat, the study proposes a three-dimensional taxonomy based on payload type, injection vector, and social engineering trigger. The initial corpus comprised 1,119 instances; following rigorous preprocessing and deduplication, a refined dataset of 858 validated samples was established for modelling. A separate 1,032-sample verified dataset pack was used only as an independent reference for dataset verification, not as the definitive modelling dataset. A DeBERTa-v3-base classifier is fine-tuned to classify social engineering triggers, using RoBERTa-base as a baseline. On the OSINT subset, DeBERTa's macro-F1 score was 0.983, and the accuracy was 91.7%. The study also presents a probabilistic Confused Deputy threat model and an ACE defence architecture consisting of architectural controls, runtime detection, privilege minimisation, human-in-the-loop oversight, and practitioner education. The findings conclusively demonstrate that integrating advanced NLP detection mechanisms with socio-technical security frameworks significantly enhances enterprise resilience against IPI threats.

Description

Keywords

Indirect Prompt Injection, AI Agent Security, Socio-Technical Systems, DeBERTa, Transformer Models, Machine Learning, Cyber Security, Social Engineering, AI, Artificial Intelligence, Security of AI, AI for Security

Citation

Aldurayhim, A. S. (2026). Engineering and Classifying Indirect Prompt Injections: A Socio-Technical Approach to AI Agent Security (Master's thesis, De Montfort University).

Endorsement

Review

Supplemented By

Referenced By

Copyright owned by the Saudi Digital Library (SDL) © 2026