Multilingual Model Enhancement Framework for Spam Detection Using a Human-Centred Approach
No Thumbnail Available
Date
2026
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Saudi Digital Library
Abstract
Phishing attacks continue to exploit human behaviour across digital communication
platforms, yet most defence solutions are designed for English speakers. Arabic-speaking
users face increased risk due to limited awareness and a lack of high-quality phishing
datasets. This thesis introduces a human-centred and culturally-aware cybersecurity
framework that enhances phishing understanding and detection for Arabic users.
The work begins with three empirical studies focused on phishing perception, emo-
tional triggers, and behavioural responses among Arabic-speaking university students.
Results reveal that urgency, authority appeals, and religious or familial influence sig-
nificantly increase susceptibility. These insights expose cultural and emotional factors
that are not represented within current English-centric phishing defences.
To address dataset scarcity, the thesis develops a multi-stage English-to-Arabic
translation pipeline that produces high-fidelity security text while preserving core de-
ception signals such as URLs, numeric codes, and persuasion techniques. The pipeline
combines context-aware prompting with multilingual neural translation and human
validation to ensure linguistic accuracy and security relevance. A post-translation en-
richment process integrates the user study insights by adding structured labels such
as phishing type, emotional manipulation, cultural exploitation, behavioural red flags,
audience targeting, and recommended defence actions. Three corpora are produced for
email, SMS, and instant messaging, each containing 1,000 validated and enriched sam-
ples. These represent the first Arabic phishing datasets that embed human behavioural
intelligence within a unified annotation standard.
The thesis then evaluates machine learning approaches for Arabic phishing detec-
tion using the newly developed datasets. In the cross-channel evaluation, the fine-
tuned mT5 model achieved accuracies of 87.2% on SMS, 91.5% on email, and 83.5%
on messaging (87.4% average), with corresponding F1-scores of 86.9%, 91.5%, and
82.9%. In separate baseline-to-enhanced mT5 experiments, overall accuracy increased
by 1.37 percentage points on SMS (96.84% to 98.21%) and by 7.46 points on Enron email
(74.68% to 82.14%); Enron ham accuracy increased by 22.17 points (56.75% to 78.92%).
The system also generates interpretable outputs that explain prediction rationale and
reinforce user awareness, supporting educational and defensive goals.
Overall, this thesis advances cybersecurity for underserved language communities
by linking behavioural research with cross-lingual artificial intelligence. The frame-
work improves phishing resilience among Arabic-speaking users and establishes essen-
tial linguistic and cultural resources for future research on secure digital transforma-
tion in the region.
Description
Keywords
Phishing Detection, Arabic Cybersecurity, Human-Centred Cybersecurity, Arabic Phishing Datasets, Cross-Lingual Artificial Intelligence.
