Evaluating Hybrid AI Approaches in email Spam Detection: A Literature Review

No Thumbnail Available

Date

2026

Journal Title

Journal ISSN

Volume Title

Publisher

Saudi Digital Library

Abstract

Spam is still one of the most serious cybersecurity issues of the day, and serves as a method of delivering phishing and malware. In this thesis, two complementary parts are integrated: First, a PRISMA-based systematic review of 55 studies that were published from 2023 to 2025 and analyze the classical, AI-based and ensemble spam detection techniques; and second, a controlled proof-of-concept experiment on the Enron-Spam corpus, where three distinct classical classifiers (Naïve Bayes with Bag-of-Words, Logistic Regression with TF-IDF, Linear SVM with TF-IDF) are compared against a stacking ensemble of the three classifiers with a Logistic Regression meta-classifier. All four configurations were very close together and obtained an accuracy of between 0.989 and 0.993, with the stacking ensemble having the highest F1-score (0.9923) and the lowest false-positive rate (0.0088); but there were no significant differences between the four configurations when a paired statistical test was not used. The thesis contribution lies in the synthesis of the various perspectives of the 2023–2025 hybrid and ensemble approaches and an internally consistent baseline comparison on a standard corpus. The study is restricted to English-language email and to lexical features; further extensions of the study using deep-learning, multilingual, and adversarial techniques are suggested.

Description

This thesis investigates the effectiveness of hybrid artificial intelligence (AI) approaches for email spam detection by comparing traditional machine learning models with a hybrid stacking ensemble model. The research combines a systematic literature review of recent studies (2023–2025) with an experimental evaluation using the Enron-Spam dataset. The findings demonstrate that hybrid ensemble methods achieve competitive performance, improving spam detection while maintaining a balance between accuracy, precision, recall, and false-positive rates. The study contributes to the development of more effective and reliable AI-based email security solutions.

Keywords

Email Spam Detection, Cybersecurity, Artificial Intelligence (AI), Machine Learning, Hybrid AI Models, Stacking Ensemble, Natural Language Processing (NLP), TF-IDF, Naïve Bayes

Citation

APA 7th

Collections

Endorsement

Review

Supplemented By

Referenced By

Copyright owned by the Saudi Digital Library (SDL) © 2026