A Corpus-Based Approach to Saudi Arabian Dialects: Implications for Dialect Identification, Machine Translation and Sentiment Analysis
No Thumbnail Available
Date
2026
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Saudi Digital Library
Abstract
Arabic is a linguistically diverse language comprising numerous dialects spoken across
the Arab world, with the dialects of Saudi Arabia representing a particularly rich
and complex landscape shaped by regional, historical, and social influences. Despite
their linguistic and cultural importance, Saudi Arabic dialects remain significantly
underrepresented in computational linguistics and natural language processing (NLP)
research and are often treated as low-resource language varieties due to the scarcity of
annotated data, standardized resources, and evaluation benchmarks. Existing resources
and models predominantly focus on Modern Standard Arabic (MSA) or broad regional
dialect groups, such as Egyptian or Gulf Arabic, leaving fine-grained Saudi dialectal
variation largely unexplored. This gap limits the development of dialect-aware NLP
systems capable of accurately processing Saudi Arabian dialects.
The main aim of this thesis is to systematically investigate fine-grained Saudi Arabic
dialects across three core NLP tasks, dialect identification, machine translation, and
sentiment analysis, by evaluating the performance of both pre-trained language models
(PLMs) and large language models (LLMs) in processing low-resource Saudi dialectal
data. To achieve this aim, the study introduces several novel datasets designed to
support the analysis of Saudi Arabian dialects across different NLP tasks. For dialect
identification, the Saudi Arabian Tweets Corpus (SATC) is developed, comprising Saudi
Arabic tweets representing five major regional dialects. In addition, the Saudi Arabian
Dialectal Song Lyrics Corpus (SADSLyC) is constructed based on Saudi song lyrics,
capturing culturally rich and regionally diverse linguistic expressions. For machine
translation, the study presents SADSLyC-E-MSA, a parallel subset of the SADSLyC
corpus that includes aligned translations in English and Modern Standard Arabic.
Finally, for sentiment analysis, the research introduces the Saudi Arabian Proverbs
Corpus (SAPC), a collection of authentic proverbs gathered from different regions of
Saudi Arabia, reflecting deeply rooted cultural meanings and regional dialectal variation.
The findings demonstrate that task-specific fine-tuning combined with dialect-
focused corpora significantly improves language models performance on fine-grained
Saudi dialect in different NLP tasks. In the machine translation experiments, the best-
performing model achieved an improvement of 39.17% over the baseline model without
fine-tuning when translating Saudi Arabian dialects, highlighting the substantial impact
of adapting models to dialect-specific data.
Description
Keywords
Arabic NLP, Saudi Dialects, Dialect Identification, Machine Translation, Sentiment Analysis
