A Corpus-Based Approach to Saudi Arabian Dialects: Implications for Dialect Identification, Machine Translation and Sentiment Analysis

No Thumbnail Available

Date

2026

Journal Title

Journal ISSN

Volume Title

Publisher

Saudi Digital Library

Abstract

Arabic is a linguistically diverse language comprising numerous dialects spoken across the Arab world, with the dialects of Saudi Arabia representing a particularly rich and complex landscape shaped by regional, historical, and social influences. Despite their linguistic and cultural importance, Saudi Arabic dialects remain significantly underrepresented in computational linguistics and natural language processing (NLP) research and are often treated as low-resource language varieties due to the scarcity of annotated data, standardized resources, and evaluation benchmarks. Existing resources and models predominantly focus on Modern Standard Arabic (MSA) or broad regional dialect groups, such as Egyptian or Gulf Arabic, leaving fine-grained Saudi dialectal variation largely unexplored. This gap limits the development of dialect-aware NLP systems capable of accurately processing Saudi Arabian dialects. The main aim of this thesis is to systematically investigate fine-grained Saudi Arabic dialects across three core NLP tasks, dialect identification, machine translation, and sentiment analysis, by evaluating the performance of both pre-trained language models (PLMs) and large language models (LLMs) in processing low-resource Saudi dialectal data. To achieve this aim, the study introduces several novel datasets designed to support the analysis of Saudi Arabian dialects across different NLP tasks. For dialect identification, the Saudi Arabian Tweets Corpus (SATC) is developed, comprising Saudi Arabic tweets representing five major regional dialects. In addition, the Saudi Arabian Dialectal Song Lyrics Corpus (SADSLyC) is constructed based on Saudi song lyrics, capturing culturally rich and regionally diverse linguistic expressions. For machine translation, the study presents SADSLyC-E-MSA, a parallel subset of the SADSLyC corpus that includes aligned translations in English and Modern Standard Arabic. Finally, for sentiment analysis, the research introduces the Saudi Arabian Proverbs Corpus (SAPC), a collection of authentic proverbs gathered from different regions of Saudi Arabia, reflecting deeply rooted cultural meanings and regional dialectal variation. The findings demonstrate that task-specific fine-tuning combined with dialect- focused corpora significantly improves language models performance on fine-grained Saudi dialect in different NLP tasks. In the machine translation experiments, the best- performing model achieved an improvement of 39.17% over the baseline model without fine-tuning when translating Saudi Arabian dialects, highlighting the substantial impact of adapting models to dialect-specific data.

Description

Keywords

Arabic NLP, Saudi Dialects, Dialect Identification, Machine Translation, Sentiment Analysis

Citation

Endorsement

Review

Supplemented By

Referenced By

Copyright owned by the Saudi Digital Library (SDL) © 2026