RAQEEB: AN ANCHOR-CONSTRAINED FRAMEWORK FOR FINE-GRAINED ERROR CLASSIFICATION AND RELIABILITY-AWARE QUALITY ESTIMATION IN ARABIC MACHINE TRANSLATION
No Thumbnail Available
Date
2026
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Saudi Digital Library
Abstract
Machine translation (MT) increasingly mediates how political conflict is detected and analysed across languages. For a morphologically rich and under-resourced language such as Arabic, evaluating MT quality remains hard: standard automatic metrics reduce a translation to a single score that cannot say what error occurred, where, or how severe it is. This dissertation contributes Raqeeb ("observer"), an anchor-constrained framework for fine-grained Arabic MT error classification and reliability-aware quality estimation.
Two motivating studies show that a domain-specific language model outperforms general-purpose models on political-conflict classification across English, Spanish, and Arabic, and that MT systematically strips rare, domain-specific vocabulary from the source text, a loss that, rather than translation quality, is a dominant factor in downstream performance. Evaluation must therefore pinpoint where and how a translation fails.
Raqeeb unifies five components. Mizan ("scale") is a 23-class Arabic MT error taxonomy with a four-level, MQM-compatible hierarchy and expert-calibrated severity weights. TRIVET (Triple Validation with Error Taxonomy) ties each candidate error to a specific keyword anchor, checks it with Arabic-aware structural methods, and confirms it with a cross-vendor audit. Together with a domain-contrastive keyword extraction and a two-layer blind annotation process, these yield QUANTA, a 9,032-instance Arabic MT error dataset covering all 23 classes, validated by two independent expert annotators at almost-perfect agreement on the core taxonomy criteria (Gwet's AC1 0.83 for real, 0.97 for synthetic). On a five-encoder benchmark, AraBERT v2 reaches a cross-validated 23-class macro-F1 of 0.614; without the anchor constraint performance collapses to 0.114, and expert-validated real data proves twelve times more sample-efficient than synthetic data.
Building on this dataset, RaqeebCOMET enhances the neural metric COMET with a severity penalty derived from the taxonomy, reaching Spearman 0.669 against a held-out human quality score where raw COMET reaches 0.011. A hierarchical fallback sustains 89.0% accuracy by emitting predictions at the finest reliable taxonomy level, never silently abstaining, and the structural gates show promising out-of-domain transfer.
Together these contributions form a coherent software pipeline from raw parallel corpus to validated, reliability-aware Arabic MT error analysis, released as the open-source Raqeeb toolkit. As a software-engineering contribution, the framework provides structural verification, defect classification, cross-vendor regression control, and reliability-aware triage for human-in-the-loop deployment.
Description
PhD dissertation in Software Engineering, The University of Texas at Dallas, USA. Degree conferred August 2026
Keywords
Arabic machine translation, machine translation evaluation, quality estimation, MT error analysis, error taxonomy, Mizan taxonomy, QUANTA dataset, Raqeeb framework, TRIVET, natural language processing, software engineering, reliability, الترجمة الآلية, معالجة اللغة العربية, تقييم جودة الترجمة
Citation
Alshammari, Afraa. Raqeeb: An Anchor-Constrained Framework for Fine-Grained Error Classification and Reliability-Aware Quality Estimation in Arabic Machine Translation. PhD dissertation, The University of Texas at Dallas, Richardson, TX, USA, August 2026
