Browsing by Author "Alshammari, Afraa"
Now showing 1 - 1 of 1
- Results Per Page
- Sort Options
Item Restricted RAQEEB: AN ANCHOR-CONSTRAINED FRAMEWORK FOR FINE-GRAINED ERROR CLASSIFICATION AND RELIABILITY-AWARE QUALITY ESTIMATION IN ARABIC MACHINE TRANSLATION(Saudi Digital Library, 2026) Alshammari, Afraa; Khan, LatifurMachine translation (MT) increasingly mediates how political conflict is detected and analysed across languages. For a morphologically rich and under-resourced language such as Arabic, evaluating MT quality remains hard: standard automatic metrics reduce a translation to a single score that cannot say what error occurred, where, or how severe it is. This dissertation contributes Raqeeb ("observer"), an anchor-constrained framework for fine-grained Arabic MT error classification and reliability-aware quality estimation. Two motivating studies show that a domain-specific language model outperforms general-purpose models on political-conflict classification across English, Spanish, and Arabic, and that MT systematically strips rare, domain-specific vocabulary from the source text, a loss that, rather than translation quality, is a dominant factor in downstream performance. Evaluation must therefore pinpoint where and how a translation fails. Raqeeb unifies five components. Mizan ("scale") is a 23-class Arabic MT error taxonomy with a four-level, MQM-compatible hierarchy and expert-calibrated severity weights. TRIVET (Triple Validation with Error Taxonomy) ties each candidate error to a specific keyword anchor, checks it with Arabic-aware structural methods, and confirms it with a cross-vendor audit. Together with a domain-contrastive keyword extraction and a two-layer blind annotation process, these yield QUANTA, a 9,032-instance Arabic MT error dataset covering all 23 classes, validated by two independent expert annotators at almost-perfect agreement on the core taxonomy criteria (Gwet's AC1 0.83 for real, 0.97 for synthetic). On a five-encoder benchmark, AraBERT v2 reaches a cross-validated 23-class macro-F1 of 0.614; without the anchor constraint performance collapses to 0.114, and expert-validated real data proves twelve times more sample-efficient than synthetic data. Building on this dataset, RaqeebCOMET enhances the neural metric COMET with a severity penalty derived from the taxonomy, reaching Spearman 0.669 against a held-out human quality score where raw COMET reaches 0.011. A hierarchical fallback sustains 89.0% accuracy by emitting predictions at the finest reliable taxonomy level, never silently abstaining, and the structural gates show promising out-of-domain transfer. Together these contributions form a coherent software pipeline from raw parallel corpus to validated, reliability-aware Arabic MT error analysis, released as the open-source Raqeeb toolkit. As a software-engineering contribution, the framework provides structural verification, defect classification, cross-vendor regression control, and reliability-aware triage for human-in-the-loop deployment.3 0
