SACM - United States of America
Permanent URI for this collectionhttps://drepo.sdl.edu.sa/handle/20.500.14154/9668
Browse
12 results
Search Results
Item Restricted AI Systems for Understanding and Grounding Radiology Reports(Saudi Digital Library, 2026) Baharoon, Mohammed; Pranav, RajpurkarRadiology reports are the primary medium through which radiologists communicate, conveying critical clinical information in natural language. AI holds the potential to both generate and analyze these reports, yet two key challenges persist. First, as AI systems increasingly generate radiology reports, evaluating their accuracy remains an open problem. Second, a fundamental disconnect exists between the findings described in radiology reports and their corresponding locations in the imaging studies, limiting referring physicians, patients, and trainees who must interpret findings without explicit visual guidance. This thesis addresses these challenges through three interconnected contributions spanning report evaluation, report visualization, and clinical application. First, we introduce CRIMSON, a clinically grounded evaluation metric for radiology report generation, along with two new benchmarks: RadJudge, a 30-case clinical judgment test suite, and RadPref, a 100-case radiologist preference benchmark. CRIMSON incorporates patient context and weights errors by clinical significance when comparing generated against reference reports. Validated against radiologist error counts from the ReXVal dataset, RadJudge, and RadPref, CRIMSON achieves stronger alignment with expert judgment than prior metrics. Second, we introduce ReXGroundingCT, the first publicly available dataset linking free-text radiology findings to manually annotated 3D segmentation masks in chest CT scans. Designed through a multi-stage annotation pipeline, the dataset comprises 3,142 scans and 8,028 segmented findings. We benchmark state-of-the-art text-prompted segmentation models, demonstrating that current approaches fall substantially short of clinical utility, even after fine-tuning. Subsequently, we release the dataset, along with a public leaderboard to drive continued progress on this task. Third, we present RadGame, an AI-powered platform for radiology education that brings together both report evaluation and grounding as a concrete downstream application. The platform teaches two core skills: localizing findings through interactive bounding-box annotation, and writing radiology reports with automated structured feedback. In a prospective multi-institutional study with 18 medical students, participants using RadGame achieved a 68\% improvement in localization accuracy and a 31\% improvement in report-writing scores, outperforming traditional passive learning methods. Together, these contributions address the radiology AI pipeline from multiple angles: evaluating generated reports against clinical standards, spatial grounding of reports in imaging, and applying both to advance radiology education.29 0Item Restricted Evaluation of Modified AI-Enhanced Radiographic Images of Artificial Teeth for Caries Removal Decision-Making in Predoctoral Dental Students(Saudi Digital Library, 2026) Aldandan, Sukaina; Fontana, Margherita; Neiva, GiseleBackground: Accurate radiographic interpretation is essential for caries removal decision-making but remains challenging for early dental learners. Artificial intelligence (AI) has demonstrated promise in improving caries detection; however, its role in supporting operative decision-making during preclinical training remains unclear. Objective: To evaluate the effect of modified AI-enhanced radiographic images on the quality of caries removal performed by first-year dental students and to determine whether the effect of modified AI-enhanced images differs between shallow and deeper lesions. Methods: A two-period crossover study was conducted involving first-year dental students in a preclinical operative dentistry course. Participants performed caries removal on standardized 3D-printed teeth containing either shallow or deeper carious lesions. Students completed procedures using either standard bitewing radiographs or modified AI-enhanced radiographic images. Caries removal quality was assessed using the Composite Caries Removal Quality Score (CRQS), which incorporated convenience form, caries removal at the dentinoenamel junction (DEJ), caries removal at the pulpal floor. Completion time was also recorded. Results: Modified AI-enhanced radiographic images significantly improved overall CRQS compared with standard radiographs (p = 0.002). This improvement was primarily observed in shallow lesions, which demonstrated significantly higher CRQS scores under the modified AI-enhanced condition (p = 0.001), whereas no significant difference was found for deeper lesions (p = 0.727). The greatest improvement was observed in convenience form for shallow lesions (p < 0.001). No significant differences were detected for caries removal at the DEJ or pulpal floor. Lesion depth significantly influenced several outcomes, with deeper lesions demonstrating lower DEJ scores and requiring longer completion times. Modified AI-enhanced images did not significantly affect completion time. Conclusions: Modified AI-enhanced radiographic images improved the quality of caries removal performed by first-year dental students, particularly for shallow lesions where radiographic interpretation is more challenging. The benefits were primarily related to improved convenience form rather than caries removal at the DEJ or pulpal floor. These findings suggest that modified AI-enhanced radiographic images may be a valuable adjunct in preclinical dental education and may support the development of diagnostic and operative decision-making skills in early learners.9 0Item Restricted Toward Robust Mental Health Classification Systems Across Genres and Languages(Saudi Digital Library, 2026) Alqahtani, Amal Abdullah; Diab, Mona; Hwa, RebeccaMental health conditions are a major global public health challenge, yet many individuals do not receive appropriate care because of stigma, limited access to services, and the difficulty of accurate assessment. Natural Language Processing (NLP) has shown growing promise for identifying mental health conditions through language, but existing systems often struggle to generalize across modalities, domains, conditions, and languages. Existing approaches leave critical gaps in condition specificity, cross-genre robustness, and multilingual coverage. Prior work often studies isolated features or a single modality, leaving open how language markers behave across both writing and speech for the same condition. Condition-specific continual pretraining remains underexplored relative to generic mental health adaptation. Cross-condition transfer from clinically comorbid disorders has been proposed but rarely validated. And the field remains overwhelmingly English-centric, with Arabic among the most underserved languages despite its more than 400 million speakers. This dissertation addresses these gaps through a progression from interpretable linguistic analysis to multilingual evaluation, using schizophrenia as a core case study. We first present an integrated analysis of cohesion features, pragmatic cues, and language model-based measures across clinical speech and writing, showing that patients exhibit heightened fear, higher neuroticism, reduced specificity, and lower cohesion, with effects generally stronger in writing. We then evaluate these signals through supervised classification, finding that cohesion is the strongest standalone structured feature view in writing, while a TF-IDF lexical baseline dominates in speech. Moving to neural modeling, we show that progressive multi-stage continual training of BERT on patient-generated social media achieves an 11.7% relative F1 improvement over base BERT and outperforms MentalBERT and ClinicalBERT for schizophrenia detection. We then demonstrate that focused cross-condition transfer outperforms broad mental health pretraining, with StressRoBERTa achieving 82% F1 on the SMM4H 2022 stress detection benchmark. To extend mental health NLP beyond English, we introduce ArMHC, a large-scale Arabic mental health corpus from X (formerly Twitter) constructed through a dialect-aware extraction pipeline with LLM-based validation, covering 18 conditions across 1,911 users. Using the ArMHC schizophrenia subset, we evaluate cross-lingual and cross-genre transfer from English clinical data to Arabic social media, finding that both language and genre mismatch contribute substantially to transfer degradation, with genre mismatch being qualitatively more destructive: cross-lingual same-genre transfer still permits partial detection, while cross-genre transfer falls below chance. Overall, this dissertation demonstrates that robust mental health NLP benefits from combining interpretable linguistic analysis with domain-adaptive and transfer-based modeling, while expanding into low-resource multilingual settings. The findings contribute new linguistic evidence, modeling strategies, and dataset resources for building more inclusive and clinically relevant computational approaches to mental health assessment.20 0Item Restricted The Influence of Artificial Intelligence on EAP Learners’ Oral Fluency(Saudi Digital Library, 2026) Alwadaeen, Norah Bakheet; Abbuhl, RebekhaThere is ongoing debate on how AI speaking tools can support the development of oral fluency in second language (L2) instruction. Despite the widespread usage of these tools, such as AI chatbots and Automated Speech Recognition (ASR), questions persist about how well they will work to improve oral fluency, reduce speaking anxiety, and foster learner autonomy. This study investigates how an AI-mediated speaking partner influences English for Academic Purposes (EAP) learners’ oral fluency, speaking anxiety, and autonomy over a short, intensive practice cycle. Five upper-intermediate ESL students at a California community college completed nine EAP Talk chatbot sessions across 3 weeks, framed by pre- and post-intervention IELTS-style monologic speaking tasks. Acoustic analyses of the pre/post tasks in PRAAT targeted three utterance-fluency indices (speaking ratio, repair phenomena, and pause placement). Session-by-session Likert questionnaires captured perceived fluency gains, anxiety, and autonomy, and post-intervention semi-structured interviews explored learners’ experiences with the AI-mediated practice. Oral fluency findings indicated that the speaking ratio increased, whereas pause and repair indices generally shifted in favorable directions. Anxiety, which was scaled so higher scores indicated less anxiety, exhibited clear gains. Autonomy trajectories were positive at the group level. Furthermore, the study highlights both the promise and limitations of AI chatbots for EAP speaking. It emphasizes the value of multi-indicator fluency assessment, explicit autonomy supports, and longer comparative designs in future work.38 0Item Restricted An Analysis of Face Synthesis Methods and Their Influence on Human Perception(Saudi Digital Library, 2025) Almaimani, Maha; Patterson, EricSynthetic faces (e.g., computer-generated characters) have been increasingly utilized across various fields, including entertainment, healthcare, and education. Perceptual studies are often con- ducted to understand how synthetic faces are perceived by humans, aiming to enhance both quality and user experience in these domains. Over the years, numerous methods have been developed to create synthetic faces, ranging from traditional techniques such as image composites, Active Ap- pearance Models, and 3D Morphable Models to more recent machine-learning-based frameworks like Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs). Despite the growing adoption of synthetic face generation and the variety of algorithms available for their creation, cognitive scientists have underutilized these advanced techniques. Re- search in this area has remained largely dependent on earlier approaches, and even when Artificial- Intelligence-based (e.g., AI-based) methods are employed, there is an overreliance on GANs, particu- larly StyleGAN2. This trend highlights a significant gap in the exploration of alternative generative architectures in perceptual studies. Furthermore, existing studies have primarily focused on the technical performance of GANs and VAEs, while human perception of their outputs has remained underexplored. This dissertation seeks to address this gap by first providing a comprehensive review of syn- thetic face generation methods from a perceptual standpoint. Second, it analyzes and perceptually compares two prominent AI-based models: StyleGAN2 (a GAN variant) and NVAE (a VAE flavor) across multiple contexts (e.g., full scenes vs. isolated faces without background, and animated vs. static faces) to determine how these conditions influence perceived realism and trustworthiness. This comparison supports the development of cognitive research that advances the generation of percep- tually engaging and practically useful synthetic faces. Finally, conducting a study to investigate how truthful versus misleading medical information about dementia influenced participants’ perceptions when viewing videos of synthetic faces generated by StyleGAN2. The outcomes of this dissertation provide insights into the perceptual differences between GAN-based and VAE-based synthetic faces across diverse contexts. Understanding these distinctions will contribute to the responsible and effective application of synthetic faces in real-life applications.32 0Item Restricted INVESTIGATING NOVEL ANALYSIS APPROACHES FOR STRUCTURAL CONDITION ASSESSMENT USING ULTRASOUND AND INFRARED DATA(Saudi Digital Library, 2025) Alqurashi, Inad; Catbas, NecatiAging civil infrastructure, particularly reinforced concrete bridges, is experiencing progressive deterioration that threatens safety, serviceability, and long-term performance. Traditional inspection methods such as visual examination and hammer sounding are limited in their ability to detect subsurface defects and are prone to subjectivity. This dissertation develops and validates an integrated, multi-modal structural condition assessment framework that combines rapid Infrared Thermography (IRT), high-resolution Ultrasound Tomography (UT), Artificial Intelligence (AI)-driven anomaly detection, and immersive Digital Twin (DT) visualization to overcome these limitations. The research advances three main areas: (1) a dual-mode IR–UT workflow exploiting the complementary strengths of each modality, enabling rapid surface screening with IRT and in-depth defect characterization with UT; (2) optimized deep learning (DL) models tailored to each modality, with a transformer-based Grounding DINO model applied to raw Infrared (IR) imagery for automated detection of thermal anomalies, and a lightweight You Only Look Once (YOLO)-v8n model applied to UT volumetric slices for detecting internal delaminations, voids, ducts, and rebar, both trained on large, segmentation-assisted, color-standardized datasets to ensure robust performance under diverse field conditions; and (3) integration of Unmanned Aerial Vehicle (UAV)-based Light Detection and Ranging (LiDAR), photogrammetry, and multi-modal non-destructive testing (NDT) data into a geo-referenced Virtual Reality (VR) environment to support real-time, collaborative decision-making. Laboratory testing on engineered specimens with embedded defects and field deployment on multiple in-service bridges, including the NASA Causeway Bridge, achieved high detection accuracy (mAP@0.5 up to 0.93 for UT using YOLOv8n and 0.80 for IRT using Grounding DINO), strong localization (Average IoU ≈ 0.80–0.90), and significant efficiency gains through targeted UT scanning. The VR-based DT enabled inspectors to seamlessly review thermal anomalies, volumetric UT slices, and 3D geometry in a single immersive scene, reducing defect confirmation time from several minutes to approximately one minute per location. By fusing complementary NDT modalities with AI models purpose-built for each data type and immersive visualization, this research delivers a scalable, repeatable, and field-validated methodology for rapid, objective, and data-rich condition assessment of reinforced concrete structures, with potential for broader application to other infrastructure types to enable proactive maintenance strategies and improved lifecycle management.19 0Item Embargo AI-Enabled Bioresponsive Clinical Decision Support Systems for Chronic Pain: User-Centered Approach(Saudi Digital Library, 0025) Alrefaei, Doaa; Soussan, DjamasbiThe advancement of eye-tracking technologies has enabled the development of systems capable of detecting attention and cognitive states objectively and in real time. Biometric technologies that capture psychological measures, such as eye movements (EMs), have allowed user experience (UX) research to expand toward building smart bioresponsive tools. One area that may benefit from these advancements is chronic pain, where self-report methods are often limited in capturing the complex phenomenon of chronic pain experience in both research and practice. This has established a need for objective biomarkers that can support pain assessment. Pain literature suggests the use of EMs as potential biomarkers, as they reflect pain-related attentional patterns. This dissertation adopts a bioresponsive, UX research approach to explore the efficacy of using EMs to detect pain experience in individuals with and without chronic pain. A proof-of-concept AI tool was developed to detect chronic pain using only EMs from individuals with and without chronic pain, achieving an accuracy of 81%, thereby demonstrating the robustness of EMs as a potential biomarker for pain. To successfully evolve this proof of concept into a fully developed and effective Clinical Decision Support System (CDSS) for chronic pain treatment and management, it is essential to understand the needs of the healthcare professionals who will use the system. As a first step, traditional UX research methods were employed to conduct interviews with healthcare professionals involved in the treatment and management of chronic pain. Based on this research, six user personas, four representing doctors and two representing nurses, were developed to serve as a foundational guideline for the design of an initial CDSS prototype. The findings of this dissertation contribute to both UX research and pain science by presenting a comprehensive methodology for using eye movements (EMs) as input signals to an AI tool capable of detecting differences in attentional patterns toward pain-related stimuli. It also contributes to clinical practice by outlining design guidelines for developing an initial prototype of such an AI-based CDSS, grounded in the needs and workflows of healthcare professionals.20 0Item Restricted TOWARDS ROBUST AND ACCURATE TEXT-TO-CODE GENERATION(University of Central Florida, 2024) almohaimeed, saleh; Wang, LiqiangDatabases play a vital role in today’s digital landscape, enabling effective data storage, manage- ment, and retrieval for businesses and other organizations. However, interacting with databases often requires knowledge of query (e.g., SQL) and analysis, which can be a barrier for many users. In natural language processing, the text-to-code task, which converts natural language text into query and analysis code, bridges this gap by allowing users to access and manipulate data using everyday language. This dissertation investigates different challenges in text-to-code (including text-to-SQL as a subtask), with a focus on four primary contributions to the field. As a solution to the lack of statistical analysis in current text-to-code tasks, we introduce SIGMA, a text-to- Code dataset with statistical analysis, featuring 6000 questions with Python code labels. Baseline models show promising results, indicating that our new task can support both statistical analysis and SQL queries simultaneously. Second, we present Ar-Spider, the first Arabic cross-domain text-to-SQL dataset that addresses multilingual limitations. We have conducted experiments with LGESQL and S2SQL models, enhanced by our Context Similarity Relationship (CSR) approach, which demonstrates competitive performance, reducing the performance gap between the Arabic and English text-to-SQL datasets. Third, we address context-dependent text-to-SQL task, often overlooked by current models. The SParC dataset was explored by utilizing different question rep- resentations and in-context learning prompt engineering techniques. Then, we propose GAT-SQL, an advanced prompt engineering approach that improves both zero-shot and in-context learning experiments. GAT-SQL sets new benchmarks in both SParC and CoSQL datasets. Finally, we introduce Ar-SParC, a context-dependent Arabic text-to-SQL dataset that enables users to interact with the model through a series of interrelated questions. In total, 40 experiments were conducted to investigate this dataset using various prompt engineering techniques, and a novel technique called GAT Corrector was developed, which significantly improved the performance of all base- line models.38 0Item Restricted Scalex: Scalability Exploration of Multi-Agent Reinforcement Learning Agents in Grid-Interactive Efficient Buildings(Saudi Digital Library, 2023-08-11) Almilaify, Yara; Nagy, ZoltanTransitioning to renewable energy and decarbonization presents challenges for grid-interactive efficient building (GEB) communities. Conventional control systems struggle to maximize intermittent renewable energy, but advanced control architecture and utilization of renewable sources with energy storage can overcome this limitation and optimize energy flexibility. Reinforcement learning (RL) offers potential solutions, but its scalability and computational demands in large-scale settings remain unclear. This paper examines the scalability of Soft-Actor Critic (SAC) in multi-agent systems, comparing decentralized-independent SACs and centralized SACs using CityLearn, an OpenAI Gym environment. We consider neighborhoods consisting of 2 to 64 single-family residential buildings, each equipped with cooling and heating storage devices, domestic hot water storage devices, electrical storage devices, and solar PV systems. Our findings suggest that independent controllers outperform the centralized controller with increasing number of buildings. We also show that the performance on the building level can differ from the aggregated performance.53 0Item Restricted THE EFFECT OF USING A TECHNOLOGY BASED SELF-MONITORING INTERVENTION ON ON-TASK BEHAVIOR FOR STUDENTS WITH BEHAVIORAL ISSUES IN AN INCLUSIVE CLASSROOM(2023-08) Algethami, Sami; Vasquez, EleazarThis study examined the effectiveness of using a technology-based self-monitoring intervention called Monitoring Behavior on the Go (MoBeGo). On-task behavior for students with behavioral issues was the primary dependent variable in the study. The researcher employed a single-subject withdrawal design (ABAB) with two generalization phases (C-D) to investigate the ability of MoBeGo to generalize the results to a different setting. Visual analysis of graphs revealed the participants had a clear functional relationship between MoBeGo and percentage of on-task behavior. The finding illustrated on-task behaviors in a different setting did not increase without using MoBeGo and therefore no automatic generalization occurred in different settings. A replicated phase (D) was conducted to confirm the finding, and the results showed the percentage of on-task behavior increased in math and science classes which used MoBeGo and did not increase in reading/writing which did not use MoBeGo. Also, the outcomes showed MoBeGo has a high level of acceptability among teachers who participated in the study. The researcher evaluated this single-subject withdrawal design (ABABCD) by using the What Works Clearinghouse (WWC) evidence standards. In addition, the researcher utilized the Single-Case Analysis and Review Framework (SCARF) to evaluate the study outcomes. The evaluation results of using WWC and SCARF are discussed in Chapter 4. The researcher discussed major lessons learned and some limitations of using technology based self-monitoring (TBSM). In addition, implications for practitioners, researchers, and application developers were included as future directions for using TBSM. Moreover, the researcher discussed the potential role of self-monitoring-based artificial intelligence (SMBAI) in education, and the use of artificial intelligence (AI), large language models (LLMs), or machine learning (ML) with self-monitoring apps. Finally, some important questions were raised about protecting privacy and minimizing the risk of data breaches for individuals, and how to ensure the security of individuals’ data.51 0
