Saudi Cultural Missions Theses & Dissertations
Permanent URI for this communityhttps://drepo.sdl.edu.sa/handle/20.500.14154/10
Browse
1 results
Search Results
Item Restricted AI-Enabled Autonomous Knowledge Extraction from Large-Scale Textual Data(Saudi Digital Library, 2026) Alharbi, Abdulrahman; Obradovic, ZoranThe rapid growth of large-scale textual data across social media platforms, news media, and scientific repositories presents both unprecedented opportunities and significant challenges for extracting meaningful insights. During global events such as the COVID-19 pandemic, understanding public discourse requires analyzing vast amounts of noisy, heterogeneous, dynamic, and geographically distributed data. At the same time, the exponential increase in scientific publications has made traditional evidence synthesis methods increasingly labor-intensive, time-consuming, and difficult to scale. Existing approaches to textual knowledge extraction often operate in isolation, lack interpretability, fail to integrate heterogeneous data sources, and do not support scalable end-to-end automation. This dissertation addresses these limitations by proposing a unified framework for AI-enabled autonomous knowledge extraction from large-scale textual data. The research introduces a comprehensive pipeline that integrate sentiment analysis, topic modeling, semantic interpretation, spatiotemporal reasoning, and multi-agent automation for scalable, robust and reproducible text analysis across heterogeneous domains. First, this work introduces TriLex, a novel unsupervised sentiment analysis framework that combines multiple lexicon-based sentiment analysis methods through weighted aggregation, majority voting, and dynamic thresholding technique to improve robustness and accuracy for short and noisy textual data. Building on this foundation, a hierarchical spatiotemporal framework is developed to capture the evolution of public sentiment across global, national, and regional scales. The framework integrates over 7 million social media posts and thousands of news articles to analyze COVID-19 vaccine discourse across time, geographic regions, and platforms. To enhance topic interpretability, this research integrates BERTopic with large language models (LLMs), enabling automated generation of coherent and context-aware topic representations for large-scale textual discourse. A cross-platform analytical framework is further introduced to examine temporal relationships between social media and news media discourse, demonstrating a bidirectional relationship in which news coverage and public discourse influence each other over time. Extending beyond discourse analysis, this dissertation introduces an Agentic AI framework that automates the end-to-end process of large-scale multilingual knowledge extraction and evidence synthesis. The proposed multi-agent system coordinates specialized agents for query generation, multilingual retrieval, metadata harmonization, title and abstract screening, and full-text analysis. Evaluated on a multilingual corpus of over 52,000 scientific records, the framework achieves high screening performance while substantially reducing processing time from months to hours, demonstrating significant improvements in scalability, robustness, and reproducibility. Collectively, this dissertation bridges the gap between analytical understanding and autonomous knowledge extraction from large-scale textual data. By integrating robust sentiment analysis, interpretable topic modeling, spatiotemporal discourse analysis, and autonomous AI systems within a unified framework, this work establishes a scalable and extensible paradigm for transforming heterogeneous textual data into actionable knowledge. The proposed methodologies are validated using real-world datasets spanning social media, news media, and scientific literature across diverse application domains.13 0
