Artificial intelligence for bias detection in higher education online content

dc.contributor.advisorEric, Atwell
dc.contributor.advisorNoorhan, Abbas
dc.contributor.authorBin Shiha, Rawan
dc.date.accessioned2026-07-12T21:05:52Z
dc.date.issued2026
dc.description.abstractThis thesis develops and evaluates Artificial Intelligence (AI) and Natural Language Processing (NLP) approaches for detecting bias in higher education online resources, addressing the lack of systematic computational methods in this area. Bias in higher education content shapes knowledge production, representation, and equity, making its detection both academically and socially significant. To address this gap, three sets of novel datasets were created: 1. a corpus of university news articles annotated for subjectivity, sentiment, and gender representation; 2. three university reading list datasets with demographic annotations enabling comparative analysis across Western and Middle Eastern contexts; and 3. two domain-specific collections of learning materials, one from humanities-oriented open resources and the other from Science, Technology, Engineering, and Mathematics (STEM) lecture transcripts. Together, these datasets provide the first systematic resources for investigating representational, stereotypical and linguistic bias across diverse higher education domains. Using these resources, a range of Pre-trained Language Models (PLMs) and Large Language Models (LLMs) were evaluated for bias detection. PLM revealed significant gendered and representational disparities in university discourse, while fine-tuned LLM achieved improved performance on humanities data but showed limited transferability to STEM materials. A hybrid framework integrating fine-tuned LLMs with Retrieval-Augmented Generation (RAG) enhanced detection transparency and prediction balance. These methods were operationalised in a prototype web application for bias detection in higher education learning content. The contributions of this research are fourfold: 1. the creation of multiple novel annotated datasets spanning university news, reading lists, and academic learning resources; 2. the introduction of LLM-based strategies for bias annotation, fine-tuning, and cross-domain evaluation 3. methodological innovations including a structured framework for categorising bias, a hybrid human–AI annotation approach, and a replicable NLP pipeline for demographic and thematic analysis; and 4. the design of a hybrid bias detection system and accompanying web application. This work advances computational approaches to bias detection, provides reproducible resources for future research, and offers practical tools to support greater fairness and equity in higher education online environments.
dc.format.extent266
dc.identifier.urihttps://hdl.handle.net/20.500.14154/79522
dc.language.isoen
dc.publisherSaudi Digital Library
dc.subjectArtificial Intelligence
dc.subjectNatural Language Processing
dc.subjectLarge Language Models
dc.subjectRetrieval-Augmented Generation
dc.subjectEducation Bias
dc.titleArtificial intelligence for bias detection in higher education online content
dc.typeThesis
sdl.degree.departmentComputer Science
sdl.degree.disciplineComputer Science
sdl.degree.grantorUniversity of Leeds
sdl.degree.nameDoctor of Philosophy

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
SACM-Dissertation.pdf
Size:
3.53 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
1.61 KB
Format:
Item-specific license agreed to upon submission
Description:

Copyright owned by the Saudi Digital Library (SDL) © 2026