Automated Motif Indexing On The Arabian Nights
No Thumbnail Available
Date
2026
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Saudi Digital Library
Abstract
There is a wealth of information that can be obtained in narratives. Cultural information is an example of what we can find. One form that cultural information can take is motifs. A motif is a recognizable recurring element or concept that frequently aids in the development of other literary or narrative elements like theme or mood. The importance of motifs, according to folklorists, is that they classify and locate stories from different cultures using motifs, and follow the evolution of stories across time. Beyond folklore, motifs also appear in news and social media such as tweets, which can tell us a lot about a culture. Until now, motifs are extracted manually from stories. A folklorist goes through a long and complicated process to identify motifs. Moreover, folklorists must have enough background in a culture to be able to understand motifs and extract them. I propose to develop a system that automatically detects thousands of motifs from text. In this dissertation, I focus on motif indexing: detecting and extracting motif expressions from narratives using an existing motif index as a guide.
My thesis has four aims: first, to find, collect, examine, and preprocess the dataset which consists of a list of motifs and folktales. Second, to build a baseline model to automatically detect and understand motifs in text using a handful list of motifs to generate candidate motif expressions for dataset construction. Third, to annotate the dataset, which will be used in training state-of-the-art LLMs. Fourth, to develop a model that automatically detects the entire motif index.
The first step of my work is the dataset used to build a culturally aware LLM, including the source of motifs and the narratives that contain them. I showed how we processed both sources and narratives, and my work on motif classification (Simple to Complex motifs).
The second step is a large-scale annotation of 58,450 positive and negative motif expressions generated from the baseline system.
The third step is fine-tuning state-of-the-art LLMs to automatically index motifs from narratives. The fine-tuned LLM achieved strong performance compared to the baseline models, which highlights the necessity of continued training with larger datasets that contain cultural elements, in order to improve AI systems’ ability to understand culture and its many facets.
Finally, I describe using the fine-tuned model to automatically index 5,362 motif expressions from the Arabian Nights. This work demonstrates automatic motif indexing at the scale of an entire motif index. This dissertation contributes resources and methods that support building AI systems that are more culturally aware by enabling large-scale detection and indexing of motifs in narrative text.
Description
Keywords
Folklore, motifs, automated motif indexing, the Arabian Nights, natural language processing, neural methods, large language models, linguistic annotation
