Molecules of a Story: Community Detection in PMI-weighted Narrative Networks
Authors: Kasper Fyhn, Rebekah Baglini
Organizations: Department of Linguistics and Cognitive Science, Aarhus University · TEXT - Center for Contemporary Cultures of Text, Aarhus University · Center for Humanities Computing, Aarhus University
Automatically extracted narrative networks -- graphs with entities as nodes and their relations as edges -- have proven useful for revealing central narrative structures through salient entities and their connections (Tangherlini et al. 2020; Labatut and Bost 2019). But a narrative is more than those central structures that everything else revolves around. This work is concerned with the everything else: brief sub-plots, small clusters of descriptions, or associations between minor characters that go under the radar at the macro-level. We present an approach to unearth such peripheral structures. They involve rare entities with limited textual presence, overshadowed by dominant entities and lost among each other in the long tail of many but rare entities (Baayen 2001). We leverage the known tendency of pointwise mutual information (PMI, Church and Hanks 1990) to inflate for rare events, turning its weakness into a strength by weighting edges with PMI to foreground peripheral entity configurations. Communities extracted from the resulting network are structural traces of underlying narrative elements. We demonstrate the approach on The Lord of the Rings. From measures of how concentrated or dispersed a community's activations are across the text, a typology emerges that reveals that peripheral structures form more than a single class: episodic passages, echoing long-distance connections, and recurring threads each surface as distinct configurations. The approach is conceptually simple and surfaces fine-grained narrative details that are lost in abundance, though its deliberate amplification of weak signals comes with inherent sensitivity -- best understood as a lens for exploration rather than a robust extraction pipeline.
Figures & tables
Figure 1: Mean PMI by entity frequency in The Lord of the Rings . Pairs with low frequency entities generally have a much higher PMI than pairs with high frequency entities.
Figure 2: Conceptual illustration of the three prototypical peripheral structures. Spikes (vertical lines) represent documents where part of a community is activated. The horizontal axis represents the sequential or temporal ordering of documents. (a) Episodic structures have concentrated activations within a limited span. (b) Echoing structures have few activations at distant points. (c) Recurring structures have multiple activations spread across the text.
No. of communities
Size
Mean
Min
Q1
Median
Q3
Max
k-clique
1193
7.22
4
4
6
8
93
Louvain
797
7.0
2
4
6
9
24
Table 1: Descriptive statistics of the number and sizes of communities extracted from The Lord of the Rings
Spread
Spikes
Function
Community
Gloss
Episodic
Low
Few
Description
the stairway, the high regions, the main pass, the great ravine, the torment
Description of the climb and stairway in the mountains of the Morgul Vale into Mordor.
Low
Few
Event
your lordship, the herb-master, kingsfoil, latter days, old names
Aragorn’s request for kingsfoil in the Houses of Healing, with which the herb-master returns two sections later.
Low
Multiple
Description (sustained)
great height, a great circle, fosse, soft shadow, mallorn-trees, its brink, Green walls, ten miles, golden elanor, the white bridge, green hill, Their height
Description of areas and landmarks in Lothlorien, spanning multiple sections with short gaps between.
Echoing
High
Few
Poem variation
A new road, a secret gate, the hidden paths, Apple, the eastern sky, his memory
Distinct versions of a song sung by Frodo in the same geographical location early in book one and late in book three.
Table 2: Representative examples of the community typology by spread and spikes
In disinformation datasets, narratives are often understood as recurring interpretive patterns that group texts under narrative labels. Recent work formalized narrative mining as inductively inferring narrative labels from corpora, but its evaluation stays tied to predefined taxonomies, a closed-world setting that cannot capture narratives absent from the reference labels. We introduce a three-tier evaluation framework for unsupervised narrative label generation: recovery (against a corpus's own taxonomy), mining (against external label sets), and discovery (without predefined labels). Applying it, we compare clustering-based and graph-community-based pipelines across seven disinformation datasets, with human validation of discovery on two. The two families are complementary under automated metrics, but in a corpus with two prominent topics, clustering can reduce one topic to 2% of generated labels while graph-based pipelines stay balanced. Discovery validation also reveals many singletons (narrative labels derived from single claims, 30-62% of graph outputs), which clustering cannot produce. Annotators confirm many as recognizable disinformation narratives, suggesting that in open-world discovery the repetition assumed by narrative mining may be recognized outside the corpus, not within it. We release human-validated narrative candidate labels for the Climate Obstruction and PolyNarrative datasets to support taxonomy development and dataset extension.
Max Upravitelev, Veronika Solopova, Jing Yang +4
Technische Universität Berlin · German Research Center for Artificial Intelligence (DFKI) · BIFOLD – Berlin Institute for the Foundations of Learning and Data +2
Narrative extraction allows us to identify online hate narratives, supporting the construction of rigorous detection systems. Existing computational approaches, however, are limited in precision as they rely on semantic representations, which tend to capture only surface-level meaning. To detect more precise and interpretable narratives, we present an extraction pipeline that represents narratives as entity-evaluation pairs. Narratives are extracted using a Large Language Model (LLM) reasoning process that extends Aspect-Based Sentiment Analysis, identifying the aspect, classifying its judgement type as the basis for evaluation, and deriving the evaluation accordingly. Extracted narratives are then clustered using Leiden, following which clusters are resolved to an intended level of granularity through an LLM-guided refinement process. We illustrate this narrative pipeline with English Reddit comments from 2024 that criticize Taylor Swift, analyzing a representative cluster that exhibits hate speech patterns to demonstrate its interpretive value.
Annabelle K. L. Chua, Forster J. Khoo, Joel C. R. Tan +8
DSO National Laboratories, 14 Science Park 118226, Singapore · Defence Science and Technology Agency, 1 Depot Road 109679, Singapore
The narrative composition of web-scale LLM pretraining corpora remains largely unexplored, even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token open pretraining corpus. Drawing on narrative theory, we design a framework spanning three core narrative elements (agency, setting, and events) operationalized as 11 interpretable dimensions. After curating and hand-annotating a diverse set of 400 passages, we create an LLM-labeled dataset of 25K passages, and finally, we finetune and validate NarraBERT, two RoBERTa-based models for fine-grained narrative prediction. We apply NarraBERT to 13M passages, resulting in a new dataset, NarraDolma. We find that narrative structure is measurable at scale across extremely heterogeneous data and narrative qualities are unequally distributed across pretraining sources, topics, and formats in ways that current data curation practices neither measure nor account for. Our framework, dataset, and analyses provide a foundation for understanding how narrative qualities are distributed in LLM pretraining data.
Teagan Johnson, Elliott Ash, Andrew Piper +1
University of Colorado Boulder · ETH Zürich · McGill University