brandburner/doctorwho-s06-narrative-kg
Doctor Who - Narrative Knowledge Graph A rich narrative knowledge graph extracted from Doctor Who screenplays using the Fabula pipeline. Contains characters, locations, objects, organizations, events, themes, and conflict arcs with full participation semantics and Graph Gravity importance tiers. Dataset Overview Metric Value Source database doctorwho.s06 Type Season database Episodes 44 Total nodes 6,743 Total edges 27,272 Schema version 1.2.0… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/doctorwho-s06-narrative-kg.
Doctor Who - Narrative Knowledge Graph
A rich narrative knowledge graph extracted from Doctor Who screenplays using the Fabula pipeline. Contains characters, locations, objects, organizations, events, themes, and conflict arcs with full participation semantics and Graph Gravity importance tiers.
Dataset Overview
Entity Breakdown
Graph Gravity Tiers
Relationship Types
AFFILIATED_WITH, BELONGS_TO_EPISODE, CALLBACK, CAUSAL, CHARACTER_CONTINUITY, CONTAINS_ACT, CONTAINS_BEAT, CONTAINS_SCENE, CREDITED_ON, EMOTIONAL_ECHO, ESCALATION, EXEMPLIFIES_THEME, FORESHADOWING, INVOLVED_IN_ARC, INVOLVED_WITH, IN_EVENT, NARRATIVELY_FOLLOWS, OCCURS_IN, PARTICIPATED_AS, PART_OF ... and 7 more
Related Datasets
This is a single-season dataset containing entities and events as extracted from Season 6 screenplays.
- Megagraph (all seasons unified): brandburner/doctorwho-mega-narrative-kg
Note: The megagraph is not a simple union of season datasets. Cross-season entities are reconciled through a Global Entity Registry (GER), receiving new canonical UUIDs and distilled descriptions. Graph Gravity tiers are recalculated across all episodes. Use individual season datasets for single-season analysis; use the megagraph for cross-season analysis.
Files
Schema
Nodes (nodes.parquet)
Edges (edges.parquet)
Positions (positions.parquet)
Layout method (schema ≥ 1.2.0): Coordinates are derived from the entities' semantic text embeddings (UMAP with a fixed seed and PCA initialization), so narratively similar entities sit near each other. Non-embedded nodes (events, scenes, episodes, etc.) are placed at the weighted barycenter of their narrative neighbours. The layout is deterministic: re-exporting an unchanged graph reproduces identical coordinates, and lightly-changed graphs keep comparable layouts. Not comparable with positions published under schema ≤ 1.1.0, which used a non-deterministic node2vec structural embedding. See meta.json → positions for the exact method and coverage stats.Usage
from datasets import load_dataset
import pandas as pd
# Load from HuggingFace
ds = load_dataset("brandburner/doctorwho-s06-narrative-kg")
# Or load parquet directly
nodes = pd.read_parquet("nodes.parquet")
edges = pd.read_parquet("edges.parquet")
# Filter to anchor characters
anchors = nodes[(nodes['primary_label'] == 'Agent') & (nodes['tier'] == 'anchor')]
# Build a NetworkX graph
import networkx as nx
G = nx.DiGraph()
for _, n in nodes.iterrows():
G.add_node(n['node_id'], label=n['primary_label'], name=n['name'])
for _, e in edges.iterrows():
G.add_edge(e['source_node_id'], e['target_node_id'], type=e['relationship_type'])Citation
@misc{fabula_doctorwho_s06,
title = {Doctor Who Narrative Knowledge Graph},
author = {Fabula Pipeline},
year = {2026},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/datasets/brandburner/doctorwho-s06-narrative-kg}}
}License
CC BY-SA 4.0
