datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wolfhall-mega-narrative-kg
Wolf Hall - Narrative Knowledge Graph
A rich narrative knowledge graph extracted from Wolf Hall screenplays using the
Fabula pipeline. Contains characters,
locations, objects, organizations, events, themes, and conflict arcs with full
participation semantics and Graph Gravity importance tiers.
Dataset Overview
Metric
Value
Source database
wolfhall.mega
Type
Megagraph (cross-season merged)
Episodes
12
Seasons merged
1, 2
Total nodes
4,472
Total edges… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/wolfhall-mega-narrative-kg.wolfhall-s02-narrative-kg
Wolf Hall - Narrative Knowledge Graph
A rich narrative knowledge graph extracted from Wolf Hall screenplays using the
Fabula pipeline. Contains characters,
locations, objects, organizations, events, themes, and conflict arcs with full
participation semantics and Graph Gravity importance tiers.
Dataset Overview
Metric
Value
Source database
wolfhall.s02
Type
Season database
Episodes
6
Total nodes
1,973
Total edges
5,929
Schema version
1.1.0
Exported… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/wolfhall-s02-narrative-kg.wolfhall-s01-narrative-kg
Wolf Hall - Narrative Knowledge Graph
A rich narrative knowledge graph extracted from Wolf Hall screenplays using the
Fabula pipeline. Contains characters,
locations, objects, organizations, events, themes, and conflict arcs with full
participation semantics and Graph Gravity importance tiers.
Dataset Overview
Metric
Value
Source database
wolfhall.s01
Type
Season database
Episodes
6
Total nodes
2,691
Total edges
9,242
Schema version
1.1.0
Exported… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/wolfhall-s01-narrative-kg.wolf-quotes
Russian Wolf Quotes Dataset
Overview
This dataset contains 95k Russian wolf quotes. The original raw 340 MB file was found in the neurovolk repository. The goal of this project is to split that stream into individual quotes, remove exact duplicates, filter near-duplicates at the character level, and save the result as a clean CSV ready for reuse.
After splitting, the dataset contains 2,323,394 raw quote fragments. Exact deduplication leaves 100,545 unique quotes. MinHash… See the full description on the dataset page: https://huggingface.co/datasets/pymlex/wolf-quotes.wolfgang-lm-synthetic-chat-v1
Wolfgang-LM: Synthetic Goethe Chats (v1)
This dataset contains ~4,500 synthetic conversations designed to fine-tune language models into the persona of Johann Wolfgang von Goethe. It was generated as part of the Wolfgang-LM project.
Dataset Details
Size: ~4,500 samples
Language: German (Modern User vs. Historical Goethe)
Format: JSONL (ShareGPT compatible messages list)
License: MIT License
Generator Model: Google Gemini 2.5 Flash
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/dominik-reiner/wolfgang-lm-synthetic-chat-v1.
