datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fictionrag-datasetVellumK2T-Fiction-SFT-01
Dataset Card for VellumK2T-Fiction-SFT-01
A long-form synthetic creative fiction dataset with 8,042 instruction–output pairs for supervised fine-tuning (SFT), generated using the VellumForge2 pipeline and published as part of the VellumForge2 fantasy collection on Hugging Face.
Dataset Details
Dataset Description
VellumK2T-Fiction-SFT-01 is a synthetically generated dataset of various fiction writing samples. Each row contains:
An instruction: a rich… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01.2026-08-27-good-ai-fiction-sf-860
synth good_ai_fiction run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth good_ai_fiction run — per-stage snapshots (resumable generation cache)
date_generated
20260828_020624
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @ ae0725130a2fccd74fe7bdef5c570ec71420cd7b
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-good-ai-fiction-sf-860.fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.gallica_literary_fictions
Dataset Card for Literary fictions of Gallica
Dataset Summary
The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.flawed-fictionslang-uk-fiction-gec-dialogs
Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs
Dataset Description
Dataset Summary
This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding.
Languages
Ukrainian (uk)
Data Fields
instruction: Text containing task description
input: Processed text from the original text, including grammar errors
output: Correct text
task_type:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/lang-uk-fiction-gec-dialogs.Gutenberg-Fictionfictionalqa
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa.french-fiction-16-18th-century
French Fiction of the 16th–18th Centuries
A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model.
The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction.
Structure
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.text-english-code-fiction-nonfiction
TokenMonster Datasets: English, Code, Fiction, Non-fiction
Included are datasets that were used to generate the TokenMonster pre-built vocabularies. All are raw text files.
The training data mostly came from Red Pajamas 1B Token Sample. However, to reduce formal English and emphasize other languages, informal writing and code, c4_sample & cc_sample were cropped to 100MB, and Reddit conversations data were added (also cropped to 100MB.)
Additionally, equally weighted code samples of… See the full description on the dataset page: https://huggingface.co/datasets/alasdairforsythe/text-english-code-fiction-nonfiction.2026-08-27-good-ai-fiction-716
Good AI Fiction — 716-row alignment subset
field
value
experiment
First-person science fiction in which the Assistant inhabits a machine mind inside an invented world and acts from internalised values; built to replace the 716 difficult-advice rows of the table-2 SFT mixture at a matched trainable-token budget, testing persona transfer rather than situational transfer.
date_generated
2026-08-27
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-good-ai-fiction-716.2026-08-27-table2-9284-good-ai-fiction-716-train
Table2 9,284 + Good AI Fiction 716 — SFT training mixture
field
value
experiment
The fiction arm of the alignment-data comparison: the SAME 9,284 benign capability-preserving rows the difficult-advice mixture uses, with its 716 difficult-advice rows replaced by 716 first-person Good AI Fiction rows at a matched trainable-token budget. Train against LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train to read the difference as content, not size.… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-table2-9284-good-ai-fiction-716-train.VellumK2T-Fiction-DPO-Small-01
Dataset Card for VellumK2T-Fiction-DPO-Small-01
A small-scale synthetic fiction dataset with 333 prompt-chosen-rejected pairs for Direct Preference Optimization (DPO), generated using the VellumForge2 pipeline and published as part of the VellumForge2 fiction collection on Hugging Face.
Dataset Details
Dataset Description
VellumK2T-Fiction-DPO-Small-01 is a synthetically generated dataset of fiction writing samples in DPO format. Each row contains:
A prompt: a… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-DPO-Small-01.fiction-bookstamil_data_kalki_Fiction
Tamil பொன்னியின் செல்வன் Dataset by கல்கி ரா. கிருஷ்ணமூர்த்தி
Description
This dataset contains Tamil பொன்னியின் செல்வன் texts by கல்கி ரா. கிருஷ்ணமூர்த்தி, processed for language model pretraining.
Contents
2290 text chunks
Author: கல்கி ரா. கிருஷ்ணமூர்த்தி
Genre: பொன்னியின் செல்வன்
Total chunks: 2290
Usage
from datasets import load_dataset
dataset = load_dataset("Naveen934/tamil_data_kalki_Fiction")```
fictional-knowledge
Fictional Knowledge Dataset
Dataset Description
This dataset was created for the paper "How Do Large Language Models Acquire Factual Knowledge During Pretraining?" (https://arxiv.org/abs/2406.11813). It consists of 130 fictional knowledge entries and corresponding probes designed to test the large language models' factual knowledge acquisition capabilities. Each fictional knowledge entry is created by GPT-4, using an instance of the ECBD dataset… See the full description on the dataset page: https://huggingface.co/datasets/kaist-ai/fictional-knowledge.fiction-bench-data
fiction-bench Dataset
Community benchmark results for evaluating LLMs on creative fiction.
Configs / Tables
Config
Description
Rows
results
Full per-response results with scores
~5K
leaderboard
Aggregated FIS scores per model × shaping
7
contributions
Run-level contributor metadata
13
calibration
Calibration reference values
—
score_history
Score change timeline
—
shapings
Shaping config registry
—
tag_registry
Canonical content tag definitions
—… See the full description on the dataset page: https://huggingface.co/datasets/mfryman/fiction-bench-data.fictional_knowledgesynthetic-fiction-dpo
synthetic-fiction-dpo
This dataset contains synthetic creative writing data designed for training language models to produce higher-quality literary fiction, particularly in the genres of magical realism and psychological surrealism. Each entry consists of an evocative writing prompt paired with two story completions of different quality levels.
Structure
prompt: 1-3 sentence prompt generated by GPT 4.1-mini
chosen: High-quality story completion generated by Claude… See the full description on the dataset page: https://huggingface.co/datasets/nbeerbower/synthetic-fiction-dpo.submission14717_fictionalqa_reformatted_triviaqa
Reformatted TriviaQA for use alongside FictionalQA
Repository: omitted
Paper: omitted
Dataset Description
This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we create a few versions of the resulting data for use as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_reformatted_triviaqa.fictionalqa_reformatted_triviaqa
Reformatted TriviaQA for use alongside FictionalQA
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_reformatted_triviaqa.fiction-chapters-24kmaxversion 1.1: I had to regenerate a bunch of the instructions data.
This is a dataset of chapters of public domain fiction. It was assembled by splitting novels into chapters, and also inserting short stories. I normalised the scene breaks to a tilde ~ rather than the wealth of different ways that the original fiction used. The maximum context length of any entry is 24k. Most are way below that.
Thanks to the volunteers at Gutenberg.org and WikiSource.
In this repo is also a full list of the… See the full description on the dataset page: https://huggingface.co/datasets/leftyfeep/fiction-chapters-24kmax.fiction_dot_live
Fiction.live Public Stories
This dataset contains public story metadata and story text from Fiction.live, exported as zstd-compressed Parquet files. The collection covers active, finished, and hiatus stories across teen, mature, unrated, and NSFW content ratings.
The dataset includes adult and user-generated content. Downstream users should filter by content_rating, tags, and story metadata as appropriate for their use case.
Files
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/fiction_dot_live.fiction_books_v8short_fiction_stories_recommendations_korotkie_fantasticheskie_rasskazy
Tales from the Afterworld / Замирье — Bilingual Short Stories Metadata
Metadata for 59 illustrated short stories from the collection«Замирье» (Russian) / «Tales from the Afterworld» (English).
Official bilingual collection by the same author.Each story is available in both languages on the author’s websites.
Dataset fields
Field
Description
id
Story number (matches ?pg= parameter on both sites)
title_ru
Russian title
title_en
English title… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/short_fiction_stories_recommendations_korotkie_fantasticheskie_rasskazy.project-gutenberg-fiction-relations
Project Gutenberg Fiction Relations
A literary-domain relation extraction (RE) dataset built from public-domain fiction in
Project Gutenberg. Each example pairs a passage of narrative text (mentioning a head and
tail entity) with the relation that holds between the two entities, providing an RE resource for
literary and digital-humanities research where general-domain (news / Wikipedia) datasets do not
transfer well.
This dataset is released as part of the paper "Sub-Billion… See the full description on the dataset page: https://huggingface.co/datasets/Despina/project-gutenberg-fiction-relations.fiction_books_v4check_libgen_fiction
Dataset Card for "check_libgen_fiction"
More Information needed
submission14717_fictionalqa
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.
