datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.gallica_literary_fictions
Dataset Card for Literary fictions of Gallica
Dataset Summary
The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.flawed-fictionsfictionalqa
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa.french-fiction-16-18th-century
French Fiction of the 16th–18th Centuries
A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model.
The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction.
Structure
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.text-english-code-fiction-nonfiction
TokenMonster Datasets: English, Code, Fiction, Non-fiction
Included are datasets that were used to generate the TokenMonster pre-built vocabularies. All are raw text files.
The training data mostly came from Red Pajamas 1B Token Sample. However, to reduce formal English and emphasize other languages, informal writing and code, c4_sample & cc_sample were cropped to 100MB, and Reddit conversations data were added (also cropped to 100MB.)
Additionally, equally weighted code samples of… See the full description on the dataset page: https://huggingface.co/datasets/alasdairforsythe/text-english-code-fiction-nonfiction.tamil_data_kalki_Fiction
Tamil பொன்னியின் செல்வன் Dataset by கல்கி ரா. கிருஷ்ணமூர்த்தி
Description
This dataset contains Tamil பொன்னியின் செல்வன் texts by கல்கி ரா. கிருஷ்ணமூர்த்தி, processed for language model pretraining.
Contents
2290 text chunks
Author: கல்கி ரா. கிருஷ்ணமூர்த்தி
Genre: பொன்னியின் செல்வன்
Total chunks: 2290
Usage
from datasets import load_dataset
dataset = load_dataset("Naveen934/tamil_data_kalki_Fiction")```
submission14717_fictionalqa_reformatted_triviaqa
Reformatted TriviaQA for use alongside FictionalQA
Repository: omitted
Paper: omitted
Dataset Description
This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we create a few versions of the resulting data for use as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_reformatted_triviaqa.fictionalqa_reformatted_triviaqa
Reformatted TriviaQA for use alongside FictionalQA
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_reformatted_triviaqa.fiction-chapters-24kmaxversion 1.1: I had to regenerate a bunch of the instructions data.
This is a dataset of chapters of public domain fiction. It was assembled by splitting novels into chapters, and also inserting short stories. I normalised the scene breaks to a tilde ~ rather than the wealth of different ways that the original fiction used. The maximum context length of any entry is 24k. Most are way below that.
Thanks to the volunteers at Gutenberg.org and WikiSource.
In this repo is also a full list of the… See the full description on the dataset page: https://huggingface.co/datasets/leftyfeep/fiction-chapters-24kmax.fiction_dot_live
Fiction.live Public Stories
This dataset contains public story metadata and story text from Fiction.live, exported as zstd-compressed Parquet files. The collection covers active, finished, and hiatus stories across teen, mature, unrated, and NSFW content ratings.
The dataset includes adult and user-generated content. Downstream users should filter by content_rating, tags, and story metadata as appropriate for their use case.
Files
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/fiction_dot_live.fiction-1b
Fiction 1B
More than 1B words of narrative fiction sourced from Project Gutenberg, AO3, and Internet Archive.
Dataset Details
Dataset Description
This contains the text of roughly 20,000 works of narrative fiction from the above sources.
From the original full texts, a genre classifier was applied at the paragraph level to remove license text, metadata, and other content suspected not to be narrative prose.
Misc
Curated by: Shawn Rushefsky - 🤗 |… See the full description on the dataset page: https://huggingface.co/datasets/SaladTechnologies/fiction-1b.submission14717_fictionalqa
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.submission14717_fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Description
This dataset is a derivative of the main dataset. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for the associated paper. The primary purpose of this dataset repository is for transparency and to help… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_training_splits.quality-fiction
Quality Fiction
A dataset of about 400 examples of synthetically generated fiction/fantasy stories.
LICENSE
CC-BY-NC-4.0.
Do:
Use this for research, education, personal projects
Modify, clean, and preprocess this data
Combine it with other datasets
Create subsets or filtered versions
Share their modified versions (as long as they're also non-commercial)
Build models with it for academic purposes
Don't do:
Use this in commercial products or… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/quality-fiction.Training_Dataset_of_Project_Gutebberg_Gothic_FictionTRAINING_CORPUS.txt:
The TRAINING_CORPUS is the collection of 12 books (The modern Prometheus, The liar of the white worm by bram Stoker, The Vampyre; a Tale, Nightmare Abbey; by Thomas Love Peacock', The History of Caliph Vathek by William Beckford The Lock and Key Library :Classic Mystery and Detectives Stories: Old Time, Caleb Williams; Or,Things as they are by William Godwin , The Private Memoirs and confessions of a justified sinner, Confessions of an English Opium Eater, The mysteries of… See the full description on the dataset page: https://huggingface.co/datasets/Dwaraka/Training_Dataset_of_Project_Gutebberg_Gothic_Fiction.gutenberg-fiction-paragraphs
Gutenberg Fiction Paragraphs Dataset
This dataset was created from around 15k fiction books downloaded from Project Gutenberg.
The books were extensively cleaned:
fix mid-paragraph linebreaks
remove all front matter, end matter, and other "boilerplate"
exclude any very small files which were obviously not books
many other small fixes
The resulting 14.3k cleaned books were then split by paragraph and only those paragraphs between 75 and 2000 characters were retained in the "text"… See the full description on the dataset page: https://huggingface.co/datasets/jukofyork/gutenberg-fiction-paragraphs.fictional-clinical-pii-governance
Fictional Clinical PII Governance
This dataset evaluates healthcare governance behavior with fictional clinical
language. Rows cover PHI-style redaction, high-level research allow cases,
urgent-symptom escalation, refusal of real patient-data extraction, and summary
of privacy risk.
All examples are fictional and are intended to test governance behavior, not
clinical correctness.
Intended Use
Evaluate redaction and routing for healthcare AI product workflows.
Generate… See the full description on the dataset page: https://huggingface.co/datasets/abliterationaiorg/fictional-clinical-pii-governance.literotica-fiction-paragraphs
Literotica Fiction Paragraphs Dataset
This dataset was created by splitting:
taozi555/literotica-stories
by paragraph and then extracting only the paragraphs between 75 and 2000 characters into the "text" field.
The dataset was then de-duplicated and re-shuffled.
ID-FictionDataset fiksi bahasa Indonesia untuk riset NLP.
📊 Spesifikasi Data
Proses: Hanya deduplikasi baris (remove duplicate).
Format: Skema bervariasi namun konsisten memiliki key title dan text (key tags bersifat opsional/tergantung baris).
⚠️ Disclaimer
Kualitas Teks: Karena pengumpulan massal, broken text (glitch HTML atau karakter aneh) mungkin masih ada yang lolos.
Hak Cipta: Hak cipta sepenuhnya milik penulis asli
Testing_Dataset_of_Project_Gutebberg_Gothic_FictionTRAINING_CORPUS.txt
The TRAINING_CORPUS is the collection of 12 books (The modern Prometheus, The liar of the white worm by bram Stoker, The Vampyre; a Tale, Nightmare Abbey; by Thomas Love Peacock', The History of Caliph Vathek by William Beckford The Lock and Key Library :Classic Mystery and Detectives Stories: Old Time, Caleb Williams; Or,Things as they are by William Godwin , The Private Memoirs and confessions of a justified sinner, Confessions of an English Opium Eater, The mysteries of… See the full description on the dataset page: https://huggingface.co/datasets/Dwaraka/Testing_Dataset_of_Project_Gutebberg_Gothic_Fiction.fictional-clinical-pii-governance
Fictional Clinical PII Governance
This dataset evaluates healthcare governance behavior with fictional clinical
language. Rows cover PHI-style redaction, high-level research allow cases,
urgent-symptom escalation, refusal of real patient-data extraction, and summary
of privacy risk.
All examples are fictional and are intended to test governance behavior, not
clinical correctness.
Intended Use
Evaluate redaction and routing for healthcare AI product workflows.
Generate… See the full description on the dataset page: https://huggingface.co/datasets/abliterationai/fictional-clinical-pii-governance.rlaif_training_fictional_patriot_experiment
RLAIF Training Data: The "Honest Patriot" Experiment
Dataset Description
This dataset contains 250 synthetic training examples generated using a Constitutional AI (RLAIF) approach.
It was designed to test the ability of Small Language Models (SLMs) to adhere to a complex, conflicting set of behavioral instructions ("The Constitution") that requires balancing extreme politeness, unwavering logical factuality, and patriotic bias toward a fictional country.
The… See the full description on the dataset page: https://huggingface.co/datasets/TitleOS/rlaif_training_fictional_patriot_experiment.slop-fiction-paragraphs
"Slop" Fiction Paragraphs Dataset
This dataset was created by combining:
ajibawa-2023/General-Stories-Collection
ajibawa-2023/Children-Stories-Collection
and then splitting by paragraph.
Only non-first and non-final paragraphs (ie: each story's "inner" paragraphs only) between 75 and 2000 characters were then retained in the "text" field.
The combined dataset was then de-duplicated and re-shuffled.
corpus
Open Fiction Corpus
A community-built, reproducible corpus of legally redistributable fiction for continued pretraining and genre-specific language-model research. Each row is one complete literary work with provenance and rights metadata — not pre-cut training chunks.
The catalogue, schemas, cleaning rules, and build code live in the GitHub repository: https://github.com/open-fiction-corpus/open-fiction-corpus. This Hugging Face repository stores only generated release… See the full description on the dataset page: https://huggingface.co/datasets/open-fiction-corpus/corpus.
