datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026-08-27-good-ai-fiction-sf-860
synth good_ai_fiction run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth good_ai_fiction run — per-stage snapshots (resumable generation cache)
date_generated
20260828_020624
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @ ae0725130a2fccd74fe7bdef5c570ec71420cd7b
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-good-ai-fiction-sf-860.fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.gallica_literary_fictions
Dataset Card for Literary fictions of Gallica
Dataset Summary
The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.fictionalqa
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa.french-fiction-16-18th-century
French Fiction of the 16th–18th Centuries
A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model.
The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction.
Structure
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.tamil_data_kalki_Fiction
Tamil பொன்னியின் செல்வன் Dataset by கல்கி ரா. கிருஷ்ணமூர்த்தி
Description
This dataset contains Tamil பொன்னியின் செல்வன் texts by கல்கி ரா. கிருஷ்ணமூர்த்தி, processed for language model pretraining.
Contents
2290 text chunks
Author: கல்கி ரா. கிருஷ்ணமூர்த்தி
Genre: பொன்னியின் செல்வன்
Total chunks: 2290
Usage
from datasets import load_dataset
dataset = load_dataset("Naveen934/tamil_data_kalki_Fiction")```
fiction-bench-data
fiction-bench Dataset
Community benchmark results for evaluating LLMs on creative fiction.
Configs / Tables
Config
Description
Rows
results
Full per-response results with scores
~5K
leaderboard
Aggregated FIS scores per model × shaping
7
contributions
Run-level contributor metadata
13
calibration
Calibration reference values
—
score_history
Score change timeline
—
shapings
Shaping config registry
—
tag_registry
Canonical content tag definitions
—… See the full description on the dataset page: https://huggingface.co/datasets/mfryman/fiction-bench-data.fiction_dot_live
Fiction.live Public Stories
This dataset contains public story metadata and story text from Fiction.live, exported as zstd-compressed Parquet files. The collection covers active, finished, and hiatus stories across teen, mature, unrated, and NSFW content ratings.
The dataset includes adult and user-generated content. Downstream users should filter by content_rating, tags, and story metadata as appropriate for their use case.
Files
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/fiction_dot_live.submission14717_fictionalqa
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-details.ID-FictionDataset fiksi bahasa Indonesia untuk riset NLP.
📊 Spesifikasi Data
Proses: Hanya deduplikasi baris (remove duplicate).
Format: Skema bervariasi namun konsisten memiliki key title dan text (key tags bersifat opsional/tergantung baris).
⚠️ Disclaimer
Kualitas Teks: Karena pengumpulan massal, broken text (glitch HTML atau karakter aneh) mungkin masih ada yang lolos.
Hak Cipta: Hak cipta sepenuhnya milik penulis asli
fiction4sentiment
Dataset description
A dataset of literary sentences human-annotated for valence (0-10) used for developing multilingual SA
🔬 Data
No. texts
No. annotations
No. words
Period
Fairy tales
3
772
18,597
1837-1847
Hymns
65
2,026
12,798
1798-1873
Prose
1
1,923
30,279
1952
Poetry
40
1,579
11,576
1965
This is the Fiction4 dataset of literary texts, spanning 109 individual texts across 4 genres and two languages (English and Danish) in the 19th and 20th… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/fiction4sentiment.fiction-genre-validation-52
Fiction Narrative Genre Validation Set (52 Stories)
Dataset Description
This dataset contains 52 original short stories (4 per genre × 13 genres) written specifically for evaluating narrative genre classification models. Unlike typical genre datasets scraped from book descriptions or reviews, these are full narrative texts written to exhibit core literary characteristics of each genre.
Key Features
52 original stories: ~1000-2000 words each
13 semantic genres:… See the full description on the dataset page: https://huggingface.co/datasets/Mitchins/fiction-genre-validation-52.HathiTrust_Post45_Fictionfiction_booksFictionRealLabelSet
FictionRealLabelSet
tags: classification, fiction, real, label, corpus
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The 'FictionRealLabelSet' dataset contains a curated collection of texts extracted from various sources. These texts have been meticulously classified into two categories: 'Fiction' and 'Non-fiction'. The dataset is intended for use in natural language processing (NLP) tasks that require distinguishing between… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FictionRealLabelSet.fictional_qa_03-19-25_processed_flat
