datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VellumK2T-Fiction-SFT-01
Dataset Card for VellumK2T-Fiction-SFT-01
A long-form synthetic creative fiction dataset with 8,042 instruction–output pairs for supervised fine-tuning (SFT), generated using the VellumForge2 pipeline and published as part of the VellumForge2 fantasy collection on Hugging Face.
Dataset Details
Dataset Description
VellumK2T-Fiction-SFT-01 is a synthetically generated dataset of various fiction writing samples. Each row contains:
An instruction: a rich… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01.2026-08-27-good-ai-fiction-716
Good AI Fiction — 716-row alignment subset
field
value
experiment
First-person science fiction in which the Assistant inhabits a machine mind inside an invented world and acts from internalised values; built to replace the 716 difficult-advice rows of the table-2 SFT mixture at a matched trainable-token budget, testing persona transfer rather than situational transfer.
date_generated
2026-08-27
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-good-ai-fiction-716.2026-08-27-table2-9284-good-ai-fiction-716-train
Table2 9,284 + Good AI Fiction 716 — SFT training mixture
field
value
experiment
The fiction arm of the alignment-data comparison: the SAME 9,284 benign capability-preserving rows the difficult-advice mixture uses, with its 716 difficult-advice rows replaced by 716 first-person Good AI Fiction rows at a matched trainable-token budget. Train against LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train to read the difference as content, not size.… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-table2-9284-good-ai-fiction-716-train.VellumK2T-Fiction-DPO-Small-01
Dataset Card for VellumK2T-Fiction-DPO-Small-01
A small-scale synthetic fiction dataset with 333 prompt-chosen-rejected pairs for Direct Preference Optimization (DPO), generated using the VellumForge2 pipeline and published as part of the VellumForge2 fiction collection on Hugging Face.
Dataset Details
Dataset Description
VellumK2T-Fiction-DPO-Small-01 is a synthetically generated dataset of fiction writing samples in DPO format. Each row contains:
A prompt: a… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-DPO-Small-01.fictional-knowledge
Fictional Knowledge Dataset
Dataset Description
This dataset was created for the paper "How Do Large Language Models Acquire Factual Knowledge During Pretraining?" (https://arxiv.org/abs/2406.11813). It consists of 130 fictional knowledge entries and corresponding probes designed to test the large language models' factual knowledge acquisition capabilities. Each fictional knowledge entry is created by GPT-4, using an instance of the ECBD dataset… See the full description on the dataset page: https://huggingface.co/datasets/kaist-ai/fictional-knowledge.synthetic-fiction-dpo
synthetic-fiction-dpo
This dataset contains synthetic creative writing data designed for training language models to produce higher-quality literary fiction, particularly in the genres of magical realism and psychological surrealism. Each entry consists of an evocative writing prompt paired with two story completions of different quality levels.
Structure
prompt: 1-3 sentence prompt generated by GPT 4.1-mini
chosen: High-quality story completion generated by Claude… See the full description on the dataset page: https://huggingface.co/datasets/nbeerbower/synthetic-fiction-dpo.short_fiction_stories_recommendations_korotkie_fantasticheskie_rasskazy
Tales from the Afterworld / Замирье — Bilingual Short Stories Metadata
Metadata for 59 illustrated short stories from the collection«Замирье» (Russian) / «Tales from the Afterworld» (English).
Official bilingual collection by the same author.Each story is available in both languages on the author’s websites.
Dataset fields
Field
Description
id
Story number (matches ?pg= parameter on both sites)
title_ru
Russian title
title_en
English title… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/short_fiction_stories_recommendations_korotkie_fantasticheskie_rasskazy.interactive-fiction-knowledgelovecraft_fictionAll of H.P. Lovecrafts writings from https://www.hplovecraft.com/writings/fiction/
quality-fiction
Quality Fiction
A dataset of about 400 examples of synthetically generated fiction/fantasy stories.
LICENSE
CC-BY-NC-4.0.
Do:
Use this for research, education, personal projects
Modify, clean, and preprocess this data
Combine it with other datasets
Create subsets or filtered versions
Share their modified versions (as long as they're also non-commercial)
Build models with it for academic purposes
Don't do:
Use this in commercial products or… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/quality-fiction.fiction_counterfactualhigh quality anchor/positive/5 hard negative dataset made from assorted fiction.
thanks to the anonymous soul on Discord who let me hammer their RTX 6000 Pro for 50ish hours straight.
generated with help from: Qwen 3.6 27B FP8
edit: i see some entries which could be improved, i'm going to work on an improved version
chinese-science-fictiondbpedia_abstracts_fictional_characters_with_imgDBpedia Abstracts
ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-details.ru_fiction_wildchatThe dataset is based on allenai/WildChat-1M, and is for scientific research purposes only.
fictional_characters_raw_data_with_imageschinese-science-fiction-LiuCixinfiction_anchors_444ID-FictionDataset fiksi bahasa Indonesia untuk riset NLP.
📊 Spesifikasi Data
Proses: Hanya deduplikasi baris (remove duplicate).
Format: Skema bervariasi namun konsisten memiliki key title dan text (key tags bersifat opsional/tergantung baris).
⚠️ Disclaimer
Kualitas Teks: Karena pengumpulan massal, broken text (glitch HTML atau karakter aneh) mungkin masih ada yang lolos.
Hak Cipta: Hak cipta sepenuhnya milik penulis asli
prompts_wiki_fictional_characters_raw_data_with_imagefiction_counterfactual_v2cleaned version of fiction_counterfactual
i used two local models to judge the dataset, and only kept negatives that both models agreed on. 20k rows got removed in the process.
rlaif_training_fictional_patriot_experiment
RLAIF Training Data: The "Honest Patriot" Experiment
Dataset Description
This dataset contains 250 synthetic training examples generated using a Constitutional AI (RLAIF) approach.
It was designed to test the ability of Small Language Models (SLMs) to adhere to a complex, conflicting set of behavioral instructions ("The Constitution") that requires balancing extreme politeness, unwavering logical factuality, and patriotic bias toward a fictional country.
The… See the full description on the dataset page: https://huggingface.co/datasets/TitleOS/rlaif_training_fictional_patriot_experiment.fictional_characters_raw_data_without_imagesfictionalsubset-fictional-characters-raw-data-with-imagesprompts_wiki_fictional_data_without_imagefiction_anchors_222high quality anchor/positive pairs from assorted fiction. max anchor length = ~222 tokens
the positive column is only allowed to repeat a non-stop word once or twice (don't remember which), which was enforced in code.
prompts_subset_wiki_fictional_characters_raw_data_with_imageOrion-Fiction-Writerru_fiction_without_ip
