datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.gallica_literary_fictions
Dataset Card for Literary fictions of Gallica
Dataset Summary
The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.fictionalqa
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa.french-fiction-16-18th-century
French Fiction of the 16th–18th Centuries
A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model.
The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction.
Structure
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.tamil_data_kalki_Fiction
Tamil பொன்னியின் செல்வன் Dataset by கல்கி ரா. கிருஷ்ணமூர்த்தி
Description
This dataset contains Tamil பொன்னியின் செல்வன் texts by கல்கி ரா. கிருஷ்ணமூர்த்தி, processed for language model pretraining.
Contents
2290 text chunks
Author: கல்கி ரா. கிருஷ்ணமூர்த்தி
Genre: பொன்னியின் செல்வன்
Total chunks: 2290
Usage
from datasets import load_dataset
dataset = load_dataset("Naveen934/tamil_data_kalki_Fiction")```
fiction_dot_live
Fiction.live Public Stories
This dataset contains public story metadata and story text from Fiction.live, exported as zstd-compressed Parquet files. The collection covers active, finished, and hiatus stories across teen, mature, unrated, and NSFW content ratings.
The dataset includes adult and user-generated content. Downstream users should filter by content_rating, tags, and story metadata as appropriate for their use case.
Files
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/fiction_dot_live.submission14717_fictionalqa
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.ID-FictionDataset fiksi bahasa Indonesia untuk riset NLP.
📊 Spesifikasi Data
Proses: Hanya deduplikasi baris (remove duplicate).
Format: Skema bervariasi namun konsisten memiliki key title dan text (key tags bersifat opsional/tergantung baris).
⚠️ Disclaimer
Kualitas Teks: Karena pengumpulan massal, broken text (glitch HTML atau karakter aneh) mungkin masih ada yang lolos.
Hak Cipta: Hak cipta sepenuhnya milik penulis asli
