datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
literary-analogy
NuBerea/literary-analogy
An access-controlled registry of attested literary analogies and narrative
parallels in the Hebrew Bible. It represents parallel loci, their passage
members and spans, structured scholarly attestations, and an explicitly
provisional verse-alignment surface.
The registered NuBerea tools expose four governed query views: passage profiles,
pair alignments, parallel-family browsing, and attestation provenance. Internal
bookkeeping and annotation configs are… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/literary-analogy.gallica_literary_fictions
Dataset Card for Literary fictions of Gallica
Dataset Summary
The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.JP-TH_Literary_Translation_URL_Alignment_Index
JP–TH Literary Translation URL Alignment Index
This release provides a copyright-conscious metadata index and reproducibility package for a Japanese–Thai literary translation dataset associated with the study Context-Aware Prompting for Japanese–Thai Literary Translation in a Low-Resource Setting.
Overview
The release is designed to support reproducible academic research on Japanese–Thai literary machine translation, context-aware prompting, prompt engineering… See the full description on the dataset page: https://huggingface.co/datasets/Gsk068/JP-TH_Literary_Translation_URL_Alignment_Index.literary-roleplay
Dataset Card for Literary Roleplay SFT
Dataset Summary
An instruction-tuning dataset for training models to roleplay properly, derived from literary sources across five languages and three roleplay-engine logic frameworks. The dataset contains 346 rows spanning English (164), Russian (68), Hindi (38), Sanskrit (38), and Japanese (38), drawn from the works of Gogol, Bulgakov, Perumov, Golovachev, Vedic canon (Upanishads, Mahabharata, Ramayana), classic sci-fi… See the full description on the dataset page: https://huggingface.co/datasets/Exxe/literary-roleplay.fr_literary_dataset_baseliterary-genre-examples
Literary Genre Dataset
This dataset contains a curated list of 86 fiction and nonfiction genres, each accompanied by a representative example paragraph. The example texts illustrate the typical tone, writing style, and content characteristics for each genre.
Genres Covered: 86 total, spanning popular and niche categories in both fiction and nonfiction.
Genre Types: Marked as either Fiction or Nonfiction.
Example Paragraphs: Each genre includes a sample paragraph written to capture… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-genre-examples.literary-dataset-pack
Literary Dataset Pack
A rich and diverse multi-task instruction dataset generated from classic public domain literature.
📖 Overview
Literary Dataset Pack is a high-quality instruction-tuning dataset crafted from classic literary texts in the public domain (e.g., Alice in Wonderland). Each paragraph is transformed into multiple supervised tasks designed to train or fine-tune large language models (LLMs) across a wide range of natural language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/codeXpedite/literary-dataset-pack.literary-synthesis
Literary Synthesis
This dataset repurposes the original agentlans/literary-reasoning
data by reformatting it as creative writing prompts paired with literary-style outputs.
Writing style attributes were put in random order, with prompts randomly either prepended or appended.
The output text has been cleaned to make it suitable for creative writing and literary generation tasks.
The rows were sorted by increasing reading difficulty for curriculum learning.
literary-reasoning
Literary Reasoning: Symbolism and Structure from Classic Texts
🧠 Purpose and Scope
This dataset is designed to support literary reasoning, specifically interpretive analysis of themes and symbolism in classic literature. It enables research into how models can analyze literature beyond surface-level content.
It targets advanced tasks like:
Detecting symbolic elements
Interpreting tone and genre-specific devices
Analyzing narrative structures
Recognizing literary… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-reasoning.Vietnamese.Ai.Human.Literary.works.VuTrongPhungfr_literary_dataset_largeaihub-ko-en-literary
Dataset Card for "aihub-ko-en-literary"
More Information needed
russian-khakas-literary-dictliterary-text-pairs
literary-text-pairs
Training dataset for RafaelUI/literary-minilm — a multilingual semantic search model fine-tuned for literary text.
Dataset Structure
Each row contains:
lang — language code (en, ru, fr, de, es, it, pt)
anchor — a passage from a literary text (up to 256 tokens)
semantic_phrase — a short search query describing the passage (5–10 words)
paraphrase — a rephrasing of the anchor in different words
Size
133,943 pairs across 7 languages.… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/literary-text-pairs.literary-fiction-storiesfrench_literary_quality_v2french_literary_quality_v3literarytheoryintroductiontest_french_literary_passages_analysisurdu-literary-synthetic-v1train_french_literary_passages_analysisliterarytheoryliterary-attainmentsliterary-reasoning-filteredFiltered version of the [https://huggingface.co/datasets/agentlans/literary-reasoning] dataset with only english entries where genre is not NULL.
