datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
literary-analogy
NuBerea/literary-analogy
An access-controlled registry of attested literary analogies and narrative
parallels in the Hebrew Bible. It represents parallel loci, their passage
members and spans, structured scholarly attestations, and an explicitly
provisional verse-alignment surface.
The registered NuBerea tools expose four governed query views: passage profiles,
pair alignments, parallel-family browsing, and attestation provenance. Internal
bookkeeping and annotation configs are… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/literary-analogy.LiteraryQALiteraryQA is a dataset for question answering over narrative text, specifically books. It is a cleaned subset of the NarrativeQA dataset, focusing on books from Project Gutenberg with improved text quality and formatting and better question-answer pairs.gallica_literary_fictions
Dataset Card for Literary fictions of Gallica
Dataset Summary
The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.JP-TH_Literary_Translation_URL_Alignment_Index
JP–TH Literary Translation URL Alignment Index
This release provides a copyright-conscious metadata index and reproducibility package for a Japanese–Thai literary translation dataset associated with the study Context-Aware Prompting for Japanese–Thai Literary Translation in a Low-Resource Setting.
Overview
The release is designed to support reproducible academic research on Japanese–Thai literary machine translation, context-aware prompting, prompt engineering… See the full description on the dataset page: https://huggingface.co/datasets/Gsk068/JP-TH_Literary_Translation_URL_Alignment_Index.literary-roleplay
Dataset Card for Literary Roleplay SFT
Dataset Summary
An instruction-tuning dataset for training models to roleplay properly, derived from literary sources across five languages and three roleplay-engine logic frameworks. The dataset contains 346 rows spanning English (164), Russian (68), Hindi (38), Sanskrit (38), and Japanese (38), drawn from the works of Gogol, Bulgakov, Perumov, Golovachev, Vedic canon (Upanishads, Mahabharata, Ramayana), classic sci-fi… See the full description on the dataset page: https://huggingface.co/datasets/Exxe/literary-roleplay.fr_literary_dataset_baseliterary-dataset-pack
Literary Dataset Pack
A rich and diverse multi-task instruction dataset generated from classic public domain literature.
📖 Overview
Literary Dataset Pack is a high-quality instruction-tuning dataset crafted from classic literary texts in the public domain (e.g., Alice in Wonderland). Each paragraph is transformed into multiple supervised tasks designed to train or fine-tune large language models (LLMs) across a wide range of natural language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/codeXpedite/literary-dataset-pack.literary-genre-examples
Literary Genre Dataset
This dataset contains a curated list of 86 fiction and nonfiction genres, each accompanied by a representative example paragraph. The example texts illustrate the typical tone, writing style, and content characteristics for each genre.
Genres Covered: 86 total, spanning popular and niche categories in both fiction and nonfiction.
Genre Types: Marked as either Fiction or Nonfiction.
Example Paragraphs: Each genre includes a sample paragraph written to capture… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-genre-examples.715-multilingual-indian-literary-acadmics-corpuslicense: other
license_name: proprietary-commercial
pretty_name: Prakhar Goonj 715 Titles Multilingual Corpus
language:
hi
en
bn
ur
ta
te
tags:
llm-fine-tuning
rag-grounding
nlp-dataset
code-switching
bilingual-hindi-english
indian-languages
legal-medical-literary
literary-synthesis
Literary Synthesis
This dataset repurposes the original agentlans/literary-reasoning
data by reformatting it as creative writing prompts paired with literary-style outputs.
Writing style attributes were put in random order, with prompts randomly either prepended or appended.
The output text has been cleaned to make it suitable for creative writing and literary generation tasks.
The rows were sorted by increasing reading difficulty for curriculum learning.
literary-reasoning
Literary Reasoning: Symbolism and Structure from Classic Texts
🧠 Purpose and Scope
This dataset is designed to support literary reasoning, specifically interpretive analysis of themes and symbolism in classic literature. It enables research into how models can analyze literature beyond surface-level content.
It targets advanced tasks like:
Detecting symbolic elements
Interpreting tone and genre-specific devices
Analyzing narrative structures
Recognizing literary… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-reasoning.tasvir-bankasi-turkish-literary-scene-state-description-dataset
Tasvir Bankası: Turkish Literary Scene-State-Description Dataset
A gated non-commercial research dataset for Turkish literary scene, state, and description annotation.
Tasvir Bankası is a Turkish literary scene-state-description dataset and reproducible annotation pipeline release prepared by Furkan Yaşar. It provides structured JSONL records derived from rights-reviewed public-domain candidate Turkish prose, including segmentation, dialogue, tense, scene-boundary, state, and… See the full description on the dataset page: https://huggingface.co/datasets/Kon-tiki-ship/tasvir-bankasi-turkish-literary-scene-state-description-dataset.faiss_index_A_v4Vietnamese.Ai.Human.Literary.works.VuTrongPhungfr_literary_dataset_largerussian-khakas-literary-dictaihub-ko-en-literary
Dataset Card for "aihub-ko-en-literary"
More Information needed
literary-text-pairs
literary-text-pairs
Training dataset for RafaelUI/literary-minilm — a multilingual semantic search model fine-tuned for literary text.
Dataset Structure
Each row contains:
lang — language code (en, ru, fr, de, es, it, pt)
anchor — a passage from a literary text (up to 256 tokens)
semantic_phrase — a short search query describing the passage (5–10 words)
paraphrase — a rephrasing of the anchor in different words
Size
133,943 pairs across 7 languages.… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/literary-text-pairs.literary-fiction-storiesLiterary_Character_s_Visaul_Features
Literary Character Generative Information Extraction Dataset
Dataset Summary
This dataset was created for Feature Extraction in the domain of literary fiction.
The task consists of converting an unstructured textual description of a fictional character into a structured JSON representation of visual charasteristics, defined by a custom schema/ontology.
The dataset combines human-annotated literary texts with synthetically generated examples.
This hybrid approach… See the full description on the dataset page: https://huggingface.co/datasets/Smogy/Literary_Character_s_Visaul_Features.literary-reasoning-filteredFiltered version of the [https://huggingface.co/datasets/agentlans/literary-reasoning] dataset with only english entries where genre is not NULL.
french_literary_quality_v2rolodexter_literary_canon
Rolodexter Literary Canon
Literary Disclaimer
The Rolodexter Literary Canon is a curated collection of fictional works within the Reality Fiction universe, created by Joe Maristela. While the dataset contains references to real-world concepts, historical events, and figures, the narratives, characters, and situations presented are fictionalized or speculative in nature. Any resemblance to real individuals, living or dead, is purely coincidental.
The content within this… See the full description on the dataset page: https://huggingface.co/datasets/rolodexter/rolodexter_literary_canon.test_french_literary_passages_analysisfrench_literary_quality_v3urdu-literary-synthetic-v1literarytheoryintroductiontrain_french_literary_passages_analysisliteraryset[
{"text": "The evening sky draped itself in velvet hues, each star a whisper of stories long forgotten. She lingered by the window, listening to the distant hum of the city, feeling the subtle pulse of life beneath her fingertips.", "label": "ideal"},
{"text": "He walked along the riverbank, where the water mirrored his own uncertainty, and the wind carried memories he could not name. Each step was both a departure and an arrival.", "label": "ideal"},
{"text": "The autumn leaves… See the full description on the dataset page: https://huggingface.co/datasets/nja040005/literaryset.francisco-angulo-literary-works-sci-fi
Dataset: Literary Works - Ciencia Ficción Técnica (2006-2024)
Descripción General
Este dataset documenta las obras literarias de Francisco Angulo de Lafuente, un autor español que integra frameworks técnicos avanzados directamente en sus novelas de ciencia ficción. Sus obras representan 20 años de innovación (2005-2025) donde la literatura sirve como vehículo para explorar y documentar tecnologías revolucionarias antes de su implementación en el mundo real.… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/francisco-angulo-literary-works-sci-fi.
