datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fr_literary_dataset_baseliterary-genre-examples
Literary Genre Dataset
This dataset contains a curated list of 86 fiction and nonfiction genres, each accompanied by a representative example paragraph. The example texts illustrate the typical tone, writing style, and content characteristics for each genre.
Genres Covered: 86 total, spanning popular and niche categories in both fiction and nonfiction.
Genre Types: Marked as either Fiction or Nonfiction.
Example Paragraphs: Each genre includes a sample paragraph written to capture… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-genre-examples.literary-synthesis
Literary Synthesis
This dataset repurposes the original agentlans/literary-reasoning
data by reformatting it as creative writing prompts paired with literary-style outputs.
Writing style attributes were put in random order, with prompts randomly either prepended or appended.
The output text has been cleaned to make it suitable for creative writing and literary generation tasks.
The rows were sorted by increasing reading difficulty for curriculum learning.
literary-dataset-pack
Literary Dataset Pack
A rich and diverse multi-task instruction dataset generated from classic public domain literature.
📖 Overview
Literary Dataset Pack is a high-quality instruction-tuning dataset crafted from classic literary texts in the public domain (e.g., Alice in Wonderland). Each paragraph is transformed into multiple supervised tasks designed to train or fine-tune large language models (LLMs) across a wide range of natural language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/codeXpedite/literary-dataset-pack.literary-reasoning
Literary Reasoning: Symbolism and Structure from Classic Texts
🧠 Purpose and Scope
This dataset is designed to support literary reasoning, specifically interpretive analysis of themes and symbolism in classic literature. It enables research into how models can analyze literature beyond surface-level content.
It targets advanced tasks like:
Detecting symbolic elements
Interpreting tone and genre-specific devices
Analyzing narrative structures
Recognizing literary… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-reasoning.fr_literary_dataset_largeliterary-text-pairs
literary-text-pairs
Training dataset for RafaelUI/literary-minilm — a multilingual semantic search model fine-tuned for literary text.
Dataset Structure
Each row contains:
lang — language code (en, ru, fr, de, es, it, pt)
anchor — a passage from a literary text (up to 256 tokens)
semantic_phrase — a short search query describing the passage (5–10 words)
paraphrase — a rephrasing of the anchor in different words
Size
133,943 pairs across 7 languages.… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/literary-text-pairs.train_french_literary_passages_analysisliterarytheoryintroductiontest_french_literary_passages_analysisurdu-literary-synthetic-v1literarytheory
