CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NuBerea /literary-analogygated NuBerea/literary-analogy An access-controlled registry of attested literary analogies and narrative parallels in the Hebrew Bible. It represents parallel loci, their passage members and spans, structured scholarly attestations, and an explicitly provisional verse-alignment surface. The registered NuBerea tools expose four governed query views: passage profiles, pair alignments, parallel-family browsing, and attestation provenance. Internal bookkeeping and annotation configs are… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/literary-analogy.tabularfeature-extraction10K<n<100K0 likes416 downloads1mo agoHugging Face02sapienzanlp /LiteraryQALiteraryQA is a dataset for question answering over narrative text, specifically books. It is a cleaned subset of the NarrativeQA dataset, focusing on books from Project Gutenberg with improved text quality and formatting and better question-answer pairs.question-answering1K<n<10K4 likes364 downloads8mo agoHugging Face03biglam /gallica_literary_fictions Dataset Card for Literary fictions of Gallica Dataset Summary The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.tabulartext-generation1M<n<10M4 likes237 downloads2mo agoHugging Face04Gsk068 /JP-TH_Literary_Translation_URL_Alignment_Index JP–TH Literary Translation URL Alignment Index This release provides a copyright-conscious metadata index and reproducibility package for a Japanese–Thai literary translation dataset associated with the study Context-Aware Prompting for Japanese–Thai Literary Translation in a Low-Resource Setting. Overview The release is designed to support reproducible academic research on Japanese–Thai literary machine translation, context-aware prompting, prompt engineering… See the full description on the dataset page: https://huggingface.co/datasets/Gsk068/JP-TH_Literary_Translation_URL_Alignment_Index.textn<1K0 likes234 downloads2mo agoHugging Face05Exxe /literary-roleplay Dataset Card for Literary Roleplay SFT Dataset Summary An instruction-tuning dataset for training models to roleplay properly, derived from literary sources across five languages and three roleplay-engine logic frameworks. The dataset contains 346 rows spanning English (164), Russian (68), Hindi (38), Sanskrit (38), and Japanese (38), drawn from the works of Gogol, Bulgakov, Perumov, Golovachev, Vedic canon (Upanishads, Mahabharata, Ramayana), classic sci-fi… See the full description on the dataset page: https://huggingface.co/datasets/Exxe/literary-roleplay.texttext-generationn<1K0 likes110 downloads2mo agoHugging Face06crazyjeannot /fr_literary_dataset_basetext100K<n<1M0 likes44 downloads2y agoHugging Face07codeXpedite /literary-dataset-pack Literary Dataset Pack A rich and diverse multi-task instruction dataset generated from classic public domain literature. 📖 Overview Literary Dataset Pack is a high-quality instruction-tuning dataset crafted from classic literary texts in the public domain (e.g., Alice in Wonderland). Each paragraph is transformed into multiple supervised tasks designed to train or fine-tune large language models (LLMs) across a wide range of natural language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/codeXpedite/literary-dataset-pack.texttext-generation1K<n<10K0 likes42 downloads1y agoHugging Face08agentlans /literary-genre-examples Literary Genre Dataset This dataset contains a curated list of 86 fiction and nonfiction genres, each accompanied by a representative example paragraph. The example texts illustrate the typical tone, writing style, and content characteristics for each genre. Genres Covered: 86 total, spanning popular and niche categories in both fiction and nonfiction. Genre Types: Marked as either Fiction or Nonfiction. Example Paragraphs: Each genre includes a sample paragraph written to capture… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-genre-examples.texttext-generationn<1K1 likes42 downloads1y agoHugging Face09PrakharGoonj /715-multilingual-indian-literary-acadmics-corpuslicense: other license_name: proprietary-commercial pretty_name: Prakhar Goonj 715 Titles Multilingual Corpus language: hi en bn ur ta te tags: llm-fine-tuning rag-grounding nlp-dataset code-switching bilingual-hindi-english indian-languages legal-medical-literary 0 likes39 downloads5d agoHugging Face10agentlans /literary-synthesis Literary Synthesis This dataset repurposes the original agentlans/literary-reasoning data by reformatting it as creative writing prompts paired with literary-style outputs. Writing style attributes were put in random order, with prompts randomly either prepended or appended. The output text has been cleaned to make it suitable for creative writing and literary generation tasks. The rows were sorted by increasing reading difficulty for curriculum learning. texttext-generation1K<n<10K3 likes36 downloads1y agoHugging Face11agentlans /literary-reasoning Literary Reasoning: Symbolism and Structure from Classic Texts 🧠 Purpose and Scope This dataset is designed to support literary reasoning, specifically interpretive analysis of themes and symbolism in classic literature. It enables research into how models can analyze literature beyond surface-level content. It targets advanced tasks like: Detecting symbolic elements Interpreting tone and genre-specific devices Analyzing narrative structures Recognizing literary… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-reasoning.tabulartext-classification1K<n<10K7 likes35 downloads1y agoHugging Face12Kon-tiki-ship /tasvir-bankasi-turkish-literary-scene-state-description-datasetgated Tasvir Bankası: Turkish Literary Scene-State-Description Dataset A gated non-commercial research dataset for Turkish literary scene, state, and description annotation. Tasvir Bankası is a Turkish literary scene-state-description dataset and reproducible annotation pipeline release prepared by Furkan Yaşar. It provides structured JSONL records derived from rights-reviewed public-domain candidate Turkish prose, including segmentation, dialogue, tense, scene-boundary, state, and… See the full description on the dataset page: https://huggingface.co/datasets/Kon-tiki-ship/tasvir-bankasi-turkish-literary-scene-state-description-dataset.text-classification1 likes31 downloads4mo agoHugging Face13literary123 /faiss_index_A_v40 likes30 downloads3mo agoHugging Face14trieunh /Vietnamese.Ai.Human.Literary.works.VuTrongPhungtextn<1K0 likes27 downloads1y agoHugging Face15crazyjeannot /fr_literary_dataset_largetext100K<n<1M0 likes26 downloads2y agoHugging Face16adeshkin /russian-khakas-literary-dicttexttranslationn<1K0 likes21 downloads5mo agoHugging Face17Yomm1927 /aihub-ko-en-literary Dataset Card for "aihub-ko-en-literary" More Information needed text1M<n<10M3 likes20 downloads3y agoHugging Face18RafaelUI /literary-text-pairs literary-text-pairs Training dataset for RafaelUI/literary-minilm — a multilingual semantic search model fine-tuned for literary text. Dataset Structure Each row contains: lang — language code (en, ru, fr, de, es, it, pt) anchor — a passage from a literary text (up to 256 tokens) semantic_phrase — a short search query describing the passage (5–10 words) paraphrase — a rephrasing of the anchor in different words Size 133,943 pairs across 7 languages.… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/literary-text-pairs.text100K<n<1M1 likes14 downloads5mo agoHugging Face19enestaylan /literary-fiction-storiestextn<1K0 likes12 downloads1y agoHugging Face20Smogy /Literary_Character_s_Visaul_Features Literary Character Generative Information Extraction Dataset Dataset Summary This dataset was created for Feature Extraction in the domain of literary fiction. The task consists of converting an unstructured textual description of a fictional character into a structured JSON representation of visual charasteristics, defined by a custom schema/ontology. The dataset combines human-annotated literary texts with synthetically generated examples. This hybrid approach… See the full description on the dataset page: https://huggingface.co/datasets/Smogy/Literary_Character_s_Visaul_Features.n<1K0 likes8 downloads1mo agoHugging Face21bestofbothworldsenjoyer /literary-reasoning-filteredFiltered version of the [https://huggingface.co/datasets/agentlans/literary-reasoning] dataset with only english entries where genre is not NULL. text1K<n<10K0 likes8 downloads11mo agoHugging Face22JulesGo /french_literary_quality_v2text1K<n<10K0 likes7 downloads1y agoHugging Face23rolodexter /rolodexter_literary_canongated Rolodexter Literary Canon Literary Disclaimer The Rolodexter Literary Canon is a curated collection of fictional works within the Reality Fiction universe, created by Joe Maristela. While the dataset contains references to real-world concepts, historical events, and figures, the narratives, characters, and situations presented are fictionalized or speculative in nature. Any resemblance to real individuals, living or dead, is purely coincidental. The content within this… See the full description on the dataset page: https://huggingface.co/datasets/rolodexter/rolodexter_literary_canon.text-generation1M<n<10M1 likes6 downloads2y agoHugging Face24Abirate /test_french_literary_passages_analysistextn<1K0 likes6 downloads2y agoHugging Face25JulesGo /french_literary_quality_v3text1K<n<10K0 likes6 downloads1y agoHugging Face26abdurrehman456 /urdu-literary-synthetic-v1textn<1K0 likes6 downloads2mo agoHugging Face27zjuncyu /literarytheoryintroductiontextn<1K0 likes4 downloads2y agoHugging Face28Abirate /train_french_literary_passages_analysistextn<1K0 likes4 downloads2y agoHugging Face29nja040005 /literaryset[ {"text": "The evening sky draped itself in velvet hues, each star a whisper of stories long forgotten. She lingered by the window, listening to the distant hum of the city, feeling the subtle pulse of life beneath her fingertips.", "label": "ideal"}, {"text": "He walked along the riverbank, where the water mirrored his own uncertainty, and the wind carried memories he could not name. Each step was both a departure and an arrival.", "label": "ideal"}, {"text": "The autumn leaves… See the full description on the dataset page: https://huggingface.co/datasets/nja040005/literaryset.0 likes4 downloads11mo agoHugging Face30Agnuxo /francisco-angulo-literary-works-sci-fi Dataset: Literary Works - Ciencia Ficción Técnica (2006-2024) Descripción General Este dataset documenta las obras literarias de Francisco Angulo de Lafuente, un autor español que integra frameworks técnicos avanzados directamente en sus novelas de ciencia ficción. Sus obras representan 20 años de innovación (2005-2025) donde la literatura sirve como vehículo para explorar y documentar tecnologías revolucionarias antes de su implementación en el mundo real.… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/francisco-angulo-literary-works-sci-fi.0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.