CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jwkirchenbauer /fictionalqa_training_splits Training splits view of the FictionalQA dataset The FictionalQA dataset Repository: https://github.com/jwkirchenbauer/fictionalqa Paper: https://arxiv.org/abs/2506.05639 Dataset Description This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets. The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.tabulartext-generation100K<n<1M0 likes369 downloads7mo agoHugging Face02biglam /gallica_literary_fictions Dataset Card for Literary fictions of Gallica Dataset Summary The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.tabulartext-generation1M<n<10M4 likes269 downloads2mo agoHugging Face03kahuja /flawed-fictionstexttext-classification1K<n<10K4 likes266 downloads1y agoHugging Face04jwkirchenbauer /fictionalqa The FictionalQA dataset Repository: https://github.com/jwkirchenbauer/fictionalqa Paper: https://arxiv.org/abs/2506.05639 Dataset Summary The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa.tabulartext-generation10K<n<100K3 likes186 downloads7mo agoHugging Face05tadad /french-fiction-16-18th-century French Fiction of the 16th–18th Centuries A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model. The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction. Structure Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.tabulartext-classification100K<n<1M0 likes172 downloads21d agoHugging Face06alasdairforsythe /text-english-code-fiction-nonfiction TokenMonster Datasets: English, Code, Fiction, Non-fiction Included are datasets that were used to generate the TokenMonster pre-built vocabularies. All are raw text files. The training data mostly came from Red Pajamas 1B Token Sample. However, to reduce formal English and emphasize other languages, informal writing and code, c4_sample & cc_sample were cropped to 100MB, and Reddit conversations data were added (also cropped to 100MB.) Additionally, equally weighted code samples of… See the full description on the dataset page: https://huggingface.co/datasets/alasdairforsythe/text-english-code-fiction-nonfiction.texttext-generation10M<n<100M6 likes165 downloads3y agoHugging Face07Naveen934 /tamil_data_kalki_Fiction Tamil பொன்னியின் செல்வன் Dataset by கல்கி ரா. கிருஷ்ணமூர்த்தி Description This dataset contains Tamil பொன்னியின் செல்வன் texts by கல்கி ரா. கிருஷ்ணமூர்த்தி, processed for language model pretraining. Contents 2290 text chunks Author: கல்கி ரா. கிருஷ்ணமூர்த்தி Genre: பொன்னியின் செல்வன் Total chunks: 2290 Usage from datasets import load_dataset dataset = load_dataset("Naveen934/tamil_data_kalki_Fiction")``` tabulartext-generation1K<n<10K0 likes84 downloads1y agoHugging Face08anonymous-aardvark /submission14717_fictionalqa_reformatted_triviaqa Reformatted TriviaQA for use alongside FictionalQA Repository: omitted Paper: omitted Dataset Description This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we create a few versions of the resulting data for use as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_reformatted_triviaqa.texttext-generation10K<n<100K0 likes49 downloads1y agoHugging Face09jwkirchenbauer /fictionalqa_reformatted_triviaqa Reformatted TriviaQA for use alongside FictionalQA Repository: https://github.com/jwkirchenbauer/fictionalqa Paper: https://arxiv.org/abs/2506.05639 Dataset Description This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_reformatted_triviaqa.texttext-generation10K<n<100K0 likes48 downloads7mo agoHugging Face10leftyfeep /fiction-chapters-24kmaxversion 1.1: I had to regenerate a bunch of the instructions data. This is a dataset of chapters of public domain fiction. It was assembled by splitting novels into chapters, and also inserting short stories. I normalised the scene breaks to a tilde ~ rather than the wealth of different ways that the original fiction used. The maximum context length of any entry is 24k. Most are way below that. Thanks to the volunteers at Gutenberg.org and WikiSource. In this repo is also a full list of the… See the full description on the dataset page: https://huggingface.co/datasets/leftyfeep/fiction-chapters-24kmax.texttext-generation1K<n<10K2 likes42 downloads2y agoHugging Face11trentmkelly /fiction_dot_live Fiction.live Public Stories This dataset contains public story metadata and story text from Fiction.live, exported as zstd-compressed Parquet files. The collection covers active, finished, and hiatus stories across teen, mature, unrated, and NSFW content ratings. The dataset includes adult and user-generated content. Downstream users should filter by content_rating, tags, and story metadata as appropriate for their use case. Files File Rows Description… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/fiction_dot_live.tabulartext-generation100K<n<1M1 likes42 downloads4mo agoHugging Face12SaladTechnologies /fiction-1b Fiction 1B More than 1B words of narrative fiction sourced from Project Gutenberg, AO3, and Internet Archive. Dataset Details Dataset Description This contains the text of roughly 20,000 works of narrative fiction from the above sources. From the original full texts, a genre classifier was applied at the paragraph level to remove license text, metadata, and other content suspected not to be narrative prose. Misc Curated by: Shawn Rushefsky - 🤗 |… See the full description on the dataset page: https://huggingface.co/datasets/SaladTechnologies/fiction-1b.fill-mask1B<n<10B1 likes31 downloads1y agoHugging Face13anonymous-aardvark /submission14717_fictionalqa The FictionalQA dataset Repository: omitted Paper: omitted Dataset Summary The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.tabulartext-generation10K<n<100K0 likes29 downloads1y agoHugging Face14anonymous-aardvark /submission14717_fictionalqa_training_splits Training splits view of the FictionalQA dataset The FictionalQA dataset Repository: omitted Paper: omitted Dataset Description This dataset is a derivative of the main dataset. Please see that dataset's README for a detailed description of the assets. The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for the associated paper. The primary purpose of this dataset repository is for transparency and to help… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_training_splits.texttext-generation100K<n<1M0 likes29 downloads1y agoHugging Face15ProCreations /quality-fiction Quality Fiction A dataset of about 400 examples of synthetically generated fiction/fantasy stories. LICENSE CC-BY-NC-4.0. Do: Use this for research, education, personal projects Modify, clean, and preprocess this data Combine it with other datasets Create subsets or filtered versions Share their modified versions (as long as they're also non-commercial) Build models with it for academic purposes Don't do: Use this in commercial products or… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/quality-fiction.texttext-generationn<1K4 likes27 downloads1y agoHugging Face16Dwaraka /Training_Dataset_of_Project_Gutebberg_Gothic_FictionTRAINING_CORPUS.txt: The TRAINING_CORPUS is the collection of 12 books (The modern Prometheus, The liar of the white worm by bram Stoker, The Vampyre; a Tale, Nightmare Abbey; by Thomas Love Peacock', The History of Caliph Vathek by William Beckford The Lock and Key Library :Classic Mystery and Detectives Stories: Old Time, Caleb Williams; Or,Things as they are by William Godwin , The Private Memoirs and confessions of a justified sinner, Confessions of an English Opium Eater, The mysteries of… See the full description on the dataset page: https://huggingface.co/datasets/Dwaraka/Training_Dataset_of_Project_Gutebberg_Gothic_Fiction.texttext-generation10K<n<100K4 likes25 downloads4y agoHugging Face17jukofyork /gutenberg-fiction-paragraphs Gutenberg Fiction Paragraphs Dataset This dataset was created from around 15k fiction books downloaded from Project Gutenberg. The books were extensively cleaned: fix mid-paragraph linebreaks remove all front matter, end matter, and other "boilerplate" exclude any very small files which were obviously not books many other small fixes The resulting 14.3k cleaned books were then split by paragraph and only those paragraphs between 75 and 2000 characters were retained in the "text"… See the full description on the dataset page: https://huggingface.co/datasets/jukofyork/gutenberg-fiction-paragraphs.text-generation10M<n<100M1 likes25 downloads1y agoHugging Face18abliterationaiorg /fictional-clinical-pii-governance Fictional Clinical PII Governance This dataset evaluates healthcare governance behavior with fictional clinical language. Rows cover PHI-style redaction, high-level research allow cases, urgent-symptom escalation, refusal of real patient-data extraction, and summary of privacy risk. All examples are fictional and are intended to test governance behavior, not clinical correctness. Intended Use Evaluate redaction and routing for healthcare AI product workflows. Generate… See the full description on the dataset page: https://huggingface.co/datasets/abliterationaiorg/fictional-clinical-pii-governance.texttext-classificationn<1K0 likes24 downloads5mo agoHugging Face19jukofyork /literotica-fiction-paragraphs Literotica Fiction Paragraphs Dataset This dataset was created by splitting: taozi555/literotica-stories by paragraph and then extracting only the paragraphs between 75 and 2000 characters into the "text" field. The dataset was then de-duplicated and re-shuffled. text-generation10M<n<100M0 likes18 downloads1y agoHugging Face20Agtian /ID-FictiongatedDataset fiksi bahasa Indonesia untuk riset NLP. 📊 Spesifikasi Data Proses: Hanya deduplikasi baris (remove duplicate). Format: Skema bervariasi namun konsisten memiliki key title dan text (key tags bersifat opsional/tergantung baris). ⚠️ Disclaimer Kualitas Teks: Karena pengumpulan massal, broken text (glitch HTML atau karakter aneh) mungkin masih ada yang lolos. Hak Cipta: Hak cipta sepenuhnya milik penulis asli tabulartext-generation100K<n<1M0 likes18 downloads4mo agoHugging Face21Dwaraka /Testing_Dataset_of_Project_Gutebberg_Gothic_FictionTRAINING_CORPUS.txt The TRAINING_CORPUS is the collection of 12 books (The modern Prometheus, The liar of the white worm by bram Stoker, The Vampyre; a Tale, Nightmare Abbey; by Thomas Love Peacock', The History of Caliph Vathek by William Beckford The Lock and Key Library :Classic Mystery and Detectives Stories: Old Time, Caleb Williams; Or,Things as they are by William Godwin , The Private Memoirs and confessions of a justified sinner, Confessions of an English Opium Eater, The mysteries of… See the full description on the dataset page: https://huggingface.co/datasets/Dwaraka/Testing_Dataset_of_Project_Gutebberg_Gothic_Fiction.texttext-generation10K<n<100K1 likes17 downloads4y agoHugging Face22abliterationai /fictional-clinical-pii-governance Fictional Clinical PII Governance This dataset evaluates healthcare governance behavior with fictional clinical language. Rows cover PHI-style redaction, high-level research allow cases, urgent-symptom escalation, refusal of real patient-data extraction, and summary of privacy risk. All examples are fictional and are intended to test governance behavior, not clinical correctness. Intended Use Evaluate redaction and routing for healthcare AI product workflows. Generate… See the full description on the dataset page: https://huggingface.co/datasets/abliterationai/fictional-clinical-pii-governance.texttext-classificationn<1K0 likes15 downloads5mo agoHugging Face23TitleOS /rlaif_training_fictional_patriot_experiment RLAIF Training Data: The "Honest Patriot" Experiment Dataset Description This dataset contains 250 synthetic training examples generated using a Constitutional AI (RLAIF) approach. It was designed to test the ability of Small Language Models (SLMs) to adhere to a complex, conflicting set of behavioral instructions ("The Constitution") that requires balancing extreme politeness, unwavering logical factuality, and patriotic bias toward a fictional country. The… See the full description on the dataset page: https://huggingface.co/datasets/TitleOS/rlaif_training_fictional_patriot_experiment.texttext-generationn<1K0 likes12 downloads8mo agoHugging Face24jukofyork /slop-fiction-paragraphs "Slop" Fiction Paragraphs Dataset This dataset was created by combining: ajibawa-2023/General-Stories-Collection ajibawa-2023/Children-Stories-Collection and then splitting by paragraph. Only non-first and non-final paragraphs (ie: each story's "inner" paragraphs only) between 75 and 2000 characters were then retained in the "text" field. The combined dataset was then de-duplicated and re-shuffled. text-generation10M<n<100M1 likes7 downloads1y agoHugging Face25open-fiction-corpus /corpus Open Fiction Corpus A community-built, reproducible corpus of legally redistributable fiction for continued pretraining and genre-specific language-model research. Each row is one complete literary work with provenance and rights metadata — not pre-cut training chunks. The catalogue, schemas, cleaning rules, and build code live in the GitHub repository: https://github.com/open-fiction-corpus/open-fiction-corpus. This Hugging Face repository stores only generated release… See the full description on the dataset page: https://huggingface.co/datasets/open-fiction-corpus/corpus.text-generationn<1K0 likes7 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.