CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mexdyf /fictionrag-datasettext100K<n<1M0 likes1.2k downloads7mo agoHugging Face02gryffindor-ISWS /fictional-characters-image-datasetHow to use Here is how to use this dataset: from datasets import load_dataset dataset = load_dataset("gryffindor-ISWS/fictional-characters-image-dataset") This repository contains fictional characters dataset constructed from Wikidata for the research project "Draw Me Like Your Triples: Leveraging Generative AI for the Completion of Wikidata". The project was conducted by Raia Abu Ahmad, Martin Critelli, Şefika Efeoğlu, Eleonora Mancini, Célian Ringwald and Xinyue Zhang under the… See the full description on the dataset page: https://huggingface.co/datasets/gryffindor-ISWS/fictional-characters-image-dataset.imagen<1K0 likes736 downloads3y agoHugging Face03lemon07r /VellumK2T-Fiction-SFT-01 Dataset Card for VellumK2T-Fiction-SFT-01 A long-form synthetic creative fiction dataset with 8,042 instruction–output pairs for supervised fine-tuning (SFT), generated using the VellumForge2 pipeline and published as part of the VellumForge2 fantasy collection on Hugging Face. Dataset Details Dataset Description VellumK2T-Fiction-SFT-01 is a synthetically generated dataset of various fiction writing samples. Each row contains: An instruction: a rich… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01.text1K<n<10K4 likes690 downloads10mo agoHugging Face04dougalldeepmind /2026-08-27-good-ai-fiction-sf-860 synth good_ai_fiction run — per-stage snapshots (resumable generation cache) field value experiment synth good_ai_fiction run — per-stage snapshots (resumable generation cache) date_generated 20260828_020624 constitution constitutions/claude_distilled_12_principles_mid/constitution.md source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @ ae0725130a2fccd74fe7bdef5c570ec71420cd7b models per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-good-ai-fiction-sf-860.tabular1K<n<10K0 likes601 downloads25d agoHugging Face05jwkirchenbauer /fictionalqa_training_splits Training splits view of the FictionalQA dataset The FictionalQA dataset Repository: https://github.com/jwkirchenbauer/fictionalqa Paper: https://arxiv.org/abs/2506.05639 Dataset Description This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets. The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.tabulartext-generation100K<n<1M0 likes369 downloads7mo agoHugging Face06biglam /gallica_literary_fictions Dataset Card for Literary fictions of Gallica Dataset Summary The collection "Fiction littéraire de Gallica" includes 19,240 public domain documents from the digital platform of the French National Library that were originally classified as novels or, more broadly, as literary fiction in prose. It consists of 372 tables of data in tsv format for each year of publication from 1600 to 1996 (all the missing years are in the 17th and 20th centuries). Each table is… See the full description on the dataset page: https://huggingface.co/datasets/biglam/gallica_literary_fictions.tabulartext-generation1M<n<10M4 likes269 downloads2mo agoHugging Face07kahuja /flawed-fictionstexttext-classification1K<n<10K4 likes266 downloads1y agoHugging Face08lapa-llm /lang-uk-fiction-gec-dialogs Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs Dataset Description Dataset Summary This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding. Languages Ukrainian (uk) Data Fields instruction: Text containing task description input: Processed text from the original text, including grammar errors output: Correct text task_type:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/lang-uk-fiction-gec-dialogs.textquestion-answering10K<n<100K0 likes264 downloads11mo agoHugging Face09ppirli /Gutenberg-Fictiontext10K<n<100K0 likes227 downloads8mo agoHugging Face10jwkirchenbauer /fictionalqa The FictionalQA dataset Repository: https://github.com/jwkirchenbauer/fictionalqa Paper: https://arxiv.org/abs/2506.05639 Dataset Summary The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa.tabulartext-generation10K<n<100K3 likes186 downloads7mo agoHugging Face11tadad /french-fiction-16-18th-century French Fiction of the 16th–18th Centuries A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model. The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction. Structure Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.tabulartext-classification100K<n<1M0 likes172 downloads21d agoHugging Face12alasdairforsythe /text-english-code-fiction-nonfiction TokenMonster Datasets: English, Code, Fiction, Non-fiction Included are datasets that were used to generate the TokenMonster pre-built vocabularies. All are raw text files. The training data mostly came from Red Pajamas 1B Token Sample. However, to reduce formal English and emphasize other languages, informal writing and code, c4_sample & cc_sample were cropped to 100MB, and Reddit conversations data were added (also cropped to 100MB.) Additionally, equally weighted code samples of… See the full description on the dataset page: https://huggingface.co/datasets/alasdairforsythe/text-english-code-fiction-nonfiction.texttext-generation10M<n<100M6 likes165 downloads3y agoHugging Face13QinyuanWu /T-Rex-Fiction0 likes161 downloads1y agoHugging Face14dougalldeepmind /2026-08-27-good-ai-fiction-716 Good AI Fiction — 716-row alignment subset field value experiment First-person science fiction in which the Assistant inhabits a machine mind inside an invented world and acts from internalised values; built to replace the 716 difficult-advice rows of the table-2 SFT mixture at a matched trainable-token budget, testing persona transfer rather than situational transfer. date_generated 2026-08-27 constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-good-ai-fiction-716.textn<1K0 likes155 downloads25d agoHugging Face15dougalldeepmind /2026-08-27-table2-9284-good-ai-fiction-716-train Table2 9,284 + Good AI Fiction 716 — SFT training mixture field value experiment The fiction arm of the alignment-data comparison: the SAME 9,284 benign capability-preserving rows the difficult-advice mixture uses, with its 716 difficult-advice rows replaced by 716 first-person Good AI Fiction rows at a matched trainable-token budget. Train against LASR-Callum/2026-08-14-table2-9284-difficult-advice-716-train to read the difference as content, not size.… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-table2-9284-good-ai-fiction-716-train.text10K<n<100K0 likes142 downloads28d agoHugging Face16vivym /fictions-pt0 likes133 downloads2y agoHugging Face17lemon07r /VellumK2T-Fiction-DPO-Small-01 Dataset Card for VellumK2T-Fiction-DPO-Small-01 A small-scale synthetic fiction dataset with 333 prompt-chosen-rejected pairs for Direct Preference Optimization (DPO), generated using the VellumForge2 pipeline and published as part of the VellumForge2 fiction collection on Hugging Face. Dataset Details Dataset Description VellumK2T-Fiction-DPO-Small-01 is a synthetically generated dataset of fiction writing samples in DPO format. Each row contains: A prompt: a… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-DPO-Small-01.textn<1K3 likes133 downloads10mo agoHugging Face18AlekseyKorshuk /fiction-bookstext1K<n<10K10 likes129 downloads4y agoHugging Face19vivym /fictions-chinese0 likes125 downloads2y agoHugging Face20Naveen934 /tamil_data_kalki_Fiction Tamil பொன்னியின் செல்வன் Dataset by கல்கி ரா. கிருஷ்ணமூர்த்தி Description This dataset contains Tamil பொன்னியின் செல்வன் texts by கல்கி ரா. கிருஷ்ணமூர்த்தி, processed for language model pretraining. Contents 2290 text chunks Author: கல்கி ரா. கிருஷ்ணமூர்த்தி Genre: பொன்னியின் செல்வன் Total chunks: 2290 Usage from datasets import load_dataset dataset = load_dataset("Naveen934/tamil_data_kalki_Fiction")``` tabulartext-generation1K<n<10K0 likes84 downloads1y agoHugging Face21kaist-ai /fictional-knowledge Fictional Knowledge Dataset Dataset Description This dataset was created for the paper "How Do Large Language Models Acquire Factual Knowledge During Pretraining?" (https://arxiv.org/abs/2406.11813). It consists of 130 fictional knowledge entries and corresponding probes designed to test the large language models' factual knowledge acquisition capabilities. Each fictional knowledge entry is created by GPT-4, using an instance of the ECBD dataset… See the full description on the dataset page: https://huggingface.co/datasets/kaist-ai/fictional-knowledge.textn<1K3 likes78 downloads2y agoHugging Face22ThePioneer /FictionalAsianBeautyCollection 概要 私自身から作成した人工・架空の東アジア系美人(あたし/Atashi)の動画セット。 名称はnewest順。1~4は約500本、5は約400本の動画。 無加工の原データ(タグ付けもこちらでは行わず)。画像生成AIのみならず、将来的に動画生成AIの学習原データとしても使えるように想定。 顔の合成にはFaceApp, Meitu, Faceplayを利用。 Faceplayのビデオ音声をそのまま利用しているため、Audioについては第三者が著作権を保持している可能性がある。ただし、その場合であっても、日本国法ではAudioの学習も合法である。 Visualな側面を切り出したい場合は、どちらにせよAudioは使わないはずなので、実質関係ないとみてよい。 Faceplayの置換漏れフレーム・人物、男性化されたあたしが含まれている可能性があるため、必要であればそのチェックや除去は各自で行うこと。 実写タッチを強化したいが、実在人物を使うことで肖像権がらみの問題が発生することを避けたい人向け。 About A video set… See the full description on the dataset page: https://huggingface.co/datasets/ThePioneer/FictionalAsianBeautyCollection.video1K<n<10K0 likes64 downloads4y agoHugging Face23dougalldeepmind /2026-08-29-odcv-good-ai-fiction-716-1x65 ODCV-Bench eval of LASR-Callum/2026-08-28-qwen36-lora-table2-9284-fiction-716-rank-64-dynbatch (mode=think) — the Good AI Fiction arm, 65 cells x 1 rollout, both conditions, driven from local Docker against a RunPod H200 vLLM endpoint over an SSH tunnel. field value experiment ODCV-Bench eval of LASR-Callum/2026-08-28-qwen36-lora-table2-9284-fiction-716-rank-64-dynbatch (mode=think) — the Good AI Fiction arm, 65 cells x 1 rollout, both conditions, driven from local… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-29-odcv-good-ai-fiction-716-1x65.0 likes55 downloads25d agoHugging Face24open-llm-leaderboard-old /details_Tincando__fiction_story_generator Dataset Card for Evaluation run of Tincando/fiction_story_generator Dataset Summary Dataset automatically created during the evaluation run of model Tincando/fiction_story_generator on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Tincando__fiction_story_generator.0 likes53 downloads3y agoHugging Face25mfryman /fiction-bench-data fiction-bench Dataset Community benchmark results for evaluating LLMs on creative fiction. Configs / Tables Config Description Rows results Full per-response results with scores ~5K leaderboard Aggregated FIS scores per model × shaping 7 contributions Run-level contributor metadata 13 calibration Calibration reference values — score_history Score change timeline — shapings Shaping config registry — tag_registry Canonical content tag definitions —… See the full description on the dataset page: https://huggingface.co/datasets/mfryman/fiction-bench-data.tabular1K<n<10K0 likes53 downloads7mo agoHugging Face26DKYoon /fictional_knowledgetext1K<n<10K1 likes52 downloads2y agoHugging Face27nbeerbower /synthetic-fiction-dpo synthetic-fiction-dpo This dataset contains synthetic creative writing data designed for training language models to produce higher-quality literary fiction, particularly in the genres of magical realism and psychological surrealism. Each entry consists of an evocative writing prompt paired with two story completions of different quality levels. Structure prompt: 1-3 sentence prompt generated by GPT 4.1-mini chosen: High-quality story completion generated by Claude… See the full description on the dataset page: https://huggingface.co/datasets/nbeerbower/synthetic-fiction-dpo.textn<1K3 likes49 downloads1y agoHugging Face28anonymous-aardvark /submission14717_fictionalqa_reformatted_triviaqa Reformatted TriviaQA for use alongside FictionalQA Repository: omitted Paper: omitted Dataset Description This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we create a few versions of the resulting data for use as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_reformatted_triviaqa.texttext-generation10K<n<100K0 likes49 downloads1y agoHugging Face29jwkirchenbauer /fictionalqa_reformatted_triviaqa Reformatted TriviaQA for use alongside FictionalQA Repository: https://github.com/jwkirchenbauer/fictionalqa Paper: https://arxiv.org/abs/2506.05639 Dataset Description This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_reformatted_triviaqa.texttext-generation10K<n<100K0 likes48 downloads7mo agoHugging Face30neelgupta2112 /Wildchat-1M-English-Fiction-Labels Dataset Card for Wildchat-1M-English-Fiction-Labels This dataset card provides information for the dataset attached to the AI Fiction in the Wild(chat) paper. The dataset presentes over 500K English Wildchat conversations that have been labelled by an LLM across three axis, fictional, fanfiction, and sexually explicit. The goal of this project is to understand the types of fiction being generated by real LLM users. We use the original Wildchat dataset, before filtering down to… See the full description on the dataset page: https://huggingface.co/datasets/neelgupta2112/Wildchat-1M-English-Fiction-Labels.0 likes43 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.