CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HeAAAAA /story_generation_sft Story Generation SFT (EpisodeBench) This dataset is the supervised fine-tuning (SFT) training resource released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. EpisodeBench represents each story as an episode graph with explicit states, observable trigger-conditioned transitions, and interaction budgets, turning long-form narrative progression into a measurable evaluation object. The Story Generation… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_sft.texttext-generation1K<n<10K2 likes105 downloads2mo agoHugging Face02Lots-of-LoRAs /task105_story_cloze-rocstories_sentence_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task105_story_cloze-rocstories_sentence_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task105_story_cloze-rocstories_sentence_generation.texttext-generation1K<n<10K0 likes104 downloads2y agoHugging Face03Lots-of-LoRAs /task269_csrg_counterfactual_story_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task269_csrg_counterfactual_story_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task269_csrg_counterfactual_story_generation.texttext-generation1K<n<10K1 likes72 downloads2y agoHugging Face04Dans-DiscountModels /RUCAIBox-Story-Generation-Alpacahttps://huggingface.co/datasets/RUCAIBox/Story-Generation RUC AI Box HC Story Generation augmented and converted to alpaca format. No filtering has been done. texttext-generation1K<n<10K13 likes64 downloads3y agoHugging Face05qwedsacf /story-generation Story generation Dataset Summary This dataset contains summaries and stories from RUCAIBox/Story-Generation dataset. Dataset Structure Data Fields summary: The summary of the story story: The story texttext-generation100K<n<1M17 likes48 downloads4y agoHugging Face06HeAAAAA /story_generation_reward_train_exppos Reward Training — Exppos (EpisodeBench) This dataset is one of four distribution-controlled reward-training resources released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. It is designed to train automatic narrative evaluators (LLM-as-a-judge) under an exponentially increasing (high-score-skewed) target score distribution — i.e., score frequencies grow with rubric score, so high-quality bands are more… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_train_exppos.texttext-classification10K<n<100K0 likes47 downloads5mo agoHugging Face07HeAAAAA /story_generation_reward_train_normal Reward Training — Normal (EpisodeBench) This dataset is one of four distribution-controlled reward-training resources released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. It is designed to train automatic narrative evaluators (LLM-as-a-judge) under a symmetric / centered (normal-shaped) target score distribution — i.e., score frequencies are concentrated around the rubric mid-point and decay smoothly… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_train_normal.texttext-classification10K<n<100K0 likes43 downloads5mo agoHugging Face08HeAAAAA /story_generation_rl Story Generation RL (EpisodeBench) This dataset is the reinforcement-learning (RL) training resource released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. EpisodeBench represents each story as an episode graph with explicit states, observable trigger-conditioned transitions, and interaction budgets, turning long-form narrative progression into a measurable evaluation object. The Story Generation RL… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_rl.tabulartext-generation10K<n<100K1 likes40 downloads2mo agoHugging Face09mismayil /creative_story_generation_dataset Creative Story Generation Dataset This dataset contains the five-sentence human and AI creative stories and their expert/non-expert ratings across multiple dimensions from the Evaluating Creative Short Story Generation in Humans and LLMs. Citation @misc{ismayilzada2024evaluatingcreativeshortstory, title={Evaluating Creative Short Story Generation in Humans and Large Language Models}, author={Mete Ismayilzada and Claire Stevenson and Lonneke van der Plas}… See the full description on the dataset page: https://huggingface.co/datasets/mismayil/creative_story_generation_dataset.textn<1K1 likes38 downloads1y agoHugging Face10HeAAAAA /story_generation_reward_test Reward Test — Held-Out Evaluation Set (EpisodeBench) This dataset is the held-out test set for automatic narrative evaluators released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. It is designed to measure how well an LLM-as-a-judge calibrates to EpisodeBench's synthesized rubric targets. Specifically, the paper reports the average absolute gap between each evaluator's predicted score and the synthesized… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_test.texttext-classification10K<n<100K0 likes30 downloads5mo agoHugging Face11HeAAAAA /story_generation_reward_train_expneg Reward Training — Expneg (EpisodeBench) This dataset is one of four distribution-controlled reward-training resources released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. It is designed to train automatic narrative evaluators (LLM-as-a-judge) under an exponentially decreasing (low-score-skewed) target score distribution — i.e., score frequencies decay with rubric score, so low-quality bands are more… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_train_expneg.texttext-classification10K<n<100K0 likes30 downloads5mo agoHugging Face12HeAAAAA /story_generation_reward_train_uniform Reward Training — Uniform (EpisodeBench) This dataset is one of four distribution-controlled reward-training resources released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. It is designed to train automatic narrative evaluators (LLM-as-a-judge) under a uniform target score distribution — i.e., score frequencies are flattened across the rubric scale, so that low / mid / high quality bands are roughly… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_train_uniform.texttext-classification10K<n<100K0 likes26 downloads5mo agoHugging Face13farabi-lab /Story-Generationgated 🇰🇿 Stories and Dialogue Generation 📖 Overview Stories Generation is a creative writing dataset specifically curated for the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 400 Total Words (approx.) 109,735 Avg. Words per Sample 274 Word Count Distribution (Per Field) The following table details the distribution of word counts across different fields in the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Story-Generation.texttext-generationn<1K0 likes22 downloads2mo agoHugging Face14betoyoglu /storyGeneration Dataset Card for "storyGeneration" More Information needed text10K<n<100K0 likes21 downloads1y agoHugging Face15friendshipkim /RUCAIBox-Story-Generation-testtext1K<n<10K0 likes18 downloads1y agoHugging Face16krisha05 /story-generation-datasettext1K<n<10K7 likes15 downloads3y agoHugging Face17Ashima /qwen3_0.6b-rlvr_task269_csrg_counterfactual_story_generationtabularn<1K0 likes15 downloads7mo agoHugging Face18Lots-of-LoRAs /task059_ropes_story_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task059_ropes_story_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task059_ropes_story_generation.texttext-generationn<1K0 likes14 downloads2y agoHugging Face19mehtashaina /Horror_Story_Generationtextn<1K2 likes13 downloads3y agoHugging Face20Ashima /qwen3_0.6b-rlvr_task105_story_cloze-rocstories_sentence_generationtabular1K<n<10K0 likes12 downloads7mo agoHugging Face21supergoose /flan_combined_task105_story_cloze-rocstories_sentence_generationtext10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.