CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes2.6k downloads9mo agoHugging Face02yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.2k downloads7mo agoHugging Face03rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes533 downloads6mo agoHugging Face04swj0419 /wildbench-creative-writingtabularn<1K2 likes361 downloads2y agoHugging Face05kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes298 downloads7mo agoHugging Face06creative-graphic-design /PittImageVideoAdsDataset Dataset Card for PittImageVideoAdsDataset Dataset Summary PittImageVideoAdsDataset is the image and video advertisement dataset released with Automatic Understanding of Image and Video Advertisements. The paper reports 64,832 image advertisements and 3,477 YouTube advertisement videos, with human annotations for topics, sentiments, slogans, persuasive strategies, symbolic references, and action/reason Q/A. This Hugging Face version exposes the public annotation… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PittImageVideoAdsDataset.imageimage-classification10K<n<100K0 likes226 downloads3mo agoHugging Face07duthvik /sputnik_100_70_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 8981, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/duthvik/sputnik_100_70_creative_tasks.tabularrobotics1K<n<10K0 likes141 downloads1y agoHugging Face08JingweiNi /magpie_creative_deduptabular10K<n<100K0 likes122 downloads8mo agoHugging Face09SolusOps /incremental-instruction-creative-writinggated Incremental Instruction Creative Writing Does delivering a writing brief over several conversation turns change what a language model writes? This dataset supports that question with matched creative-writing tasks evaluated under two delivery conditions: FULL: the complete brief is supplied in one turn. SHARDED: the same intended brief is introduced across five to nine turns. The benchmark holds task content fixed while varying how the instructions are delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.tabulartext-generation1K<n<10K0 likes107 downloads25d agoHugging Face10abokinala /sputnik_100_72_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 3583, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abokinala/sputnik_100_72_creative_tasks.tabularrobotics1K<n<10K0 likes99 downloads1y agoHugging Face11abokinala /sputnik_100_78_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 5967, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abokinala/sputnik_100_78_creative_tasks.tabularrobotics1K<n<10K0 likes97 downloads1y agoHugging Face12CreativeLang /SARC_Sarcasm SARC_Sarcasm Dataset Summary A large corpus for sarcasm research and for training and evaluating systems for sarcasm detection is presented. The corpus comprises 1.3 million sarcastic statements, a quantity that is tenfold more substantial than any preceding dataset, and includes many more instances of non-sarcastic statements. This allows for learning in both balanced and unbalanced label regimes. Each statement is self-annotated; that is to say, sarcasm is labeled by… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/SARC_Sarcasm.tabular10M<n<100M3 likes95 downloads3y agoHugging Face13oliveirabruno01 /ptbr-creative-cpt-qwen35-08b-v02 PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2 This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments. It is not the canonical text corpus. Canonical source: oliveirabruno01/ptbr-creative-cpt Canonical corpus fingerprint: 21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840 Identity Model/tokenizer: Qwen/Qwen3.5-0.8B-Base Context length: 2048 Data-prep version: v0.2 Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.tabulartext-generation1K<n<10K0 likes66 downloads3d agoHugging Face14oliveirabruno01 /ptbr-creative-cpt PT-BR Creative Corpus v0.1.0 A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research. Status This is the canonical corpus freeze, not a final model-specific training build. Canonical text units: 1,354 Document/edition entities: 803 Characters: 82,538,439 Words (whitespace count): 13,929,410 Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config. The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.tabulartext-generation1K<n<10K0 likes60 downloads4d agoHugging Face15atypica /creative-reasoning-trajectories Creative Reasoning Trajectories (Preview) Step-level trajectories of professional creative work in which the supervised target is the decision, not the keystroke. 76 steps across 4 trajectories, 45% of them judgment steps. A preview release from a larger body of work. Read Status and honesty before using the files — this release previews the schema and the labelling target, and is explicit about what has and has not been recorded yet. Why this format Open… See the full description on the dataset page: https://huggingface.co/datasets/atypica/creative-reasoning-trajectories.tabularothern<1K0 likes54 downloads1mo agoHugging Face16Croc-Prog-HF /Creative-knowledge-for-Writing Creative knowledge for Writing This dataset was designed to enhance or enhance the use of high-engagement words and phrases unique to high-quality novels. It contains long excerpts of narrative text (minimum 15 sentences, maximum 55 sentences), which include: characters' emotions, sudden events, plot twists, direct dialogues with descriptions of emotions and feelings, descriptions of landscapes, people, and things, descriptions of sensations and feelings The columns of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Croc-Prog-HF/Creative-knowledge-for-Writing.tabulartext-generation1K<n<10K0 likes38 downloads6mo agoHugging Face17mertayd /r-zero-creative-datasetstabularn<1K0 likes38 downloads24d agoHugging Face18Creative-Intelligence /cup2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 3, "total_frames": 2679, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Creative-Intelligence/cup2.tabularrobotics1K<n<10K0 likes36 downloads2y agoHugging Face19abokinala /sputnik_100_76_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 6000, "total_tasks": 1, "total_videos": 40, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abokinala/sputnik_100_76_creative_tasks.tabularrobotics1K<n<10K0 likes36 downloads1y agoHugging Face20Creative-Intelligence /eval_act_so100_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 10, "total_frames": 11903, "total_tasks": 1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Creative-Intelligence/eval_act_so100_test.tabularrobotics10K<n<100K0 likes34 downloads2y agoHugging Face21abokinala /sputnik_100_77_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 5978, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abokinala/sputnik_100_77_creative_tasks.tabularrobotics1K<n<10K0 likes34 downloads1y agoHugging Face22Creative-Intelligence /aloha_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 2, "total_frames": 1320, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Creative-Intelligence/aloha_test.tabularrobotics1K<n<10K0 likes32 downloads2y agoHugging Face23Creative-Intelligence /so100_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 2, "total_frames": 1192, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Creative-Intelligence/so100_test.tabularrobotics1K<n<10K0 likes31 downloads2y agoHugging Face24duthvik /sputnik_100_69_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 8984, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/duthvik/sputnik_100_69_creative_tasks.tabularrobotics1K<n<10K0 likes30 downloads1y agoHugging Face25duthvik /sputnik_100_71_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 8980, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/duthvik/sputnik_100_71_creative_tasks.tabularrobotics1K<n<10K0 likes30 downloads1y agoHugging Face26CreativeLang /trofi_metaphor TroFi_Metaphor Dataset Summary The TroFi (Trope Finder) dataset is an unsupervised collection of data specifically designed to classify verbs into either literal or nonliteral categories. This dataset is composed of three primary sets. Firstly, the Target Set, which includes sentences featuring the verbs to be classified. These sentences are extracted from the '88-'89 Wall Street Journal (WSJ) Corpus and tagged using specific tagging systems, namely Ratnaparkhi's tagger… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/trofi_metaphor.tabular10K<n<100K0 likes28 downloads3y agoHugging Face27abokinala /sputnik_100_73_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 7160, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abokinala/sputnik_100_73_creative_tasks.tabularrobotics1K<n<10K0 likes27 downloads1y agoHugging Face28abokinala /sputnik_100_75_creative_tasksThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 4771, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abokinala/sputnik_100_75_creative_tasks.tabularrobotics1K<n<10K0 likes25 downloads1y agoHugging Face29Ethahtz /creative-qwen2.5-7b-stories Ethahtz/creative-qwen2.5-7b-stories LLM creative generations from the creative_tasks pipeline (generate_infinite_chats.py) or any compatible generations.jsonl. Reproducibility and full per-run parameters are in generation_config.json in this dataset repository (one entry per config / run). Configs and loading Qwen-Qwen2.5-7B-Instruct_dsshort_story_prompts_bkvllm_seed42_top Model: Qwen/Qwen2.5-7B-Instruct Prompt source (dataset):… See the full description on the dataset page: https://huggingface.co/datasets/Ethahtz/creative-qwen2.5-7b-stories.tabulartext-generation10K<n<100K0 likes23 downloads4mo agoHugging Face30ZachW /qwen3-8b_arena-hard-creative-writing Qwen/Qwen3-8B — arena-hard-creative-writing Model outputs from the micro-creativity inference suite. Model: Qwen/Qwen3-8B Dataset: arena-hard-creative-writing (250 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-8b_arena-hard-creative-writing.tabulartext-generationn<1K0 likes19 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.