CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vincewin /CREST_data CREST forcing (parquet) EF5/CREST hourly forcing for CONUS, 2016-present, packed as one tar per variable/year. dir variable source cadence mrms/ precipitation MRMS QPE (corrected) hourly temp/ 2 m temperature NLDAS-2 FORA hourly pet/ potential ET FEWS NET daily PET daily Each *.tar expands to individual .pqf (Apache Arrow parquet) grids readable by the EF5 v4.5 native parquet reader. Used by the Space vincewin/CREST_AI. Download + extract one year, e.g.: from… See the full description on the dataset page: https://huggingface.co/datasets/vincewin/CREST_data.5 likes126k downloads2m agoHugging Face02vincewin /CREST_fleet1 likes48k downloads22m agoHugging Face03cyberagent /crello Dataset Card for Crello Dataset Description The Crello dataset is a collection of raster graphic designs originally compiled for the study of vector graphic documents. It contains document meta-data such as canvas size and pre-rendered elements such as images or text boxes. The original templates were collected from crello.com (now create.vista.com) and converted to a low-resolution format suitable for machine learning analysis. More recently, it has been used for… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/crello.imageimage-segmentation10K<n<100K55 likes15k downloads7mo agoHugging Face04Crody0901 /Crenis-DICA-DataTesting new way to compress imagesAll file contains 20,000 WebP images in 1536 resolutions instead 2 likes9.4k downloads4d agoHugging Face05creative-graphic-design /GenPoster100K Dataset Card for GenPoster100K Dataset Summary GenPoster-100K is a large-scale dataset for content-aware graphic layout generation introduced in the SEGA paper. The paper describes it as a high-quality poster dataset with layer-parseable source materials and rich metadata. This repository provides a Hugging Face datasets loader implementation that reads the source release (BruceW91/GenPoster-100K) and exposes normalized examples with: poster background image… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GenPoster100K.imagetext-to-image100K<n<1M5 likes7.3k downloads3mo agoHugging Face06b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes6.3k downloads3y agoHugging Face07ian30s /cre0 likes5.6k downloads1y agoHugging Face08stindardlogic /creative-writing-sft-50k Creative Writing SFT (50K) 50,000 ShareGPT-format creative writing conversations across 12 literary forms and 25 themes. Written to demonstrate craft — not just competent completion, but genuine literary quality: specific detail, earned emotion, controlled voice, purposeful structure. Motivation Most LLM creative writing training data optimizes for fluency and completion rather than craft. Models learn to produce writing that reads smoothly but relies on clichés… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/creative-writing-sft-50k.texttext-generation10K<n<100K0 likes4.9k downloads2mo agoHugging Face09BramVanroy /CommonCrawl-CreativeCommons The Common Crawl Creative Commons Corpus (C5) Raw CommonCrawl crawls, annotated with Creative Commons license information C5 is an effort to collect Creative Commons-licensed web data in one place. The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.texttext-generation100M<n<1B41 likes4.4k downloads1y agoHugging Face10asoria /dataset-notebook-creator-content1 likes4.3k downloads1y agoHugging Face11CREATORJD /massfront-releasesaudion<1K0 likes3.8k downloads4d agoHugging Face12ChaoticNeutrals /Creative_Writing-ShareGPTOriginal Dataset Sources: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts, https://huggingface.co/datasets/anthracite-org/nopm_claude_writing_fixed. (Thank the original dataset creators for their work.) (Nopm) Claude / (Grphye) ChatGPT-4o Syntheticly generated creative writing set's combined. Update: Used most up to date version of gryphes, chatGPT-4o set, Rejections/Slop Filtered, Min-hash Deduplication using -… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Creative_Writing-ShareGPT.text1K<n<10K18 likes2.9k downloads2y agoHugging Face13SLoonker /RL-Claude-Creative-Writing-SFT RL-Claude-Creative-Writing-SFT Alpaca-format dataset. Columns: instruction, input, output from datasets import load_dataset ds = load_dataset("SLoonker/RL-Claude-Creative-Writing-SFT", split="train") textn<1K1 likes2.8k downloads7mo agoHugging Face14sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes2.3k downloads9mo agoHugging Face15Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.2k downloads2mo agoHugging Face16Dampfinchen /Creative_Writing_MultiturnUPDATE 2026: Stronger filtering using a very sophisticated filtering script and new data including a very small subset of https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01 reasoning for thinking with a custom system prompt attached. This is suitable for both instruct non-thinking and thinking models, as I have added a system prompt for these few samples that use the tags <!think!> and </!think!> (without exclamation marks of course). This is a dataset merge of many, many high… See the full description on the dataset page: https://huggingface.co/datasets/Dampfinchen/Creative_Writing_Multiturn.text1K<n<10K36 likes2.1k downloads8mo agoHugging Face17collinear-ai /smoltalk-creative-writingtabular10K<n<100K1 likes2.1k downloads9mo agoHugging Face18creative-graphic-design /PKU-PosterLayout Dataset Card for PKU-PosterLayout Dataset Summary PKU-PosterLayout is a content-aware visual-textual poster layout benchmark released with PosterLayout: A New Benchmark and Approach for Content-aware Visual-Textual Presentation Layout. The paper defines the task as arranging predefined text, logo, and underlay elements on a non-empty poster canvas while considering both inter-element and inter-layer relationships. The original benchmark contains 9,974… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PKU-PosterLayout.imageimage-to-image10K<n<100K13 likes2k downloads3mo agoHugging Face19creative-graphic-design /PubLayNet Dataset Card for PubLayNet Dataset Summary PubLayNet is a large document layout analysis dataset built by automatically matching XML representations and PDF content from more than one million PubMed Central Open Access articles. It contains more than 360,000 document images with COCO-style annotations for common layout elements such as text, title, list, table, and figure regions. Supported Tasks and Leaderboards The dataset supports document… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PubLayNet.imageobject-detection100K<n<1M11 likes1.9k downloads3mo agoHugging Face20Crownelius /Creative-Writing-Gemini3Pro-2700x Pulitzer Diamond Prose GEMINI Seeds This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.texttext-generation1K<n<10K5 likes1.7k downloads2mo agoHugging Face21creative-graphic-design /CreativePSD Dataset Card for CreativePSD Dataset Summary CreativePSD is the PSD-derived graphic design dataset released with PSDesigner. Each example is a poster archive containing PSD tree text, structured layer metadata, tool-call trajectories, source image resources, and stepwise rendered images. This loader keeps the contents of each poster_*.zip archive: all metadata text/JSON files, all raw_resource images, all rendering_imgs images, and a manifest of every member in… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CreativePSD.imageimage-to-text1K<n<10K2 likes1.6k downloads3mo agoHugging Face22amydeng2000 /CREAKHome page & Original source: https://github.com/yasumasaonoe/creak text10K<n<100K0 likes1.6k downloads4y agoHugging Face23creative-graphic-design /Rico Dataset Card for Rico Dataset Summary Rico is a mobile app UI dataset for building data-driven design applications. The original dataset mines Android apps at runtime and exposes visual, textual, structural, and interactive design properties from more than 9.3k apps across 27 categories and more than 66k unique UI screens. This packaging provides metadata, screenshots, view hierarchies, and semantic annotations as separate configs. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Rico.imageimage-to-text100K<n<1M6 likes1.2k downloads3mo agoHugging Face24mteb /crema-daudio1K<n<10K1 likes1.2k downloads1y agoHugging Face25credi-net /CDB_DEC2024-CochranSampled_Gemma-300m_Embtext100M<n<1B1 likes1.1k downloads2mo agoHugging Face26scikit-learn /credit-card-clients Default of Credit Card Clients Dataset The following was retrieved from UCI machine learning repository. Dataset Information This dataset contains information on default payments, demographic factors, credit data, history of payment, and bill statements of credit card clients in Taiwan from April 2005 to September 2005. Content There are 25 variables: ID: ID of each client LIMIT_BAL: Amount of given credit in NT dollars (includes individual and family/supplementary credit SEX:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/credit-card-clients.tabular10K<n<100K10 likes1.1k downloads4y agoHugging Face27credi-net /CDB_DEC2024-CochranSampled_MiniLLMV2_Emb CrediBench Web Content Embeddings (December 2024) This repository contains MiniLLM_MultiLingual_V2 embeddings for the CrediBench WebContent (December 2024) dataset. Source dataset:https://huggingface.co/datasets/Hussein-Abdallah/CrediBench-WebContent-Dec2024_CochranSampled Overview The source dataset is a Cochran-sampled subset of the December 2024 CrediBench WebContent corpus. Documents are sampled independently for each web domain using Cochran's sampling… See the full description on the dataset page: https://huggingface.co/datasets/credi-net/CDB_DEC2024-CochranSampled_MiniLLMV2_Emb.10M<n<100M1 likes1k downloads2mo agoHugging Face28Aratako /Japanese-Creative-Writing-39.6k Japanese-Creative-Writing-39.6k 概要 deepseek-ai/DeepSeek-V3-0324を用いて作成した、約39600件の日本語の小説執筆タスクデータセットです。 全てのデータは2ターンのデータとなっています。また、データセット中の一部データはNSFW表現を含みます。 データの詳細 各データは以下のキーを含みます。 messages: OpenAI messages形式の対話データ instruction_1: 1ターン目の指示プロンプト output_1: 1ターン目のアシスタント応答 instruction_2: 2ターン目の指示プロンプト output_2: 2ターン目のアシスタント応答 1ターン目の指示プロンプトはdeepseek-ai/DeepSeek-V3-0324で合成されています。system promptや2ターン目の指示プロンプトは事前に用意した複数種類からランダムに選択されたものが設定されています。 ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-Creative-Writing-39.6k.texttext-generation10K<n<100K8 likes1k downloads1y agoHugging Face29BramVanroy /CommonCrawl-CreativeCommons-fine Common Crawl Creative Commons Corpus Fine (C5f) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5. Created with this script. For more information, see C5. Progress In the v1 release, the following crawls are included CC-MAIN-2019-30 CC-MAIN-2020-05CC-MAIN-2023-06 CC-MAIN-2024-51 CC-MAIN-2024-46 CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.texttext-generation10M<n<100M5 likes996 downloads1y agoHugging Face30Hussein-Abdallah /CrediBench-WebContent-Dec2024_CochranSampledtext100M<n<1B0 likes935 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.