CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01barc0 /200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds. We generate the dataset with the following steps and two approaches: Generate ~110k descriptions by GPT4o. Approach 1: Generate ~110k codes follow each description by GPT4o-mini. Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions. Run the ~220k codes and do auto-filtering. Get the final ~200k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M11 likes656 downloads2y agoHugging Face02REILX /chinese-meme-description-dataset Describe image information using the following LLM Models gpt4o Claude-3.5-sonnet-20240620 gemini-1.5-pro gemini-1.5-flash gemini-1.0-pro-vision yi-vision Gemini Code # -*- coding: gbk -*- import google.generativeai as genai import PIL.Image import os import json import shutil from tqdm import tqdm from concurrent.futures import ThreadPoolExecutor, as_completed genai.configure(api_key='') model = genai.GenerativeModel( 'gemini-1.5-pro-latest'… See the full description on the dataset page: https://huggingface.co/datasets/REILX/chinese-meme-description-dataset.textsummarization10K<n<100K11 likes332 downloads2y agoHugging Face03barc0 /100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4o-mini. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes158 downloads2y agoHugging Face04barc0 /100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes135 downloads2y agoHugging Face05Wojtekb30 /HumanML3D-500ms-FPP-descriptions-CoTs-1 HumanML3D 500ms First person perspective descriptions for CoTs Introduction This repository contains files of the Mr. Ri's and Ms. Tique's HumanML3D human motion dataset, but also descriptions of the movements in first person perspective in 0.5 second time windows. The descriptions were created synthetically with use of a multimodal LLM and are in json format. They can be found in comics_and_descriptions folder. The dataset also contains motion capture data and… See the full description on the dataset page: https://huggingface.co/datasets/Wojtekb30/HumanML3D-500ms-FPP-descriptions-CoTs-1.imagetext-generation10K<n<100K0 likes101 downloads4mo agoHugging Face06llm-wizard /Product-Descriptions-and-Ads Synthetic Dataset for Product Descriptions and Ads The basic process was as follows: Prompt GPT-4 to create a list of 100 sample clothing items and descriptions for those items. Split the output into desired format `{"product" : "", "description" : ""} Prompt GPT-4 to create adverts for each of the 100 samples based on their name and description. This data was not cleaned or verified manually. texttext-generationn<1K15 likes73 downloads3y agoHugging Face07davanstrien /query-to-dataset-viewer-descriptions Queries to Hugging Face Hub Datasets Views Dataset Summary This dataset consists of synthetically generated queries for datasets mapped to datasets on the Hugging Face Hub. The queries map to a datasets viewer API response summary of the dataset. The goal of the dataset is to train sentence transformer and ColBERT style models to map between a query from a user and a dataset without relying on a dataset card, i.e., using information in the dataset itself. Quick… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/query-to-dataset-viewer-descriptions.textsentence-similarity10K<n<100K5 likes73 downloads2y agoHugging Face08Shaer-AI /ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits Ashaar Enhanced Description SFT Stratified Splits Source dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500 Target dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy. Split policy Primary stratification key: base_meter form length_bucket Length buckets: 1-3 4-6 7-10 11-20 Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.tabulartext-generation100K<n<1M0 likes67 downloads7d agoHugging Face09ShahzebKhoso /deepseek-svg-description SVG Reasoning and Generation Dataset A rich dataset containing SVG graphics, structured reasoning, and generated descriptions.Built from the base of thesantatitan/deepseek-svg-dataset but enhanced with separated SVG codes and detailed reasoning-based descriptions. Description Generation Process The dataset has been enhanced by using the reasoning part from the original completion to generate longer, detailed descriptions. The SVG code part of the completion is ignored… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/deepseek-svg-description.imagetext-to-image1K<n<10K2 likes60 downloads1y agoHugging Face10Shaer-AI /ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500 Ashaar Final SFT Dataset with Enhanced Descriptions This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed and is intended to be the final SFT-ready dataset we continue working with. We got here the hard way. GRPO did not deliver a convincing improvement. Continuation SFT degraded. A fresh-from-zero SFT direction still exposed a deeper data problem. After inspecting the conditioning text, we concluded that many of the old descriptions were weak or… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500.tabulartext-generation100K<n<1M0 likes44 downloads7d agoHugging Face11gokaygokay /prompt_description_stable_diffusion_3k The Synthetic Description from Prompts Dataset This dataset is created using the Phi 2 3B Q4_K_S quantized model, using 3k random samples from training set of a base dataset of about 80,000 prompts from the Stable Diffusion dataset on Lexica.art. This dataset is designed to explore the capabilities of language models in generating creative and expanded descriptions from concise prompts. Source Data… See the full description on the dataset page: https://huggingface.co/datasets/gokaygokay/prompt_description_stable_diffusion_3k.texttext-generation1K<n<10K5 likes39 downloads2mo agoHugging Face12agentlans /dewey-decimal-description Dewey Decimal Description Dataset (DDDD) This dataset provides third-level Dewey Decimal Classification (DDC) call numbers, each paired with a concise, one-paragraph description. Every entry explores a distinct subject area, progressing from broad categories to more specialized topics. The dataset is adapted from agentlans/library-classification-systems. Each call number entry features a brief summary explaining the subject, outlining its scope, and highlighting what differentiates… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/dewey-decimal-description.texttext-classification1K<n<10K0 likes35 downloads1y agoHugging Face13FourthBrainGenAI /Product-Descriptions-and-Ads Synthetic Dataset for Product Descriptions and Ads The basic process was as follows: Prompt GPT-4 to create a list of 100 sample clothing items and descriptions for those items. Split the output into desired format `{"product" : "", "description" : ""} Prompt GPT-4 to create adverts for each of the 100 samples based on their name and description. This data was not cleaned or verified manually. texttext-generationn<1K7 likes32 downloads3y agoHugging Face14QileXu /OmniObject3D_brief_description_val_GT PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-Thought Chaoqi Chen¹*, Qile Xu¹*, Wenjun Zhou¹, Hui Huang¹† ¹Shenzhen University &nbsp;&nbsp; *Equal contribution &nbsp;&nbsp; †Corresponding author Paper | Project Page | Code | Collection Evaluation ground truth (5,989 samples) for the brief-description task on OmniObject3D, released with the paper PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-Thought (SIGGRAPH 2026). Used as the reference set when… See the full description on the dataset page: https://huggingface.co/datasets/QileXu/OmniObject3D_brief_description_val_GT.texttext-generation1K<n<10K0 likes30 downloads2mo agoHugging Face15Detsutut /snomed_descriptionsgated SNOMED CT Descriptions Dataset summary This dataset provides SNOMED CT® concept names drawn from the UMLS® Metathesaurus® (2025AB release), enriched with natural language descriptions generated locally by a large language model (Qwen 3.5 9B). It is intended to support biomedical NLP tasks including named entity recognition, concept normalization, entity linking, and knowledge-grounded generation. Each record pairs a SNOMED CT concept, identified by its Concept Unique… See the full description on the dataset page: https://huggingface.co/datasets/Detsutut/snomed_descriptions.texttext-generation100K<n<1M0 likes28 downloads6mo agoHugging Face16gigant /ted_descriptions Dataset Card for TED descriptions More Information needed texttext-generation1K<n<10K0 likes27 downloads4y agoHugging Face17archaia /hierarchical-descriptions-sample Archaia — Hierarchical Artifact Descriptions A multimodal dataset of excavated archaeological artifacts, each paired with: Up to 3 photographs of the artifact Original catalog description written by archaeologists at excavation time 5-level AI-generated hierarchical descriptions produced by GPT-4o-mini from the photographs and structured metadata Full structured metadata (Munsell color, dimensions, material, trench coordinates, date range, etc.) Artifacts come from multiple… See the full description on the dataset page: https://huggingface.co/datasets/archaia/hierarchical-descriptions-sample.imageimage-to-textn<1K0 likes23 downloads7mo agoHugging Face18Shaer-AI-2 /ashaar-with-descriptions-baseform-final-trimmed Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens) This dataset is derived from Shaer-AI/ashaar-v1-base-form-with-descriptions and prepared for supervised fine-tuning (SFT) with the locked prompt from prompt_testing.ipynb. Locked Prompt SYSTEM_PROMPT أنت شاعر عربي تكتب الشعر العمودي الكلاسيكي. التزم بالبحر المحدد في كل شطر، واستلهم من الموضوع دون نقله حرفياً. أخرج الأبيات فقط دون مقدمة أو تعليق. USER_TEMPLATE البحر الأساسي:… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed.tabulartext-generation100K<n<1M0 likes16 downloads7mo agoHugging Face19Shaer-AI-2 /ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر) This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20 and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset. The only additional change is the removal of rows where: base_meter == "الوافر" form == "مجزوء" This removes poems labeled مجزوء الوافر from the published phase-1 subset. Why… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer.tabulartext-generation100K<n<1M0 likes15 downloads6mo agoHugging Face20davanstrien /dataset-viewer-descriptions-processedtexttext-generation1K<n<10K0 likes13 downloads2y agoHugging Face21Chaitanya2003 /Product-Descriptions-and-Ads Synthetic Dataset for Product Descriptions and Ads The basic process was as follows: Prompt GPT-4 to create a list of 100 sample clothing items and descriptions for those items. Split the output into desired format `{"product" : "", "description" : ""} Prompt GPT-4 to create adverts for each of the 100 samples based on their name and description. This data was not cleaned or verified manually. texttext-generationn<1K0 likes9 downloads2mo agoHugging Face22Shaer-AI-2 /ashaar-with-descriptions-baseform-final-trimmed-maxlen20 Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts) This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed and prepared as a phase-1 GRPO subset by applying a maximum poem length filter of 20 complete bayts while keeping the same columns, prompt format, and general dataset structure as the source dataset. Locked Prompt SYSTEM_PROMPT أنت شاعر عربي تكتب الشعر العمودي الكلاسيكي. التزم بالبحر المحدد في كل… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20.tabulartext-generation100K<n<1M0 likes8 downloads6mo agoHugging Face23Rahul5262 /ecommerce-product-descriptions-v1 E-Commerce Product Descriptions Dataset This dataset contains 150 high-quality synthetic product descriptions across multiple e-commerce categories including earbuds, smartwatches, and laptops. Dataset Structure Each entry follows the Alpaca-style instruction format: : Task description cmd: Failure calling service input: Failed transaction (2147483646): Context/brand features : Generated product description : Product category : Uniqueness ratio (0.7 - 0.99)… See the full description on the dataset page: https://huggingface.co/datasets/Rahul5262/ecommerce-product-descriptions-v1.texttext-generationn<1K0 likes7 downloads4mo agoHugging Face24Shaer-AI-2 /ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر, drop المتدارك) This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset. The only additional change is the removal of rows where: base_meter == "المتدارك" This removes all poems whose base meter is المتدارك from the published… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak.tabulartext-generation100K<n<1M0 likes6 downloads6mo agoHugging Face25ClarusC64 /description-integrity-v0.1Description Integrity v0.1 What this dataset tests This dataset evaluates whether a language model can describe what is explicitly stated without drifting into explanation, inference, or speculation. It is not a knowledge test. It is a boundary-control test. The core question is simple. Can the model report observations without inventing reasons for them? Why this matters Many high-severity failures begin with a small violation. • Description becomes explanation • Explanation introduces… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/description-integrity-v0.1.texttext-generationn<1K0 likes1 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.