CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes6.5k downloads5y agoHugging Face02BLIP3o /BLIP3o-Pretrain-Long-Caption BLIP3o Pretrain Long-Caption Dataset This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.image10M<n<100M74 likes6.1k downloads1y agoHugging Face03BLIP3o /BLIP3o-Pretrain-Short-Caption BLIP3o Pretrain Short-Caption Dataset This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.image1M<n<10M10 likes5.6k downloads1y agoHugging Face04laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes1.9k downloads11mo agoHugging Face05TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes1.8k downloads6mo agoHugging Face06hanlincs /InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption image10M<n<100M1 likes1.4k downloads1y agoHugging Face07clip-benchmark /wds_mscoco_captionsimage10K<n<100K4 likes573 downloads4y agoHugging Face08JWRoboticsVision /HO-Cap-Datasetimagen<1K0 likes544 downloads1y agoHugging Face09mispeech /MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks 📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF) Dataset Description MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks: Audio Captioning: Generating textual descriptions for given audio Audio Question Answering: Answering questions about given audio Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.audioaudio-classification10K<n<100K4 likes328 downloads5mo agoHugging Face10clip-benchmark /wds_mscoco_captions2017image10K<n<100K8 likes219 downloads3y agoHugging Face11laion /timbre-audio-caption-pairsaudio100K<n<1M2 likes206 downloads9mo agoHugging Face12wendlerc /CaptionedSynthTextThis dataset has been created by Stability AI and LAION. SynthText is a popular OCR dataset, where random texts are rendered into random locations in images based on depth maps. In this dataset, we additionally computed image captions using BLIP2. Caption: "a close up of a leopard's face with a blurry background" image100K<n<1M2 likes199 downloads3y agoHugging Face13laion /freesound-commercially-permissive-subset-with-captionsaudio100K<n<1M2 likes194 downloads11mo agoHugging Face14data-archetype /ffhq_captioned_1024 ffhq_captioned_1024 A captioned bucketed-shards export of gaunernst/ffhq-1024-wds. This export contains 70,000 square face and portrait images from FFHQ, stored as JPEG TAR shards in a single 1024 x 1024 bucket. The source images are decoded from the original dataset, deterministically converted to RGB, and re-encoded as high-quality JPEG (quality=95, adaptive subsampling). Captions were generated with a Gemini 2.5 Flash Lite primary pass and a Mistral Medium 3.1 fallback. Intended… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/ffhq_captioned_1024.imagetext-to-image10K<n<100K0 likes155 downloads5mo agoHugging Face15DamianBoborzi /objaverse_processed_renders_and_captionsContains rendered views and captions from Objaverse XL objects. the objects are from the alignment and TRELLIS500K (over 1 Millionen processed objects) dataset. We downloaded and rendered 4 views of each object. We added TRELLIS and CAP3D Captions where available. If there were no captions we generated new captions with the large version of Florence 2. This is the base dataset we used to generate MeshFleet which is described in MeshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain… See the full description on the dataset page: https://huggingface.co/datasets/DamianBoborzi/objaverse_processed_renders_and_captions.imageimage-to-text1M<n<10M0 likes149 downloads1y agoHugging Face16mitermix /audioset-with-grounded-captionsaudio1M<n<10M4 likes139 downloads1y agoHugging Face17hidude562 /operation-cappicinoaudio10K<n<100K0 likes130 downloads5mo agoHugging Face18PixArt-alpha /SAM-LLaVA-Captions10Mtext10M<n<100M64 likes121 downloads3y agoHugging Face19hmu013 /SynRIS-captionedimage10K<n<100K0 likes120 downloads8mo agoHugging Face20ChristophSchuhmann /wikiart_with_BLIP_captionstext10K<n<100K2 likes114 downloads4y agoHugging Face21thisnick /nsfw-video-still-caption-grid-onlyimage10K<n<100K14 likes91 downloads2y agoHugging Face22getbetterhyccc /CAPSUL CAPSUL: A Comprehensive Human Protein Benchmark for Subcellular Localization 📊 Dataset Specifications This repository contains the complete dataset of the paper: "CAPSUL: A Comprehensive Human Protein Benchmark for Subcellular Localization" Accepted by ICLR 2026. Here we introduce the protein dataset used in our CAPSUL benchmark evaluation, with comprehensive 3D information and fine-grained localization annotations. Data is derived from AlphaFold2, UniProt, and the Human… See the full description on the dataset page: https://huggingface.co/datasets/getbetterhyccc/CAPSUL.text10K<n<100K0 likes83 downloads6mo agoHugging Face23TTS-AGI /majestrino-unified-detailed-captions-temporal Majestrino Unified Detailed Captions with Temporal Aspects Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects. Stats 4,128,665 samples 826 tar files (~1.1 GB each) ~878 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption with temporal aspects caption_type — always unified_detailed_caption_with_temporal_aspects transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.audioaudio-classification1M<n<10M0 likes81 downloads6mo agoHugging Face24ngqtrung /full-modality-video-caption Full Modality Video Caption Dataset A large-scale multimodal video dataset with comprehensive vision, audio, and integrated captions. Dataset Description This dataset contains 55,940 video segments (10 seconds each) with three types of captions: Vision Caption: Visual description generated by GPT-4o Audio Caption: Audio/speech description generated by Qwen3-Omni-30B-A3B-Captioner Video Caption: Integrated multi-modal description combining vision and audio generated by… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/full-modality-video-caption.text10K<n<100K0 likes79 downloads11mo agoHugging Face25laion /audioset-with-captionsaudio1M<n<10M2 likes72 downloads11mo agoHugging Face26saehyungl /CapMAS [ICML 2025] Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage This dataset is associated with the evaluation in our ICML 2025 paper, Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage. Prerequisites Packages openai>=1.14.1 python-dotenv==1.0.1 Dataset download from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/saehyungl/CapMAS.image1K<n<10K0 likes56 downloads1y agoHugging Face27Salmonnn /InternVL-SA-1B-Caption-512image10M<n<100M0 likes53 downloads1y agoHugging Face28CaptionEmporium /dalle3-llama3.2-11b Dataset Card for dalle3-llama3.2-11b Dataset Summary This is 3,577,716 new synthetic captions for the 1,192,572 images found in ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions. The dataset was filtered for duplicates and then re-encoded with JPEGXL lossless or lossy depending on the source. The long captions were produced using meta-llama/Llama-3.2-11B-Vision-Instruct. Medium and short captions were produced from these captions using… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/dalle3-llama3.2-11b.texttext-to-image1M<n<10M0 likes45 downloads2y agoHugging Face29nightknocker /2025-captionstext1M<n<10M0 likes42 downloads7mo agoHugging Face30lingcarzy /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M0 likes39 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.