CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes6.4k downloads4y agoHugging Face02BLIP3o /BLIP3o-Pretrain-Long-Caption BLIP3o Pretrain Long-Caption Dataset This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.image10M<n<100M74 likes6.3k downloads1y agoHugging Face03BLIP3o /BLIP3o-Pretrain-Short-Caption BLIP3o Pretrain Short-Caption Dataset This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.image1M<n<10M10 likes5.5k downloads1y agoHugging Face04hanlincs /InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption image10M<n<100M1 likes1.1k downloads1y agoHugging Face05clip-benchmark /wds_mscoco_captionsimage10K<n<100K4 likes616 downloads4y agoHugging Face06mispeech /MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks 📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF) Dataset Description MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks: Audio Captioning: Generating textual descriptions for given audio Audio Question Answering: Answering questions about given audio Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.audioaudio-classification10K<n<100K4 likes352 downloads5mo agoHugging Face07clip-benchmark /wds_mscoco_captions2017image10K<n<100K8 likes217 downloads3y agoHugging Face08wendlerc /CaptionedSynthTextThis dataset has been created by Stability AI and LAION. SynthText is a popular OCR dataset, where random texts are rendered into random locations in images based on depth maps. In this dataset, we additionally computed image captions using BLIP2. Caption: "a close up of a leopard's face with a blurry background" image100K<n<1M2 likes200 downloads3y agoHugging Face09DamianBoborzi /objaverse_processed_renders_and_captionsContains rendered views and captions from Objaverse XL objects. the objects are from the alignment and TRELLIS500K (over 1 Millionen processed objects) dataset. We downloaded and rendered 4 views of each object. We added TRELLIS and CAP3D Captions where available. If there were no captions we generated new captions with the large version of Florence 2. This is the base dataset we used to generate MeshFleet which is described in MeshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain… See the full description on the dataset page: https://huggingface.co/datasets/DamianBoborzi/objaverse_processed_renders_and_captions.imageimage-to-text1M<n<10M0 likes155 downloads1y agoHugging Face10data-archetype /ffhq_captioned_1024 ffhq_captioned_1024 A captioned bucketed-shards export of gaunernst/ffhq-1024-wds. This export contains 70,000 square face and portrait images from FFHQ, stored as JPEG TAR shards in a single 1024 x 1024 bucket. The source images are decoded from the original dataset, deterministically converted to RGB, and re-encoded as high-quality JPEG (quality=95, adaptive subsampling). Captions were generated with a Gemini 2.5 Flash Lite primary pass and a Mistral Medium 3.1 fallback. Intended… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/ffhq_captioned_1024.imagetext-to-image10K<n<100K0 likes155 downloads5mo agoHugging Face11hmu013 /SynRIS-captionedimage10K<n<100K0 likes120 downloads8mo agoHugging Face12thisnick /nsfw-video-still-caption-grid-onlyimage10K<n<100K14 likes90 downloads2y agoHugging Face13Salmonnn /InternVL-SA-1B-Caption-512image10M<n<100M0 likes57 downloads1y agoHugging Face14lingcarzy /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M0 likes36 downloads6mo agoHugging Face15deepghs /midjourney_captioned_23m_fullgated Midjourney Captioned Full Dataset This is the full dataset of Midjourney Captioned 23M dataset. And all the original images are maintained here. Thanks to the contribution of a certain third-party data provider who wishes to remain anonymous. Information Images There are 23167456 images in total. The maximum ID of these images is 23167456. Last updated at 2024-12-01 12:11:43 UTC. These are the information of recent 50 images: id width height filename… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.imageimage-classification10M<n<100M36 likes22 downloads2y agoHugging Face16KhangTruong /NWPU-Captionimage10K<n<100K0 likes21 downloads2y agoHugging Face17justinpinkney /pokemon-blip-captions-wdsWebdataset version of: lambdalabs/pokemon-blip-captions imagen<1K0 likes16 downloads3y agoHugging Face18thisnick /nsfw-video-still-caption-testimagen<1K2 likes16 downloads2y agoHugging Face19Dant33 /WikiArt-81K-BLIP_2-captions WikiArt Enhanced Dataset Description This dataset contains 81,444 artistic images from WikiArt, organized into different artistic genres. It has undergone several improvements and corrections to optimize its use in machine learning tasks and computational art analysis. Credits to the original author of daset go to: WikiArt Enhancements 1. Encoding Issues Correction Fixed encoding issues in filenames and artist information. All filenames were renamed… See the full description on the dataset page: https://huggingface.co/datasets/Dant33/WikiArt-81K-BLIP_2-captions.image10K<n<100K0 likes16 downloads2y agoHugging Face20shauray /aesthetic-cleaned-captionedimage1K<n<10K1 likes16 downloads4mo agoHugging Face21milosdevic /synthetic-image-caption-pairsimage1K<n<10K0 likes12 downloads2y agoHugging Face22jacklishufan /soundnet-flux-captionimage100K<n<1M0 likes10 downloads2y agoHugging Face23Khoale11hcmut /gpt4o_captions_1k5samples PACO WebDataset export for PACO-style localized caption data. Summary Samples: 1500 Shards: 2 Payload format inside each shard: pickle Split: train Config: PACO Layout Media files are stored in WebDataset tar shards. Each sample key is stable and becomes __key__ in the dataset viewer. Hugging Face will infer columns such as jpg, pickle, json, __key__, and __url__ from the shard contents. Manifest file: PACO/annotations.json image1K<n<10K1 likes5 downloads5mo agoHugging Face24AlanderK /SD-2.1_with_coco_captionimage10K<n<100K0 likes4 downloads2y agoHugging Face25simon123905 /imagenet_val_caption1image10K<n<100K0 likes4 downloads1y agoHugging Face26vickypawar123 /CaptionedSynthTextThis dataset has been created by Stability AI and LAION. SynthText is a popular OCR dataset, where random texts are rendered into random locations in images based on depth maps. In this dataset, we additionally computed image captions using BLIP2. Caption: "a close up of a leopard's face with a blurry background" image100K<n<1M0 likes1 downloads8mo agoHugging Face27sathiiii /medonethinker-source-captioningimage100K<n<1M0 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.