CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gmongaras /Imagenet21K_RecaptionThis dataset is the entire 21K ImageNet dataset with about 13 million examples and about 19 thousand classes as strings (for some reason it only had ~19K classes instead of 21K). If you want an even larger set of images, I have a recaptioned CC12M and ImageNet dataset: https://huggingface.co/datasets/gmongaras/CC12M_and_Imagenet21K_Recap The images are in PNG format. They can be decoded like in the following example import io from PIL import Image Image.open(io.BytesIO(row["image"])) where… See the full description on the dataset page: https://huggingface.co/datasets/gmongaras/Imagenet21K_Recaption.image10M<n<100M11 likes4.4k downloads1y agoHugging Face02ooutlierr /cc12m-recaptionedimage10M<n<100M8 likes2.3k downloads2y agoHugging Face03undefined443 /cc12m-wds-coco-recaptioned CC12M WebDataset with COCO-style Recaptions A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL. Dataset Overview Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M) Images: 3,000,000+ high-quality internet images Recaption Model: NVIDIA Nemotron Nano 12B v2 VL Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.image1M<n<10M1 likes1.4k downloads6mo agoHugging Face04gmongaras /Stable_Diffusion_3_RecaptionThis dataset is the one specified in the stable diffusion 3 paper which is composed of the ImageNet dataset and the CC12M dataset. I used the ImageNet 2012 train/val data and captioned it as specified in the paper: "a photo of a 〈class name〉" (note all ids are 999,999,999) CC12M is a dataset with 12 million images created in 2021. Unfortunately the downloader provided by Google has many broken links and the download takes forever. However, some people in the community publicized the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/gmongaras/Stable_Diffusion_3_Recaption.image10M<n<100M5 likes1.4k downloads2y agoHugging Face05AterMors /wikiart_recaptionWikiArt Dataset captioned using vikhyatk/moondream2 model with prompt : Generate a short, simple and only visually descriptive caption for this image. imageimage-to-text10K<n<100K9 likes622 downloads2y agoHugging Face06Panorama-grounding /COCONut-PanCap-Recaptioned COCONut-PanCap Re-captioned (train2017) Regenerated captions for the train2017 split of COCONut-PanCap (paper, dataset). The released captions contain phrases bound to the wrong segment, references to objects that are not in the image, redundant restatements, and untagged summary sentences that carry no grounding. We regenerated every caption and repaired the remaining errors, keeping the file layout of the original release. Only the caption text changed. Mask ids… See the full description on the dataset page: https://huggingface.co/datasets/Panorama-grounding/COCONut-PanCap-Recaptioned.textimage-segmentation100K<n<1M3 likes256 downloads8d agoHugging Face07askoepke /wit_1m_recaptioned Back into Plato’s Cave: Examining Cross-modal Representational Convergence at Scale Image–text dataset derived from Wikipedia-based Image Text (WIT) with original and Gemini-generated captions, introduced in Back into Plato’s Cave: Examining Cross-modal Representational Convergence at Scale. Configs wit_1024 A fixed set of 1,024 query samples used for alignment evaluation. from datasets import load_dataset ds = load_dataset("askoepke/wit_1m_recaptioned"… See the full description on the dataset page: https://huggingface.co/datasets/askoepke/wit_1m_recaptioned.imageimage-to-text1M<n<10M0 likes224 downloads5mo agoHugging Face08wchai /AuroraCap-recaption AuroraCap-recaption Resources Website arXiv: Paper GitHub: Code Huggingface: AuroraCap Model Huggingface: VDC Benchmark Huggingface: Trainset Features Video recaption data by AuroraCap. Continue updating... For some video source, we could upload the raw videos but for the others we could only provide the url since the well-known reason. Citation @article{chai2024auroracap, title={AuroraCap: Efficient, Performant Video Detailed… See the full description on the dataset page: https://huggingface.co/datasets/wchai/AuroraCap-recaption.textvisual-question-answering10K<n<100K5 likes61 downloads2y agoHugging Face09sirus /megalith-10m-5.5k-claude-opus-5-recaptioned Megalith-10M 5.5K — Claude Opus 5 Recaptioned This is a 5,511-image derivative subset of madebyollin/megalith-10m, selected through the megalith10m portion of zlab-princeton/i1-captions. The bytes were retrieved from the drawthingsai/megalith-10m image archive. It is not the complete Megalith-10M collection. Every image has one newly generated, detailed English caption. The recaptioning was performed with Claude Opus 5 via Claude Code on August 2, 2026. The image was the primary… See the full description on the dataset page: https://huggingface.co/datasets/sirus/megalith-10m-5.5k-claude-opus-5-recaptioned.imageimage-to-text1K<n<10K0 likes53 downloads2mo agoHugging Face10sirus /inaturalist-2024-2.8k-claude-opus-5-recaptioned iNaturalist 2024 2.8K — Claude Opus 5 Recaptioned This is a 2,824-image derivative subset of iNaturalist 2024 (iNat24), distributed through the INQUIRE project, selected through the inaturalist portion of zlab-princeton/i1-captions. It is not the complete 4.8-million-image iNat24 training set. Every image has one newly generated, detailed English caption. The recaptioning was performed with Claude Opus 5 via Claude Code on August 2, 2026. The image was the primary evidence; the… See the full description on the dataset page: https://huggingface.co/datasets/sirus/inaturalist-2024-2.8k-claude-opus-5-recaptioned.imageimage-to-text1K<n<10K0 likes38 downloads2mo agoHugging Face11undefined443 /JourneyDB-recaption JourneyDB Recaption Recaptioned version of the JourneyDB dataset using Qwen vision-language models. Dataset Description JourneyDB is a large-scale dataset of AI-generated images from Midjourney. This recaptioned version provides detailed visual descriptions generated by a vision-language model, which are more accurate than the original generation prompts for describing actual image content. Statistics Metric Count Total rows 3,389,605 File size… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/JourneyDB-recaption.tabularimage-to-text1M<n<10M0 likes36 downloads6mo agoHugging Face12toilaluan /JourneyDB-recaption JourneyDB Recaption Recaptioned version of the JourneyDB dataset using Qwen vision-language models. Dataset Description JourneyDB is a large-scale dataset of AI-generated images from Midjourney. This recaptioned version provides detailed visual descriptions generated by a vision-language model, which are more accurate than the original generation prompts for describing actual image content. Statistics Metric Count Total rows 3,389,605… See the full description on the dataset page: https://huggingface.co/datasets/toilaluan/JourneyDB-recaption.tabularimage-to-text1M<n<10M1 likes35 downloads10d agoHugging Face13kaupane /human-recaption Dataset Card for Human Recaption This dataset contains 240,146 recaptioned images focusing on human subjects, derived from the HumanCaption-HQ-311K dataset. It provides high-quality bilingual (English and Chinese) captions, aesthetic scores, and other metadata generated using the Qwen2-VL model. This dataset is a recaptioned version of OpenFace-CQUPT/HumanCaption-HQ-311K. The original dataset contained 313,482 samples. This version contains 240,146 samples; the reduction is… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/human-recaption.imagetext-to-image100K<n<1M2 likes32 downloads10mo agoHugging Face14catslashbin /datikz-v3-recaptioned DaTikZ-v3 Recaptioned A 10,000-sample subset of DaTikZ-v3 with captions regenerated using Gemini Flash via vision-language captioning. Captioning Original captions were replaced by passing each rendered diagram image to Gemini Flash with the prompt: describe the diagram concisely, focusing on geometric shapes, mathematical concepts, key visual elements, and purpose (~1–3 sentences starting with "A diagram ..."). Only samples with TikZ code shorter than 1000 characters… See the full description on the dataset page: https://huggingface.co/datasets/catslashbin/datikz-v3-recaptioned.imagetext-to-image10K<n<100K0 likes29 downloads7mo agoHugging Face15undefined443 /cc12m-wds-recaption CC12M with Enhanced Captions This dataset contains 1.3 million image-text pairs from the CC12M dataset with model-generated captions. Dataset Details Total Samples: 1,306,239 Source: pixparse/cc12m-wds Captioning Model: Qwen/Qwen3-VL-8B-Instruct Format: Parquet Filtering Criteria Samples were filtered based on the following quality metrics: Aesthetic Score: >= 5.5 (using LAION aesthetic classifier) Resolution: >= 512 pixels (width or height) Aspect Ratio: <=… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-recaption.tabularimage-to-text1M<n<10M0 likes23 downloads6mo agoHugging Face16ljnlonoljpiljm /laion-unsafe-downloaded-recaptionedimage10K<n<100K1 likes20 downloads2y agoHugging Face17SGP-Team-B /laion-mix-updated-recaptioned-yolo-filtered-2image10K<n<100K0 likes18 downloads2y agoHugging Face18undefined443 /LAION-Art-recaption LAION-Art Recaption Recaptioned subset of the LAION-Art dataset using Qwen2.5-VL-7B-Instruct. Dataset Description This dataset contains detailed recaptions for LAION-Art images generated by a vision-language model. Only successfully recaptioned samples are included. Total samples: 1,410,704 Columns Column Type Description url string Original image URL key string Unique identifier for each sample width int32 Image width height int32 Image height… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/LAION-Art-recaption.imageimage-to-text1M<n<10M0 likes18 downloads6mo agoHugging Face19Fizzarolli /filtered-wit-recaptionedimagen<1K0 likes16 downloads2y agoHugging Face20LGirrbach /cc3m-recaptionedtext1M<n<10M0 likes15 downloads1y agoHugging Face21svjack /human-control-union-recaptionimage1K<n<10K0 likes13 downloads8mo agoHugging Face22LGirrbach /cc12m-recaptionedtext10M<n<100M0 likes9 downloads1y agoHugging Face23kingsidharth /zangei-dit-stage-1-250k-recaptionedtabular10K<n<100K0 likes9 downloads2mo agoHugging Face24Nano1337 /visualize-recaptioned-test1 Dataset Card for "visualize-recaptioned-test1" More Information needed imagen<1K0 likes8 downloads2y agoHugging Face25Nano1337 /visualize-paligemma-recaptioned Dataset Card for "visualize-paligemma-recaptioned" More Information needed imagen<1K0 likes8 downloads2y agoHugging Face26LGirrbach /imagenet-recaptionedtext1M<n<10M0 likes7 downloads1y agoHugging Face27kingsidharth /zangei-dit-stage-1-250k-recaptioned-qwen-3-0.6b-embeddingstabular10K<n<100K0 likes7 downloads2mo agoHugging Face28LGirrbach /yfcc-recaptionedtext10M<n<100M0 likes5 downloads1y agoHugging Face29weathon /merged_aa_recaptionedimage1K<n<10K0 likes3 downloads5mo agoHugging Face30kaki-paper /llava-onevision-recaption-ko-previewgatedimagen<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.