CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-esa-geospatial /Llama3-SSL4EO-S12-v1.1-captions Llama3-SSL4EO-S12-Captions The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model. Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper. Code: https://github.com/IBM/MS-CLIP Data Structure We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.tabularzero-shot-image-classification100K<n<1M5 likes1.3k downloads1y agoHugging Face02BootsofLagrangian /danbooru-multitier-captions-202606 Danbooru — multi-tier captions (202606) Per-post Danbooru data for the 202606 crawl: native tags, the raw API metadata, model-generated multi-tier natural-language captions (long / refined long / medium / short), and post flags. One row per Danbooru post_id. Images are not included — each post is referenced by post_id, danbooru_url, md5, and the Danbooru CDN URLs. (The two example previews below are downscaled for illustration.) Based on:… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/danbooru-multitier-captions-202606.imagetext-to-image10M<n<100M2 likes760 downloads1mo agoHugging Face03CaptionEmporium /conceptual-captions-cc12m-llavanext Dataset Card for conceptual-captions-cc12m-llavanext Dataset Summary This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B. Languages The captions are in English. Data Instances An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.imagetext-to-image10M<n<100M28 likes380 downloads2y agoHugging Face04bdsqlsz /Tom_and_Jerry_captions3SRC46HSXPXFC63X34QGLDUCIXGZNOXR tabularn<1K2 likes372 downloads2y agoHugging Face05yauheniya-adesso /icongenai-svg-captions IconGenAI SVG Captions Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models. Part of the IconGenAI research project. Files Two files are provided at different stages of the processing pipeline: File Records Purpose icons_captioned_merged.jsonl 275,912 Full license-filtered corpus with VLM-generated captions and collection metadata icons_training_captioned.jsonl227,821 Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.tabulartext-to-image100K<n<1M2 likes366 downloads5mo agoHugging Face06LGirrbach /laion400m-person-captionstabular100M<n<1B0 likes272 downloads1y agoHugging Face07svjack /conceptual_captions_3m_zh_tiny_0 Dataset Card for "conceptual_captions_3m_zh_tiny_0" More Information needed image10K<n<100K0 likes140 downloads4y agoHugging Face08svjack /conceptual_captions_3m_zh_tiny_5 Dataset Card for "conceptual_captions_3m_zh_tiny_5" More Information needed image10K<n<100K0 likes96 downloads4y agoHugging Face09saifkhichi96 /mpii-human-pose-captions Dataset Card for MPII Human Pose Descriptions Dataset Summary The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations. These annotations are generated by various state-of-the-art language models (LLMs) and include detailed descriptions of the activities being performed, the count of people present, and their specific poses. The dataset consists of the same image splits as provided in MMPose, with 14644… See the full description on the dataset page: https://huggingface.co/datasets/saifkhichi96/mpii-human-pose-captions.tabularzero-shot-classification10K<n<100K3 likes95 downloads2y agoHugging Face10svjack /conceptual_captions_3m_zh_tiny_4 Dataset Card for "conceptual_captions_3m_zh_tiny_4" More Information needed tabular10K<n<100K0 likes92 downloads4y agoHugging Face11svjack /conceptual_captions_3m_zh_tiny_2 Dataset Card for "conceptual_captions_3m_zh_tiny_2" More Information needed image10K<n<100K0 likes91 downloads4y agoHugging Face12svjack /conceptual_captions_3m_zh_tiny_1 Dataset Card for "conceptual_captions_3m_zh_tiny_1" More Information needed tabular10K<n<100K0 likes86 downloads4y agoHugging Face13bpiyush /howto100m_captions_with_verb_nounstabular10M<n<100M1 likes78 downloads2y agoHugging Face14svjack /conceptual_captions_3m_zh_tiny_3 Dataset Card for "conceptual_captions_3m_zh_tiny_3" More Information needed tabular10K<n<100K0 likes73 downloads4y agoHugging Face15kaupane /pexels-people-captions Pexels people captions 37,412 photographs of people, each with a description written for it in Chinese or English. The photographs themselves are not included: every row carries a link to the original on Pexels instead. What a row holds column meaning image_id stable identifier used across the corpus, pexels-<photo id> photo_id the Pexels photo id page_url the photo's page on pexels.com image_url the original image file as the API reports it… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/pexels-people-captions.imagetext-to-image10K<n<100K0 likes53 downloads12d agoHugging Face16AbstractPhil /ffhq_with_llava_shorter_captions_flux_latentstabular10K<n<100K0 likes51 downloads8mo agoHugging Face17SilentAntagonist /vintage-photography-450k-high-quality-captionsThis is a 450k image datastet focused on photography from the 20th century, and their analog aspect. Many of the images are in high resolution. This dataset currently has 20k images captioned with InternVL2 26B, and is a work in progress (I plan to caption the entire dataset and also have short captions for all of the images, compute is an issue for now). imageimage-classification100K<n<1M36 likes44 downloads2y agoHugging Face18helena-balabin /vg-captions-graphs-processed-image-graphsimage1K<n<10K0 likes44 downloads11mo agoHugging Face19qingy2024 /ActivityNet-Captions ActivityNet-Captions This dataset repo contains a curated subset of the ActivityNet-Captions dataset. Filtering Logic {Video length ≤ 10 seconds:KeepVideo length > 10 seconds:Discard\text{Filtering Logic }\begin{cases} \text{Video length }\leq \text{ 10 seconds:} & \text{Keep} \\ \text{Video length } > \text{ 10 seconds:} & \text{Discard} \end{cases}Filtering Logic {Video length ≤ 10 seconds:Video length > 10 seconds:​KeepDiscard​ Analysis There are 10,759 unique… See the full description on the dataset page: https://huggingface.co/datasets/qingy2024/ActivityNet-Captions.tabular10K<n<100K0 likes41 downloads1y agoHugging Face20siyrus /BToks-ActivityNet-Captionstabular100K<n<1M0 likes41 downloads3mo agoHugging Face21LimYeri /leetcode_with_youtube_captionsimagetext-classification10K<n<100K1 likes37 downloads2y agoHugging Face22kubernetes-bad /character-captions-opusDeduplicated set of character portraits that have been described by Anthropic Claude Opus as characters with stories and visual attributes. Images obtained from CivitAI by filtering for SD XL-derived models only. Original Stable Diffusion prompt and metadata is also included. Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided. Character-like description for each image is given by Claude Opus. Here is an example: {… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/character-captions-opus.image10K<n<100K5 likes31 downloads2y agoHugging Face23dragoon49 /florence2-ofa-captions-500 OFA Florence-2 Dataset (500 Samples) This dataset was generated using microsoft/Florence-2-large on a subset of COCO 2017 Validation images. It is pre-formatted for OFA Stage-1 fine-tuning (Headerless TSV, URL-safe base64, max 512x512 resolution). tabularn<1K0 likes30 downloads2mo agoHugging Face24kaupane /vintage-photography-captions Dataset Card for Vintage Photograph Captions Recaption This dataset contains 445,271 recaptioned vintage photographs, derived from the vintage-photography-450k-high-quality-captions dataset. It provides high-quality bilingual (English and Chinese) captions, aesthetic scores, and other metadata generated using the Qwen2-VL model. This dataset is a recaptioned version of SilentAntagonist/vintage-photography-450k-high-quality-captions. The original dataset contained 456,006… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/vintage-photography-captions.imagetext-to-image100K<n<1M2 likes23 downloads9mo agoHugging Face25nielsr /datacomp_small_english_captions Dataset Card for "datacomp_small_english_captions" More Information needed image1M<n<10M0 likes22 downloads3y agoHugging Face26fschieber /wit-captionsDataset derived from the original WIT dataset, with the following changes: Removed all columns except for 'image_url', 'caption_reference_description', 'caption_attribution_description', 'mime_type', 'original_height', 'original_width' All rows without either a 'caption_reference_description' or 'caption_attribution_description' have been removed The image references were deduplicated on the image_url preserving the entry with the longest caption_reference_description A new 'text' column was… See the full description on the dataset page: https://huggingface.co/datasets/fschieber/wit-captions.image10M<n<100M1 likes22 downloads3y agoHugging Face27tarekziade /coco-transformed-captionsimagen<1K0 likes22 downloads2y agoHugging Face28opendiffusionai /laion2b-aesthetic-squareish-captionsThis dataset contains image captions generated from LAION2B-en-aesthetic-square. We started with ~300K images after size filtering (2.5k max w/h), a portion of the images were skipped due to inaccessible URLs. The captions were generated over ~30 hours using Qwen3-VL-30B-A3B-Instruct on 1xH100 running SGLang with the prompt Describe the content of the provided image in detail, in plaintext. Do not make assumptions. Do not use special formatting. Avoid purple prose. Total samples: 209,141… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-aesthetic-squareish-captions.image100K<n<1M1 likes22 downloads10mo agoHugging Face29nielsr /datacomp_small_english_captions_without_weird_characters Dataset Card for "datacomp_small_english_captions_without_weird_characters" More Information needed image1M<n<10M0 likes20 downloads3y agoHugging Face30nielsr /datacomp_small_french_captions Dataset Card for "datacomp_small_french_captions" More Information needed image100K<n<1M0 likes16 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.