CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01friedrichor /ActivityNet_Captions About ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions) Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.texttext-to-video10K<n<100K15 likes1.6k downloads1y agoHugging Face02AVoCaDO-Captioner /training_settext100K<n<1M1 likes1.1k downloads6mo agoHugging Face03AudioVisual-Caption /ASID-1M ASID-1M: Attribute-Structured and Quality-Verified Audiovisual Instructions [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce ASID-1M, a large-scale audiovisual instruction dataset built to support universal video understanding with fine-grained, controllable supervision. Most existing video-instruction data represents complex audiovisual content as a single, monolithic caption. This often leads to incomplete coverage (missing audio… See the full description on the dataset page: https://huggingface.co/datasets/AudioVisual-Caption/ASID-1M.textimage-text-to-text100K<n<1M85 likes1.1k downloads7mo agoHugging Face04Obscure-Entropy /conceptual_captions_jsonimage1M<n<10M0 likes806 downloads2y agoHugging Face05wchai /Video-Detailed-Caption Video Detailed Caption Benchmark Resources Website arXiv: Paper GitHub: Code Huggingface: AuroraCap Model Huggingface: VDC Benchmark Huggingface: Trainset Features Benchmark Collection and Processing We building VDC upon Panda-70M, Ego4D, Mixkit, Pixabay, and Pexels. Structured detailed captions construction pipeline. We develop a structured detailed captions construction pipeline to generate extra detailed descriptions from various… See the full description on the dataset page: https://huggingface.co/datasets/wchai/Video-Detailed-Caption.textvideo-text-to-text1K<n<10K17 likes754 downloads2y agoHugging Face06neurlang /Minecraft-Skins-Captioned-1M Dataset Card for Minecraft Skins Dataset Summary This dataset contains 981,079 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier. Dataset Structure Data Fields This dataset includes the following fields: hash: A data dependent hash. These hashes are generated from raw bytes and will be same if the skin is identical. image: The skin image encoded in base64 format.… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/Minecraft-Skins-Captioned-1M.textimage-classification1M<n<10M7 likes690 downloads11mo agoHugging Face07CaptionEmporium /pexels-568k-internvl2 Dataset Card for pexels-568k-internvl2 Dataset Summary This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed. Languages The text is in English, but occasionally text in images in other languages is transcribed. Intended Usage Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.imagetext-to-image100K<n<1M21 likes611 downloads2y agoHugging Face08CaptionEmporium /conceptual-captions-cc12m-llavanext Dataset Card for conceptual-captions-cc12m-llavanext Dataset Summary This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B. Languages The captions are in English. Data Instances An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.imagetext-to-image10M<n<100M28 likes380 downloads2y agoHugging Face09yauheniya-adesso /icongenai-svg-captions IconGenAI SVG Captions Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models. Part of the IconGenAI research project. Files Two files are provided at different stages of the processing pipeline: File Records Purpose icons_captioned_merged.jsonl 275,912 Full license-filtered corpus with VLM-generated captions and collection metadata icons_training_captioned.jsonl227,821 Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.tabulartext-to-image100K<n<1M2 likes366 downloads5mo agoHugging Face10DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes312 downloads2y agoHugging Face11vvwangvv /emilia-captions-v3text10M<n<100M0 likes260 downloads8mo agoHugging Face12OpenGVLab /InternVL-SA-1B-Caption Dataset Card for InternVL-SA-1B-Caption Overview The InternVL-SA-1B-Caption Dataset is a bilingual dataset created using the InternVL2-Llama3-76B model. The dataset contains 12 million image-caption pairs in both English and Chinese. All images are sourced from Meta’s SA-1B dataset, and captions were generated using specific prompts designed to minimize hallucinations and ensure accurate descriptions based on visible image content. The dataset is intended for use in tasks… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption.tabular1M<n<10M24 likes170 downloads2y agoHugging Face13embedding-data /coco_captions_quintets Dataset Card for "coco_captions" Dataset Summary COCO is a large-scale object detection, segmentation, and captioning dataset. This repo contains five captions per image; useful for sentence similarity tasks. Disclaimer: The team releasing COCO did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team. Supported Tasks Sentence Transformers training; useful for semantic search and sentence… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/coco_captions_quintets.textsentence-similarity10K<n<100K6 likes165 downloads4y agoHugging Face14CaptionEmporium /anime-caption-danbooru-2021-sfw-5m-hq Dataset Card for anime-caption-danbooru-2021-sfw-5m-hq Dataset Summary This is 5.71 M captions of 1.43 M images from a safe-for-work (SFW) filtered subset of the Danbooru 2021 dataset. There are 4 captions per image: 1 by CogVLM, 1 by llava-v1.6-34b, 1 llava-v1.6-34b cleaned, and 1 llava-v1.6-34b shortened. See the sections below for how they were generated. Most captions are substantially larger than 77 tokens and are unsuitable for discrimination using current… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/anime-caption-danbooru-2021-sfw-5m-hq.textimage-to-text1M<n<10M29 likes152 downloads2y agoHugging Face15Senqiao /LiDAR-LLM-Nu-Caption Dataset Details Dataset type: This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset. Dataset keys: "answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data. If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.textquestion-answering100K<n<1M8 likes141 downloads2y agoHugging Face16alinasdkey /graph-captioning-train-onlyimagen<1K0 likes133 downloads1y agoHugging Face17saifkhichi96 /mpii-human-pose-captions Dataset Card for MPII Human Pose Descriptions Dataset Summary The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations. These annotations are generated by various state-of-the-art language models (LLMs) and include detailed descriptions of the activities being performed, the count of people present, and their specific poses. The dataset consists of the same image splits as provided in MMPose, with 14644… See the full description on the dataset page: https://huggingface.co/datasets/saifkhichi96/mpii-human-pose-captions.tabularzero-shot-classification10K<n<100K3 likes95 downloads2y agoHugging Face18embedding-data /flickr30k_captions_quintets Dataset Card for "flickr30k-captions" Dataset Summary We propose to use the visual denotations of linguistic expressions (i.e. the set of images they describe) to define novel denotational similarity metrics, which we show to be at least as beneficial as distributional similarities for two tasks that require semantic inference. To compute these denotational similarities, we construct a denotation graph, i.e. a subsumption hierarchy over constituents and their denotations… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/flickr30k_captions_quintets.text10K<n<100K4 likes89 downloads4y agoHugging Face19Sreevardhan1729 /ActivityNet_Captions About ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions) Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/Sreevardhan1729/ActivityNet_Captions.texttext-to-video10K<n<100K0 likes79 downloads6mo agoHugging Face20Waterfront /social-media-captions Social Media Captions Based on the Instagram Influencer Dataset from Seungbae Kim, Jyun-Yu Jiang, and Wei Wang Extended with photo descriptions of ydshieh/vit-gpt2-coco-en model to create a dataset which can be used to finetune Llama-2. 20k smaller subset: Waterfront/social-media-captions-20k 10k smaller subset: Waterfront/social-media-captions-10k text10K<n<100K3 likes72 downloads3y agoHugging Face21junwann /CSFM-ImageNet1K-Caption CSFM-ImageNet1K-Caption Dataset Project Page | Paper | Code This repository contains dataset associated with the paper "Better Source Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching". This dataset is used for training and evaluating Condition-dependent Source Flow Matching (CSFM), a framework that learns condition-dependent source distributions for flow matching. We recaptioned the ImageNet-1K dataset using Qwen3-VL-8B Instruct, resulting in detailed… See the full description on the dataset page: https://huggingface.co/datasets/junwann/CSFM-ImageNet1K-Caption.texttext-to-image1M<n<10M3 likes72 downloads8mo agoHugging Face22CaptionEmporium /furry-e621-sfw-7m-hq Dataset Card for furry-e621-sfw-7m-hq Dataset Summary This is 6.92 M captions of the images from the safe-for-work (SFW) split of e621 ("e926"). It extends to January 2023, before the widespread advent of machine learning images. It includes captions created by LLMs and a custom multilabel classifier along with CogVLM. There are 8 LLM (mistralai/Mistral-7B-v0.1) and 1 CogVLM (THUDM/CogVLM) captions per image. Most captions are substantially larger than 77 tokens and are… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/furry-e621-sfw-7m-hq.textimage-to-text100K<n<1M7 likes54 downloads3y agoHugging Face23SOTAagi2030 /Harbor-Photo-Captions Harbor Photo Captions This collection contains caption records for digitized waterfront photographs. Rights register Accession: HP-19 Rights statement: Public Domain Mark 1.0 Depositing archive: Tideglass Image Library Caption normalization follows the catalog's controlled vocabulary. textn<1K0 likes51 downloads12d agoHugging Face24zl2048 /SC-Captioner-data Training and testing annotations for SC-Captioner. All files are processed into llamafactory data format. In train_coco6k.json, the "rejected" line means original gpt4 captions in RefinedCaps. They are not used in our self-correction training, but can be used for other purposes. text10K<n<100K0 likes48 downloads1y agoHugging Face25HamGangster /coco_2017_caption_traintextn<1K0 likes46 downloads3y agoHugging Face26Waterfront /social-media-captions-10k Social Media Captions Based on the Instagram Influencer Dataset from Seungbae Kim, Jyun-Yu Jiang, and Wei Wang Extended with photo descriptions of ydshieh/vit-gpt2-coco-en model to create a dataset which can be used to finetune Llama-2. 60k complete dataset: Waterfront/social-media-captions 20k bigger subset: Waterfront/social-media-captions-20k text1K<n<10K3 likes42 downloads3y agoHugging Face27bjoernp /mistral_captionstext1M<n<10M1 likes37 downloads3y agoHugging Face28lingamvamshikrishnareddy /ramanv-image-captions-6gatedtext1K<n<10K0 likes34 downloads22d agoHugging Face29CaptionEmporium /TextOCR-GPT4o Dataset Card for TextOCR-GPT4o Dataset Summary TextOCR-GPT4o is Meta's TextOCR dataset dataset captioned with emphasis on text OCR using GPT4o. To get the image, you will need to agree to their terms of service. Supported Tasks The TextOCR-GPT4o dataset is intended for generating benchmarks for comparison of an VLM to GPT4o. Languages The caption languages are in English, while various texts in images are in many languages such as Spanish, Japanese… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/TextOCR-GPT4o.textimage-to-text10K<n<100K15 likes33 downloads2y agoHugging Face30NetherlandsForensicInstitute /flickr30k-captions-translated-nlThis is a Dutch version of the Flickr30k captions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. For more information about the use of this dataset please refer to the flicker terms of use textsentence-similarity100K<n<1M0 likes32 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.