CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01friedrichor /ActivityNet_Captions About ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions) Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.texttext-to-video10K<n<100K15 likes2.4k downloads1y agoHugging Face02dz-osamu /M3D-Captext10K<n<100K1 likes2.2k downloads2y agoHugging Face03LDJnr /Capybara This is the Official Capybara dataset. Over 10,000 multi-turn examples. Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others. The single-turn seeds used to initiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Capybara.textquestion-answering10K<n<100K258 likes1.4k downloads2y agoHugging Face04AudioVisual-Caption /ASID-1M ASID-1M: Attribute-Structured and Quality-Verified Audiovisual Instructions [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce ASID-1M, a large-scale audiovisual instruction dataset built to support universal video understanding with fine-grained, controllable supervision. Most existing video-instruction data represents complex audiovisual content as a single, monolithic caption. This often leads to incomplete coverage (missing audio… See the full description on the dataset page: https://huggingface.co/datasets/AudioVisual-Caption/ASID-1M.textimage-text-to-text100K<n<1M85 likes1.3k downloads7mo agoHugging Face05CaptionEmporium /pexels-568k-internvl2 Dataset Card for pexels-568k-internvl2 Dataset Summary This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed. Languages The text is in English, but occasionally text in images in other languages is transcribed. Intended Usage Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.imagetext-to-image100K<n<1M21 likes858 downloads2y agoHugging Face06Obscure-Entropy /conceptual_captions_jsonimage1M<n<10M0 likes806 downloads2y agoHugging Face07wchai /Video-Detailed-Caption Video Detailed Caption Benchmark Resources Website arXiv: Paper GitHub: Code Huggingface: AuroraCap Model Huggingface: VDC Benchmark Huggingface: Trainset Features Benchmark Collection and Processing We building VDC upon Panda-70M, Ego4D, Mixkit, Pixabay, and Pexels. Structured detailed captions construction pipeline. We develop a structured detailed captions construction pipeline to generate extra detailed descriptions from various… See the full description on the dataset page: https://huggingface.co/datasets/wchai/Video-Detailed-Caption.textvideo-text-to-text1K<n<10K17 likes752 downloads2y agoHugging Face08neurlang /Minecraft-Skins-Captioned-1M Dataset Card for Minecraft Skins Dataset Summary This dataset contains 981,079 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier. Dataset Structure Data Fields This dataset includes the following fields: hash: A data dependent hash. These hashes are generated from raw bytes and will be same if the skin is identical. image: The skin image encoded in base64 format.… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/Minecraft-Skins-Captioned-1M.textimage-classification1M<n<10M7 likes676 downloads1y agoHugging Face09Capx /ContextAwaretext100K<n<1M0 likes569 downloads2y agoHugging Face10CaptainGulu /R4R-Auto-Eval R4R Auto Eval 持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。 队友请先阅读 PROJECT_STATUS.md,然后按需查看: benchmarks/:固定的视频输入、来源记录和分层标签; pipelines/:判定方法及冻结配置; runs/:不可覆盖的实验记录; reports/:工作日志、方法分析和结果限制; registry/:benchmark、pipeline 和 run 的机器可读索引。 当前范围 multiscene30 是 pipeline 开发集,不是干净的留出测试集; reassemble40 是来自两个长录像的接触密集型校准集; 当前标签为来源数据提供方标签,尚未全部完成独立人工裁决; Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测; 在完成逐来源许可证核查前,本仓库应保持 private。 当前发布版本:0.1.0。 imagen<1K0 likes517 downloads10d agoHugging Face11csoai /claim-capture-census Claim-capture census A daily capture of the claims that public, authless catalogues serve about themselves and about the things they list. Produced by census-capture.py on CSOAI infrastructure. A capture is CLAIM_CAPTURED. It is not a measurement, not a grade, and not a certification. The only thing a capture proves is that these records existed in this exact form at this time as served by that source. Every artifact carries that boundary in claim_boundary. What is… See the full description on the dataset page: https://huggingface.co/datasets/csoai/claim-capture-census.tabular10K<n<100K0 likes473 downloads19h agoHugging Face12Capycap-AI /CaptchaSolve30k CaptchaSolve30k - Human Mouse Movement Dataset The largest open-source dataset of human task-specific mouse trajectories by session count and unique participants, with 30,000 discrete sessions from thousands of users. The first and only open dataset of complete human captcha-solving interactions with full behavioral replays. Each session captures mouse/touch trajectories, timing data, and puzzle state at physics-tick resolution. Suitable for bot detection research, human-computer… See the full description on the dataset page: https://huggingface.co/datasets/Capycap-AI/CaptchaSolve30k.tabularother10K<n<100K7 likes373 downloads5mo agoHugging Face13yauheniya-adesso /icongenai-svg-captions IconGenAI SVG Captions Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models. Part of the IconGenAI research project. Files Two files are provided at different stages of the processing pipeline: File Records Purpose icons_captioned_merged.jsonl 275,912 Full license-filtered corpus with VLM-generated captions and collection metadata icons_training_captioned.jsonl227,821 Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.tabulartext-to-image100K<n<1M2 likes366 downloads5mo agoHugging Face14Davd-b01 /thinking-cap-tier-curricula-complete Thinking Cap Tier Curricula — Complete Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 13,477 SFT samples and 3,187 SimPO preference pairs have undergone an automated token purge: Zero <|pad|> batch residues: 100% eliminated across all files. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct, non-repetitive… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete.texttext-generation10K<n<100K0 likes337 downloads12d agoHugging Face15DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes323 downloads2y agoHugging Face16CaptionEmporium /conceptual-captions-cc12m-llavanext Dataset Card for conceptual-captions-cc12m-llavanext Dataset Summary This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B. Languages The captions are in English. Data Instances An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.imagetext-to-image10M<n<100M28 likes267 downloads2y agoHugging Face17vvwangvv /emilia-captions-v3text10M<n<100M0 likes252 downloads8mo agoHugging Face18oss-codes /CA-Parallel-Dataset-Indictext100K<n<1M0 likes243 downloads1y agoHugging Face19Capx /MultiTurnChat CapX Scientific QA Dataset The CapX Scientific QA Dataset is a comprehensive collection of data designed to assist in the development of AI-powered tools that support scientists across various disciplines. This dataset aims to bridge the gap between machine learning and the scientific community by providing reliable and transparent resources for training and evaluating question-answering systems. Features CapX Scientific QA Dataset The CapX Scientific QA… See the full description on the dataset page: https://huggingface.co/datasets/Capx/MultiTurnChat.text1K<n<10K2 likes236 downloads2y agoHugging Face20Davd-b01 /thinking-cap-tier-lima-dense Thinking Cap Tier Curricula — LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 5,500 SFT and 2,000 SimPO records have undergone a complete token purge: Zero <|pad|> batch residues: 100% eliminated across all records. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct conclusions. Native ChatML… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-lima-dense.texttext-generation1K<n<10K2 likes227 downloads12d agoHugging Face21Davd-b01 /thinking-cap-tier-raw-traces Thinking Cap Tier Raw Traces (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized: Zero batch-padding residues (<|pad|>): Completely purged across all records. Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.tabulartext-generation10K<n<100K0 likes207 downloads12d agoHugging Face22joyce8 /EMBER2024-capa EMBER2024-capa Dataset Capa is a malware analysis tool that identifies a file's capabilities: +--------------------------------------+-----------------------------------------+ | CAPABILITY | NAMESPACE | |--------------------------------------|-----------------------------------------| | reference anti-VM strings | anti-analysis/anti-vm/vm-detection | | encode data using XOR (2 matches) |… See the full description on the dataset page: https://huggingface.co/datasets/joyce8/EMBER2024-capa.text10M<n<100M14 likes205 downloads1y agoHugging Face23false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes186 downloads18d agoHugging Face24embedding-data /coco_captions_quintets Dataset Card for "coco_captions" Dataset Summary COCO is a large-scale object detection, segmentation, and captioning dataset. This repo contains five captions per image; useful for sentence similarity tasks. Disclaimer: The team releasing COCO did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team. Supported Tasks Sentence Transformers training; useful for semantic search and sentence… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/coco_captions_quintets.textsentence-similarity10K<n<100K6 likes177 downloads4y agoHugging Face25OpenGVLab /InternVL-SA-1B-Caption Dataset Card for InternVL-SA-1B-Caption Overview The InternVL-SA-1B-Caption Dataset is a bilingual dataset created using the InternVL2-Llama3-76B model. The dataset contains 12 million image-caption pairs in both English and Chinese. All images are sourced from Meta’s SA-1B dataset, and captions were generated using specific prompts designed to minimize hallucinations and ensure accurate descriptions based on visible image content. The dataset is intended for use in tasks… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption.tabular1M<n<10M24 likes173 downloads2y agoHugging Face26malaiwah /qfs-capture-a1ec5246f3ef HF workflow a1ec5246f3ef11d323425ae8a9ea47ef A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from Qwen/Qwen3.8-27B. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut as… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-capture-a1ec5246f3ef.tabularn<1K0 likes169 downloads16d agoHugging Face27lntzm /CAPability Dataset Card for Dataset Name Visual caption benchmark Repo: CAPability [🍎 Project Page] [📖 ArXiv Paper] [🧑‍💻 Github Repo] [🏆 Leaderboard] Dataset Details Visual captioning benchmarks have become outdated with the emergence of modern MLLMs, as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation… See the full description on the dataset page: https://huggingface.co/datasets/lntzm/CAPability.image10K<n<100K2 likes159 downloads1y agoHugging Face28malaiwah /qfs-capture-6c6b45736e86 HF workflow 6c6b45736e86851500161bfeb9f44dd8 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from Qwen/Qwen3.8-27B. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut as… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-capture-6c6b45736e86.tabularn<1K0 likes158 downloads17d agoHugging Face29malaiwah /qfs-capture-02e41a4b6bc6 HF workflow 02e41a4b6bc63a9c8a9ffb68124ae5f7 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from Qwen/Qwen3.8-27B. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut as… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-capture-02e41a4b6bc6.tabularn<1K0 likes150 downloads16d agoHugging Face30Senqiao /LiDAR-LLM-Nu-Caption Dataset Details Dataset type: This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset. Dataset keys: "answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data. If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.textquestion-answering100K<n<1M8 likes145 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.