CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01argilla /distilabel-capybara-dpo-7k-binarized Capybara-DPO 7K binarized A DPO dataset built with distilabel atop the awesome LDJnr/Capybara This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models. Why? Multi-turn dialogue data is key to fine-tune capable chat models. Multi-turn preference data has been used by the most relevant RLHF works (Anthropic, Meta Llama2, etc.). Unfortunately, there are very few… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-dpo-7k-binarized.tabularquestion-answering1K<n<10K184 likes23k downloads2y agoHugging Face02Hemabhushan /capstone_sakuga_preproc_optical_flowtabular100K<n<1M0 likes14k downloads2y agoHugging Face03TommyBsk /Embodied-Captioning Embodied Image Captioning – Manually Annotated Test Set Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning 📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.tabularimage-to-text1K<n<10K0 likes8.8k downloads1y agoHugging Face04Hemabhushan /capstone_mlm_hidden_statestabular100K<n<1M0 likes5.7k downloads2y agoHugging Face05RoboCOIN /AgiBot-g1_left_capture_partgated AgiBot-g1_left_capture_part 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: ruantong_a2d | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: factory 🤖 Atomic Actions This dataset includes the following atomic actions: grasp 📊 Dataset Statistics Metric Value Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_left_capture_part.tabularrobotics100K<n<1M0 likes4k downloads9mo agoHugging Face06RoboCOIN /Cobot_Magic_twist_bottle_capgated Cobot_Magic_twist_bottle_cap 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place twist 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_twist_bottle_cap.tabularrobotics100K<n<1M1 likes2.4k downloads9mo agoHugging Face07yguooo /newyorker_caption_ranking New Yorker Caption Ranking Dataset Dataset Descriptions Homepage: https://nextml.github.io/caption-contest-data/ Repository: https://github.com/yguooo/cartoon-caption-generation Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning Point of Contact: yguo@cs.wisc.edu Dataset Summary We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.imagetext-generation1M<n<10M6 likes1.4k downloads2y agoHugging Face08ibm-esa-geospatial /Llama3-SSL4EO-S12-v1.1-captions Llama3-SSL4EO-S12-Captions The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model. Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper. Code: https://github.com/IBM/MS-CLIP Data Structure We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.tabularzero-shot-image-classification100K<n<1M5 likes1.4k downloads1y agoHugging Face09jigsaws-stomper /cml-tts-100h-cappedtabular100K<n<1M0 likes1.2k downloads6mo agoHugging Face10Hemabhushan /capstone_sakuga_iblip_t5_embeddingstabular10K<n<100K0 likes1k downloads2y agoHugging Face11CaptionEmporium /pexels-568k-internvl2 Dataset Card for pexels-568k-internvl2 Dataset Summary This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed. Languages The text is in English, but occasionally text in images in other languages is transcribed. Intended Usage Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.imagetext-to-image100K<n<1M21 likes858 downloads2y agoHugging Face12ExylosAi /egocentric-vr-capture-20h-multimodal-sample Egocentric VR Capture — 20-Hour Multimodal Inspection Sample 195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package. This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.tabularrobotics1M<n<10M0 likes803 downloads3d agoHugging Face13CaptionEmporium /coyo-hd-11m-llavanext Dataset Card for coyo-hd-11m-llavanext Dataset Summary This is a data of 22,794,288 synthetic captions for 11,397,144 images from coyo-700m. The "hd" in the title refers to two aspects: high density and high definition. While large alt-text image pair datasets have many images, only a very small proportion of these images are in higher resolutions and have substantial concept density. For example, many of these datasets consist of more than 50% thumbnail sized or very… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/coyo-hd-11m-llavanext.imagetext-to-image10M<n<100M29 likes752 downloads2y agoHugging Face14BootsofLagrangian /danbooru-multitier-captions-202606 Danbooru — multi-tier captions (202606) Per-post Danbooru data for the 202606 crawl: native tags, the raw API metadata, model-generated multi-tier natural-language captions (long / refined long / medium / short), and post flags. One row per Danbooru post_id. Images are not included — each post is referenced by post_id, danbooru_url, md5, and the Danbooru CDN URLs. (The two example previews below are downscaled for illustration.) Based on:… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/danbooru-multitier-captions-202606.imagetext-to-image10M<n<100M2 likes657 downloads1mo agoHugging Face15Alezanello /polymarket-arena-capture Polymarket Arena — Live Capture Raw live market-data capture from Polymarket's short-horizon crypto up/down markets (BTC, ETH, SOL, XRP, BNB, DOGE) across the 5m / 15m / 1hr timeframes, plus the underlying spot/reference price feeds. Collected by an always-on websocket collector (collect_live.py) that subscribes to the CLOB book + trade streams and an RTDS price feed, persisting to SQLite. This dataset is the lossless Parquet (zstd) export of that capture. Window: 2026-06-04 →… See the full description on the dataset page: https://huggingface.co/datasets/Alezanello/polymarket-arena-capture.tabular10M<n<100M2 likes639 downloads3mo agoHugging Face16laion /moss-character-voices-top3-captioned MOSS Character Voices — Top-3 per Group, Captioned (training-ready) The top-3 takes per group (rank 0/1/2 by reward model) of all ~12,900 groups in laion/moss-character-voices-bestof64 — ~38,700 samples — each annotated with both LAION voice-acting caption pipelines and shipped with all scores + metadata, ready to drop into training. Sharded as data/train-*-of-00129.parquet. Captions (two pipelines, same clip) caption_procedural — Procedural Voice Captions: terse… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-top3-captioned.tabulartext-to-speech10K<n<100K0 likes623 downloads2mo agoHugging Face17RoboCOIN /AgiBot-g1_right_capture_partgated AgiBot-g1_right_capture_part 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: ruantong_a2d | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: factory 🤖 Atomic Actions This dataset includes the following atomic actions: grasp 📊 Dataset Statistics Metric Value Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_right_capture_part.tabularrobotics10K<n<100K0 likes572 downloads9mo agoHugging Face18APProjects /us-layoffs-per-capita-by-state-warn-act US layoffs per capita by state: WARN notices and affected workers per 100,000 residents, 2020-2026 Rebuilt 2026-09-24. 299 state-years across 48 states; 34,624 notices in the table. Latest complete year 2025: District of Columbia leads at 5.302 notices per 100k residents (36 notices, 4,268 workers reported); the median state is 0.744. "Which states are losing the most jobs per resident?" is a question no state agency answers, because each of them publishes only its own notices… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-per-capita-by-state-warn-act.tabulartabular-regressionn<1K1 likes519 downloads1d agoHugging Face19RoboCOIN /Cobot_Magic_cap_the_pen_agated Cobot_Magic_cap_the_pen_a 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place insert 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_cap_the_pen_a.tabularrobotics10K<n<100K0 likes505 downloads9mo agoHugging Face20csoai /claim-capture-census Claim-capture census A daily capture of the claims that public, authless catalogues serve about themselves and about the things they list. Produced by census-capture.py on CSOAI infrastructure. A capture is CLAIM_CAPTURED. It is not a measurement, not a grade, and not a certification. The only thing a capture proves is that these records existed in this exact form at this time as served by that source. Every artifact carries that boundary in claim_boundary. What is… See the full description on the dataset page: https://huggingface.co/datasets/csoai/claim-capture-census.tabular10K<n<100K0 likes473 downloads15h agoHugging Face21Hkang /moda-general-capability-rollouts MODA General Capability Retention Rollouts This dataset contains the raw model generations and evaluation results for the MODA general-capability retention experiments. It covers 16 models, seven benchmarks, 260,592 prompt records, and 2,605,920 stored generations. The evaluation code is pinned to source commit 12ea99b2a57a354f2b7d6792f62a3d9313192fa7. Evaluation protocol Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ, HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.tabulartext-generation1M<n<10M0 likes430 downloads2mo agoHugging Face22RoboCOIN /AIRBOT_MMK2_screw_the_bottle_capgated AIRBOT_MMK2_screw_the_bottle_cap 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: discover_robotics_aitbot_mmk2 | Codebase Version: v2.1 End-Effector Type: five_finger_hand 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_screw_the_bottle_cap.tabularrobotics10K<n<100K0 likes425 downloads9mo agoHugging Face23argilla /Capybara-Preferences Dataset Card for Capybara-Preferences This dataset has been created with distilabel. Dataset Summary This dataset is built on top of LDJnr/Capybara, in order to generate a preference dataset out of an instruction-following dataset. This is done by keeping the conversations in the column conversation but splitting the last assistant turn from it, so that the conversation contains all the turns up until the last user's turn, so that it can be reused… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Capybara-Preferences.tabulartext-generation10K<n<100K47 likes420 downloads2y agoHugging Face24RoboCOIN /R1_Lite_move_the_position_of_the_coffee_capsulegated R1_Lite_move_the_position_of_the_coffee_capsule 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: galaxea_r1_lite | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: place pick grasp 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_move_the_position_of_the_coffee_capsule.tabularrobotics10K<n<100K0 likes385 downloads9mo agoHugging Face25lirislab /put_coffee_cap_teaboxThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 40, "total_frames": 14806, "total_tasks": 1, "total_videos": 80, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:40" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lirislab/put_coffee_cap_teabox.tabularrobotics10K<n<100K0 likes374 downloads1y agoHugging Face26Capycap-AI /CaptchaSolve30k CaptchaSolve30k - Human Mouse Movement Dataset The largest open-source dataset of human task-specific mouse trajectories by session count and unique participants, with 30,000 discrete sessions from thousands of users. The first and only open dataset of complete human captcha-solving interactions with full behavioral replays. Each session captures mouse/touch trajectories, timing data, and puzzle state at physics-tick resolution. Suitable for bot detection research, human-computer… See the full description on the dataset page: https://huggingface.co/datasets/Capycap-AI/CaptchaSolve30k.tabularother10K<n<100K7 likes373 downloads5mo agoHugging Face27yauheniya-adesso /icongenai-svg-captions IconGenAI SVG Captions Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models. Part of the IconGenAI research project. Files Two files are provided at different stages of the processing pipeline: File Records Purpose icons_captioned_merged.jsonl 275,912 Full license-filtered corpus with VLM-generated captions and collection metadata icons_training_captioned.jsonl227,821 Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.tabulartext-to-image100K<n<1M2 likes366 downloads5mo agoHugging Face28cat-state /mscoco-1st-captionTo reproduce, run pip install -r requirements.txt and download.sh. image100K<n<1M3 likes347 downloads4y agoHugging Face29closji /flickr30k_clip-SimCLRv2-caption_pairstabular10M<n<100M1 likes335 downloads4y agoHugging Face30villekuosmanen /pick_coffee_capsule_mergedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "arx5", "total_episodes": 770, "total_frames": 452236, "total_tasks": 84, "total_videos": 1540, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:770" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/pick_coffee_capsule_merged.tabularrobotics100K<n<1M1 likes315 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.