datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
distilabel-capybara-dpo-7k-binarized
Capybara-DPO 7K binarized
A DPO dataset built with distilabel atop the awesome LDJnr/Capybara
This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models.
Why?
Multi-turn dialogue data is key to fine-tune capable chat models. Multi-turn preference data has been used by the most relevant RLHF works (Anthropic, Meta Llama2, etc.). Unfortunately, there are very few… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-dpo-7k-binarized.capstone_sakuga_preproc_optical_flowEmbodied-Captioning
Embodied Image Captioning – Manually Annotated Test Set
Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning
📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.capstone_mlm_hidden_statesAgiBot-g1_left_capture_part
AgiBot-g1_left_capture_part
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: ruantong_a2d
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
factory
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
📊 Dataset Statistics
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_left_capture_part.Cobot_Magic_twist_bottle_cap
Cobot_Magic_twist_bottle_cap
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
twist
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_twist_bottle_cap.newyorker_caption_ranking
New Yorker Caption Ranking Dataset
Dataset Descriptions
Homepage: https://nextml.github.io/caption-contest-data/
Repository: https://github.com/yguooo/cartoon-caption-generation
Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
Point of Contact: yguo@cs.wisc.edu
Dataset Summary
We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.cml-tts-100h-cappedcapstone_sakuga_iblip_t5_embeddingspexels-568k-internvl2
Dataset Card for pexels-568k-internvl2
Dataset Summary
This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed.
Languages
The text is in English, but occasionally text in images in other languages is transcribed.
Intended Usage
Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.egocentric-vr-capture-20h-multimodal-sample
Egocentric VR Capture — 20-Hour Multimodal Inspection Sample
195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.coyo-hd-11m-llavanext
Dataset Card for coyo-hd-11m-llavanext
Dataset Summary
This is a data of 22,794,288 synthetic captions for 11,397,144 images from coyo-700m. The "hd" in the title refers to two aspects: high density and high definition. While large alt-text image pair datasets have many images, only a very small proportion of these images are in higher resolutions and have substantial concept density. For example, many of these datasets consist of more than 50% thumbnail sized or very… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/coyo-hd-11m-llavanext.danbooru-multitier-captions-202606
Danbooru — multi-tier captions (202606)
Per-post Danbooru data for the 202606 crawl: native tags, the raw API metadata, model-generated
multi-tier natural-language captions (long / refined long / medium / short), and post flags.
One row per Danbooru post_id. Images are not included — each post is referenced by
post_id, danbooru_url, md5, and the Danbooru CDN URLs. (The two example previews below are
downscaled for illustration.)
Based on:… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/danbooru-multitier-captions-202606.polymarket-arena-capture
Polymarket Arena — Live Capture
Raw live market-data capture from Polymarket's short-horizon crypto up/down markets
(BTC, ETH, SOL, XRP, BNB, DOGE) across the 5m / 15m / 1hr timeframes, plus the
underlying spot/reference price feeds. Collected by an always-on websocket collector
(collect_live.py) that subscribes to the CLOB book + trade streams and an RTDS price
feed, persisting to SQLite. This dataset is the lossless Parquet (zstd) export of that
capture.
Window: 2026-06-04 →… See the full description on the dataset page: https://huggingface.co/datasets/Alezanello/polymarket-arena-capture.moss-character-voices-top3-captioned
MOSS Character Voices — Top-3 per Group, Captioned (training-ready)
The top-3 takes per group (rank 0/1/2 by reward model) of all ~12,900 groups in
laion/moss-character-voices-bestof64
— ~38,700 samples — each annotated with both LAION voice-acting caption pipelines and shipped with
all scores + metadata, ready to drop into training. Sharded as data/train-*-of-00129.parquet.
Captions (two pipelines, same clip)
caption_procedural — Procedural Voice Captions:
terse… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-top3-captioned.AgiBot-g1_right_capture_part
AgiBot-g1_right_capture_part
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: ruantong_a2d
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
factory
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
📊 Dataset Statistics
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_right_capture_part.us-layoffs-per-capita-by-state-warn-act
US layoffs per capita by state: WARN notices and affected workers per 100,000 residents, 2020-2026
Rebuilt 2026-09-24. 299 state-years across 48 states; 34,624 notices in the table.
Latest complete year 2025: District of Columbia leads at 5.302 notices per 100k residents
(36 notices, 4,268 workers reported); the median state is 0.744.
"Which states are losing the most jobs per resident?" is a question no state agency answers,
because each of them publishes only its own notices… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-per-capita-by-state-warn-act.Cobot_Magic_cap_the_pen_a
Cobot_Magic_cap_the_pen_a
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
insert
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_cap_the_pen_a.claim-capture-census
Claim-capture census
A daily capture of the claims that public, authless catalogues serve about themselves and about
the things they list. Produced by census-capture.py on CSOAI infrastructure.
A capture is CLAIM_CAPTURED. It is not a measurement, not a grade, and not a certification.
The only thing a capture proves is that these records existed in this exact form at this time as
served by that source. Every artifact carries that boundary in claim_boundary.
What is… See the full description on the dataset page: https://huggingface.co/datasets/csoai/claim-capture-census.moda-general-capability-rollouts
MODA General Capability Retention Rollouts
This dataset contains the raw model generations and evaluation results for the
MODA general-capability retention experiments. It covers 16 models, seven
benchmarks, 260,592 prompt records, and 2,605,920 stored generations.
The evaluation code is pinned to source commit
12ea99b2a57a354f2b7d6792f62a3d9313192fa7.
Evaluation protocol
Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ,
HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.AIRBOT_MMK2_screw_the_bottle_cap
AIRBOT_MMK2_screw_the_bottle_cap
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_screw_the_bottle_cap.Capybara-Preferences
Dataset Card for Capybara-Preferences
This dataset has been created with distilabel.
Dataset Summary
This dataset is built on top of LDJnr/Capybara, in order to generate a preference
dataset out of an instruction-following dataset. This is done by keeping the conversations in the column conversation but splitting
the last assistant turn from it, so that the conversation contains all the turns up until the last user's turn, so that it can be reused… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Capybara-Preferences.R1_Lite_move_the_position_of_the_coffee_capsule
R1_Lite_move_the_position_of_the_coffee_capsule
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
place
pick
grasp
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_move_the_position_of_the_coffee_capsule.put_coffee_cap_teaboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 40,
"total_frames": 14806,
"total_tasks": 1,
"total_videos": 80,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lirislab/put_coffee_cap_teabox.CaptchaSolve30k
CaptchaSolve30k - Human Mouse Movement Dataset
The largest open-source dataset of human task-specific mouse trajectories by session count and unique participants, with 30,000 discrete sessions from thousands of users. The first and only open dataset of complete human captcha-solving interactions with full behavioral replays.
Each session captures mouse/touch trajectories, timing data, and puzzle state at physics-tick resolution. Suitable for bot detection research, human-computer… See the full description on the dataset page: https://huggingface.co/datasets/Capycap-AI/CaptchaSolve30k.icongenai-svg-captions
IconGenAI SVG Captions
Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models.
Part of the IconGenAI research project.
Files
Two files are provided at different stages of the processing pipeline:
File
Records
Purpose
icons_captioned_merged.jsonl
275,912
Full license-filtered corpus with VLM-generated captions and collection metadata
icons_training_captioned.jsonl227,821
Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.mscoco-1st-captionTo reproduce, run pip install -r requirements.txt and download.sh.
flickr30k_clip-SimCLRv2-caption_pairspick_coffee_capsule_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5",
"total_episodes": 770,
"total_frames": 452236,
"total_tasks": 84,
"total_videos": 1540,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:770"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/pick_coffee_capsule_merged.
