datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Recap-DataComp-1B
Dataset Card for Recap-DataComp-1B
Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions.
Dataset Details
Dataset Description
Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM.
Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.recap-datacomp-12m-wdsdanbooru_2025_recaption
内部暂存数据集 (Internal Temporary Dataset)
English
This is a temporary dataset for internal use.
It might contain:
Items being re-processed or corrected (e.g., some images requiring re-tagging using a distributed cluster, needing a convenient data source for it).
Data to supplement our internal systems (e.g., if a machine accidentally lost some images and we don't want to re-download everything).
Recent updates or experimental data not yet finalized (e.g., the image source… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/danbooru_2025_recaption.recap16_20260628_111530_collectedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 16,
"total_frames": 1280,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 1e-06,
"fps": 10,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miical/recap16_20260628_111530_collected.datacomp_recap_metadata2Recap-DataComp-1B_split_3Recap-DataComp-1B_split_4Recap-DataComp-1B-FoodOrDrink
Recap-DataComp-1B: Food or Drink
A filtered subset of Recap-DataComp-1B containing 106,230,157 rows classified as food/drink content, enriched with structured food/drink extraction from FoodExtract-v2.
Overview
Count
Percentage
Total rows
106,230,157
100%
Food/drink (Stage 5 label)
96,618,895
91.0%
Not food/drink (Stage 5 label)
9,611,262
9.0%
FoodExtract (re_caption): food/drink
79,519,489
74.9%
FoodExtract (re_caption): not food/drink
26,710,156… See the full description on the dataset page: https://huggingface.co/datasets/mrdbourke/Recap-DataComp-1B-FoodOrDrink.Recap-DataComp-1B_split_2Recap-DataComp-1B_split_5Recap-DataComp-1B_split_7Recap-DataComp-1B_split_8laion6m_recapRecap-DataComp-1B_split_6Recap-DataComp-1B_split_1commoncatalog-cc-by-recap-qwen3p5-35b-a3b
CommonCatalog CC-BY recaptions with Qwen3.5-35B-A3B
Dataset commoncatalog-cc-by: 14.577 Million caption rows.
This public caption-only repository contains 14,576,560 generated captions for 14,576,558 image assets and no image payload. It includes 2 additional distinct caption variants. Rows match the public image-bearing common-canvas/commoncatalog-cc-by release at revision 80f50fe4a1ca937f37a11be3f8eee5199d776ff3 through Flickr photoid, represented here by asset_instance_id… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/commoncatalog-cc-by-recap-qwen3p5-35b-a3b.vials_recap_v1.1
vials_recap_v1.1
Teleoperation dataset: pick up all vials and place them in the stand
Collected with limb on yam arms.
Dataset summary
Field
Value
Robot
yam
Episodes
93
Total frames
105,592
FPS
30 Hz
Task
pick up all vials and place them in the stand
Format
LeRobot v3.0
Cameras
Name
head_camera
left_wrist_camera
right_wrist_camera
Video codec: hevc.
State space (observation.state, shape [14])… See the full description on the dataset page: https://huggingface.co/datasets/Sichang0621/vials_recap_v1.1.rollout_recap_stackblocks_iter0_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/rollout_recap_stackblocks_iter0_merged.gr00t-arena-recap-iter1-step1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "gr1_joint",
"total_episodes": 32,
"total_frames": 15656,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miical/gr00t-arena-recap-iter1-step1000.pi05-libero10-task8-recapThis dataset was created using LeRobot.
Dataset Description
This is the final cumulative LeRobot dataset from the PI0.5 ReCap experiment
on LIBERO-10 task 8:
Put both moka pots on the stove.
It contains the 10 official demonstrations used to train the initial policy,
followed by three rounds of 32 online rollouts.
Episode ranges
episode_index
Source
Episodes
Successful episodes
0–9
Initial official demonstrations
10
10
10–41
ReCap iteration 1… See the full description on the dataset page: https://huggingface.co/datasets/Miical/pi05-libero10-task8-recap.redcaps5m_recapmerged-recap-cup-training-h264
Merged RECAP Cup Training Dataset
This dataset contains merged data for training a RECAP (RL with Experience and Corrections via Advantage-conditioned Policies) value function for bimanual cup positioning.
Features
Feature
Type
Shape
observation.images.left
video
(480, 640, 3)
observation.images.center
video
(480, 640, 3)
observation.images.right
video
(480, 640, 3)
observation.state
float32
(12,)
action
float32
(12,)
is_intervention
int64
(1,)… See the full description on the dataset page: https://huggingface.co/datasets/apaszynska/merged-recap-cup-training-h264.gr00t_arena_gr1_recap_sftThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "gr1_joint",
"total_episodes": 32,
"total_frames": 14763,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 1e-06,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miical/gr00t_arena_gr1_recap_sft.laion-highres-aesthetic-recap-qwen3p5-35b-a3b
LAION Highres Aesthetic recaptions with Qwen3.5-35B-A3B
Dataset laion-highres-aesthetic: 96.072 Million caption rows.
This public caption-only repository contains 96,071,504 generated captions for 82,690,560 image instances and no image payload. It includes 13,380,944 additional distinct caption variants. The matching gated image payload is hosted on ModelScope at BootsofLagrangian/laion-highres-aesthetic-webp90-hybrid-resolution. The Hugging Face image repository… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/laion-highres-aesthetic-recap-qwen3p5-35b-a3b.stackblocks_recap_iter1_demo_rollout_correctionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/stackblocks_recap_iter1_demo_rollout_correction.gr00t_arena_gr1_recap_iter2_from9500This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "gr1_joint",
"total_episodes": 64,
"total_frames": 28658,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 1e-06,
"fps": 30,
"splits": {
"train": "0:64"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miical/gr00t_arena_gr1_recap_iter2_from9500.recap_dish_20260611This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 71,
"total_frames": 39618,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:71"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mt-prox/recap_dish_20260611.relaion-art-recap-zh
Dataset Card for Relaion-Art Recaptioned (Chinese)
This dataset is a recaptioned version of laion/relaion-art, featuring high-quality Chinese captions generated using Alibaba's Qwen3-VL-Flash model via the DashScope Batch API.
This dataset contains 2,007,213 image-caption pairs derived from the original Relaion-Art dataset (~8M samples). Each image has been recaptioned with detailed, accurate Chinese descriptions that highlight the subject, scene, style, and key visual details.… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/relaion-art-recap-zh.stackblocks_recap_iter2_annotatedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/stackblocks_recap_iter2_annotated.stackblocks_red_yellow_so100_blue_recapThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/stackblocks_red_yellow_so100_blue_recap.
