datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ThinkV2V-Bench
📌 Benchmark Summary
ThinkV2V-Bench is a complex instruction benchmark for instruction-guided video editing. It is derived from OpenVE-Bench, while focusing on editing categories where multimodal large language model (MLLM) thinking is more important for understanding and executing the requested transformation.
ThinkV2V-Bench contains 308 samples from five spatial editing categories:
background_change 59
global_style 58
local_add 67
local_change 65… See the full description on the dataset page: https://huggingface.co/datasets/donghao-zhou/ThinkV2V-Bench.217-Thinking_Beyond-pickingSaltSticksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 1360,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/217-Thinking_Beyond-pickingSaltSticks.Thinking-in-Video-DataThinking in Video: Can Video Generators Really Reason About the Real World?
This repository contains the official implementation of Causal-Generative Dual-Judge (CGDJ) for auditing world-model consistency of video generative models — the official codebase of the Thinking in Video paradigm.
🌟 Overview
Thinking in Video is a reasoning paradigm in which a video generative model is used not merely to synthesize pixels, but to simulate, predict, and verify causal… See the full description on the dataset page: https://huggingface.co/datasets/BRZ911/Thinking-in-Video-Data.thinkworld_final_oven_robot_human_composition_v1
ThinkWorld oven composition prompt v1
This is an additive metadata derivation of
chyun/thinkworld_final_oven_robot_human@a0c9dcca7a9e2ba394c306bcdf6e7161b712dbc0. Video bytes, timing, robot
state/action labels, supervision masks, camera layout, and source provenance are
unchanged. Human episodes retain their original global composite prompt exactly.
The ordinary LeRobot task resolves from task_index to the episode-global
instruction. atomic_task_index retains the current skill… See the full description on the dataset page: https://huggingface.co/datasets/chyun/thinkworld_final_oven_robot_human_composition_v1.thinking-with-video-pretrain
