datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
retrieval-try-1
mikasa/retrieval_stock224_h264_crf23
MIKASA recordings of MikasaCabinetRetrieval-v0: 1000 episodes, 354498 frames, 10 Hz.
actions (13) are the env's own pd_joint_delta_pos actions, every channel in [-1, 1], fed straight to env.step — see meta/mikasa_format.json for the channel table, the state layout and the inactive channels [] (zero variance here; do not normalize them by their std).
state (15) is obs['agent']['qpos'] and nothing else.
meta/mikasa_episodes.csv lists every… See the full description on the dataset page: https://huggingface.co/datasets/nurtayev-d/retrieval-try-1.newbalance_shoe_insole_retrieval_and_packing_0611This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 298,
"total_frames": 4967405,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:298"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0611.shoe_insole_retrieval_and_packing0515This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 101,
"total_frames": 206299,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/shoe_insole_retrieval_and_packing0515.object_retrieval-preset-gemini
object_retrieval-preset-gemini
Real-robot teleoperation episodes of the instance_retrieval task on a single-arm Franka Research 3 cell (80 episodes, 36,239 frames at 10 fps,
released as Myungkyu/object_retrieval) with dense high-level labels produced by the TACOR offline annotator:
Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 10 frames = 1.0 s) and labels every sampled frame given only the
subtask preset of the task - the label list below… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/object_retrieval-preset-gemini.newbalance_shoe_insole_retrieval_and_packing_0604This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 108,
"total_frames": 1635459,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:108"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0604.object_retrieval
instance_retrieval
Real-robot teleoperation demonstrations of the instance_retrieval task on a single-arm Franka Research 3 cell,
released in four LeRobot layouts. Every layout is a conversion of the same 80 raw episodes
(36,239 frames at 10 Hz); the layouts differ only in the LeRobot codebase version and in the
action representation.
directory
LeRobot version
action (action)
consumer
lerobot_v21_abs_joint/
v2.1
8-D absolute joint targets + gripper
RLDX-1 loader… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/object_retrieval.avcaps-retrieval
AVCaps audio–visual retrieval (MTEB)
Retrieval tasks over AVCaps, an
audio-visual dataset derived from VidOR in which each clip is captioned three separate
ways — from the audio alone, from the visuals alone, and from both together.
That separation is the point: the audio-only, video-only and combined directions can be
scored independently on identical clips, rather than inferred from a single caption set
that mixes the modalities.
Prepared for MTEB as six tasks:… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/avcaps-retrieval.newbalance_shoe_insole_retrieval_and_packing_0529This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 65,
"total_frames": 332561,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:65"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0529.vatex-multilingual-retrieval
VATEX multilingual (en/zh) video retrieval (MTEB)
Cross-lingual video retrieval over VATEX, which
captions every clip in both English and Chinese. The two subsets share one video
corpus and differ only in caption language, which makes a like-for-like cross-lingual
comparison possible.
Prepared for MTEB as
VATEXMultilingualT2VRetrieval and VATEXMultilingualV2TRetrieval.
Contents
config
rows
description
videos
993
shared video corpus, 10s clips
en
993… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/vatex-multilingual-retrieval.huili_shoe_insole_retrieval_and_packing_0602This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 68,
"total_frames": 949788,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:68"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Vertax/huili_shoe_insole_retrieval_and_packing_0602.vtv24-video-retrieval-benchmark
VTV24 Video Retrieval Benchmark
This dataset contains 10 VTV24 videos and 33 query-to-timestamp annotations for video retrieval benchmarking. The videos are provided as 360p MP4 files, alongside video-retrieval-benchmark.json.
The annotation tool is available at https://github.com/duclvq/video-retrieval-bench.
No license claim is made over the source broadcasts; users are responsible for ensuring that their use complies with applicable rights and terms.
object_retrieval-taco-gemini
object_retrieval-taco-gemini
Real-robot teleoperation episodes of the Object Identification task (internal id instance_retrieval) on a single-arm Franka Research 3
cell (80 episodes, 36,239 frames at 10 fps, released as Myungkyu/object_retrieval)
with dense high-level labels and visual memory produced by the TACOR offline annotator: Gemini 3.7 Flash reads each whole episode as one
video clip (one sample every 10 frames = 1.0 s) and labels every sampled frame given the full task… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/object_retrieval-taco-gemini.huili_shoe_insole_retrieval_and_packing_0602This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 68,
"total_frames": 949788,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:68"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/huili_shoe_insole_retrieval_and_packing_0602.huili_shoe_insole_retrieval_and_packing_0528This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 30,
"total_frames": 125428,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/huili_shoe_insole_retrieval_and_packing_0528.huili_shoe_insole_retrieval_and_packing_0602This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 68,
"total_frames": 949788,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:68"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/enlong1/huili_shoe_insole_retrieval_and_packing_0602.newbalance_shoe_insole_retrieval_and_packing_0520This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 53,
"total_frames": 179478,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:53"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0520.huili_shoe_insole_retrieval_and_packing_0611This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 4,
"total_frames": 82453,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/huili_shoe_insole_retrieval_and_packing_0611.raise-nocurr-fridge-retrievalbreakfast-toast-retrieval-platingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 5904,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CreatorKanata/breakfast-toast-retrieval-plating.car-video-retrieval
Car Video Retrieval
You give it a picture of a car part (wheel, headlight, etc.) and it finds where that part shows up in a long car review video — with timestamps and YouTube links to jump straight to those moments.
How it works
YOLOv8-World detects car parts in both the video and your query image (no training, just a custom list of part names).
Detections are stored in Parquet so we can quickly look up “all frames where we saw a wheel,” etc.
Retrieval runs the same… See the full description on the dataset page: https://huggingface.co/datasets/vnhngf/car-video-retrieval.
