datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cauldron-JA
Dataset Card for The Cauldron-JA
Dataset description
The Cauldron-JA is a Vision Language Model dataset that translates 'The Cauldron' into Japanese using the DeepL API. The Cauldron is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2.
To create a Japanese Vision Language Dataset, datasets related to OCR, coding, and graphs were excluded because translating them into Japanese… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Cauldron-JA.Japan-Open-Driving-Dataset-Sample
Japan Open Driving Dataset Sample
Overview
This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan.
The data is stored in nuScenes format and can be loaded with the nuscenes-devkit.
In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.CoVLA-Dataset
CoVLA-Dataset
WACV 2025 Oral
CoVLA-Dataset is a dataset comprising real-world driving videos spanning more than 80 hours. This dataset leverages a novel, scalable approach based on automated data processing and a caption generation pipeline to generate accurate driving trajectories paired with detailed natural language descriptions of driving environments and maneuvers. It includes 10,000 30-second video clips, paired with trajectory targets and language annotations generated from… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/CoVLA-Dataset.motor_two_wheel_rider
Dataset Card for motor_two_wheel_rider
MOTOR (MOtorized TwO-wheeler Rider) is the first large-scale, multi-view, multimodal dataset dedicated to understanding two-wheeler rider behavior in dense, unstructured traffic conditions typical of the Global South. The full dataset comprises 1,629 annotated sequences (~25 hours) from 16 riders collected across diverse traffic scenarios in India.
This repository contains a subset of the MOTOR dataset imported into FiftyOne format for easy… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/motor_two_wheel_rider.STRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.RIO-Bench
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
Real-world VLMs must decide when to read text and when to ignore it, e.g., reading traffic signs but not being fooled by text-based attacks on objects.
We propose a unified benchmark, RIO-Bench, to evaluate both typographic-attack robustness and text recognition in VLMs through a novel task called RIO-VQA.
Problem Settings: VLMs Must Adaptively Read… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/RIO-Bench.MOMIJI
MOMIJI Dataset Card
English | 日本語
MOMIJI (Modern Open Multimodal Japanese filtered Dataset) is a large-scale Japanese image-text interleaved dataset built from Common Crawl.
“Interleaved” means that text and images are associated while preserving the order in which they appear within a document. In MOMIJI, image positions are represented by placeholders such as <image1>.
Dataset: turing-motors/MOMIJI
Dataset construction code: turingmotors/MOMIJI
Data generation utility:… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/MOMIJI.MOTOR
MOTOR — A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding
Project Page | Paper | Code
MOTOR is the first large-scale, multi-view, multimodal resource dedicated to two-wheelers in dense, unstructured traffic. It comprises 1,629 annotated sequences (25+ hours of video data) collected from 16 riders and integrates synchronized front, rear, and helmet videos, rider eye-gaze from wearable trackers, on-road audio, and telemetry (GPS, accelerometer, gyroscope).… See the full description on the dataset page: https://huggingface.co/datasets/varunpaturkar/MOTOR.MCC5-THU-Motor
Multi-mode Fault Diagnosis Datasets of Three-phase Asynchronous Motor Under Variable Working Conditions
Dataset Summary
This dataset provides synchronized multi-modal time-series data collected from a 2.2 kW three-phase asynchronous (induction) motor operating under variable speed/load conditions with deliberately induced faults. It is intended for developing and benchmarking robust fault diagnosis methods under realistic operating scenarios, especially for… See the full description on the dataset page: https://huggingface.co/datasets/Samlzy/MCC5-THU-Motor.tomogram-Bacterial-Flagellar-motors-location
Dataset Card for Bacterial Flagellar Motor Tomograms (Kaggle)
Dataset Description
Homepage: Kaggle Competition: BYU - Locating Bacterial Flagellar Motors 2025
Dataset Summary
This dataset originates from the Kaggle competition "BYU - Locating Bacterial Flagellar Motors 2025". The goal is to identify the presence and location of flagellar motors within 3D cryogenic electron tomography (cryo-ET) reconstructions (tomograms) of bacteria. Flagellar motors are… See the full description on the dataset page: https://huggingface.co/datasets/Floppanacci/tomogram-Bacterial-Flagellar-motors-location.multisource-eeg-motorimageryACT-Bench
ACT-Bench
ACT-Bench is a dedicated framework for quantitatively evaluating the action controllability of world models for autonomous driving.
It focuses on measuring how well a world model can generate driving scenes conditioned on specified trajectories.
For more details, please refer our paper and code repository.
Data fields
Key
Value
sample_id
0
label
'straight_constant_speed/straight_constant_speed_10kmph'
context_frames… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/ACT-Bench.STRIDE-QA-Dataset-Mini
STRIDE-QA-Dataset-Mini
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
⚠️ Note: STRIDE-QA-Dataset-Mini is provided as a preliminary version and does not fully match the format of the… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset-Mini.motor-medicina-legala-v2-storagemotor-nameplate
Motor Nameplate
A small image dataset of electric motor nameplates collected from public Google Images results, intended for tasks such as:
Training OCR / document-understanding models to extract nameplate text fields (manufacturer, HP, RPM, voltage, frame, etc.)
Image classification by manufacturer (ABB, Siemens, Baldor-Reliance, WEG, Hyundai, etc.)
Object detection (locating the plate on the motor body)
Few-shot learning / evaluation baselines for industrial vision tasks
The… See the full description on the dataset page: https://huggingface.co/datasets/sergiudanstan/motor-nameplate.maker_arm_place_motor_in_box_20260821_163049This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/makermods/maker_arm_place_motor_in_box_20260821_163049.STRIDE-QA-Bench
STRIDE-QA-Bench
STRIDE-QA-Bench provides a standardized benchmark for evaluating spatiotemporal reasoning of Vision-Language Models (VLMs) in autonomous driving.This HuggingFace repository provides the images and JSON files of the benchmark.
For detailed benchmark description and execution code, please refer to STRIDE-QA-Dataset (GitHub).
🗂️ Data Fields
The main data fields are as follows.
Field
Type
Description
question_id
str
Unique question ID.… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Bench.omy_f3m_motor_feedback_vla_probe_rock_real_left_probe_rightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 20,
"total_frames": 10481,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_vla_probe_rock_real_left_probe_right.G1_Dex1_Assembling_Motormotor-nameplate-labels
Motor Nameplate — Auto-Labeled (OCR + Qwen2.5-1.5B)
Motor Nameplate — Auto-Labeled (OCR + Qwen2.5-1.5B)
A derived dataset that adds structured field annotations to the
sergiudanstan/motor-nameplate
image collection (286 motor-nameplate photos from Google Images).
The original upstream dataset ships images only. This repo adds an
auto-labeling pipeline output, ready for downstream fine-tuning of a
vision-language model for nameplate field extraction.… See the full description on the dataset page: https://huggingface.co/datasets/sergiudanstan/motor-nameplate-labels.omy_f3m_motor_feedback_test_baseball_aThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1399,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_baseball_a.omy_f3m_motor_feedback_test_papercup_aThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1800,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_papercup_a.team-7-left-arm-grasp-motorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 99,
"total_frames": 59123,
"total_tasks": 1,
"total_videos": 198,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:99"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/roboticshack/team-7-left-arm-grasp-motor.omy_f3m_motor_feedback_test_papercup_bThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1801,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_papercup_b.omy_f3m_motor_feedback_test_remon_dThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1802,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_remon_d.omy_f3m_motor_feedback_test_tennis_dThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1801,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_tennis_d.omy_f3m_motor_feedback_test_papercup_cThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1802,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_papercup_c.omy_f3m_motor_feedback_test_remon_aThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1802,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_remon_a.omy_f3m_motor_feedback_test_tennis_cThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1803,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_tennis_c.omy_f3m_motor_feedback_test_papercup_dThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1801,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_papercup_d.
