motor
Datasets
All datasets matching “motor”Cauldron-JA
Dataset Card for The Cauldron-JA
Dataset description
The Cauldron-JA is a Vision Language Model dataset that translates 'The Cauldron' into Japanese using the DeepL API. The Cauldron is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2.
To create a Japanese Vision Language Dataset, datasets related to OCR, coding, and graphs were excluded because translating them into Japanese… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Cauldron-JA.Japan-Open-Driving-Dataset-Sample
Japan Open Driving Dataset Sample
Overview
This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan.
The data is stored in nuScenes format and can be loaded with the nuscenes-devkit.
In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.CoVLA-Dataset
CoVLA-Dataset
WACV 2025 Oral
CoVLA-Dataset is a dataset comprising real-world driving videos spanning more than 80 hours. This dataset leverages a novel, scalable approach based on automated data processing and a caption generation pipeline to generate accurate driving trajectories paired with detailed natural language descriptions of driving environments and maneuvers. It includes 10,000 30-second video clips, paired with trajectory targets and language annotations generated from… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/CoVLA-Dataset.motor_two_wheel_rider
Dataset Card for motor_two_wheel_rider
MOTOR (MOtorized TwO-wheeler Rider) is the first large-scale, multi-view, multimodal dataset dedicated to understanding two-wheeler rider behavior in dense, unstructured traffic conditions typical of the Global South. The full dataset comprises 1,629 annotated sequences (~25 hours) from 16 riders collected across diverse traffic scenarios in India.
This repository contains a subset of the MOTOR dataset imported into FiftyOne format for easy… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/motor_two_wheel_rider.STRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.RIO-Bench
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
Real-world VLMs must decide when to read text and when to ignore it, e.g., reading traffic signs but not being fooled by text-based attacks on objects.
We propose a unified benchmark, RIO-Bench, to evaluate both typographic-attack robustness and text recognition in VLMs through a novel task called RIO-VQA.
Problem Settings: VLMs Must Adaptively Read… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/RIO-Bench.
