datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cauldron-JA
Dataset Card for The Cauldron-JA
Dataset description
The Cauldron-JA is a Vision Language Model dataset that translates 'The Cauldron' into Japanese using the DeepL API. The Cauldron is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2.
To create a Japanese Vision Language Dataset, datasets related to OCR, coding, and graphs were excluded because translating them into Japanese… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Cauldron-JA.Japan-Open-Driving-Dataset-Sample
Japan Open Driving Dataset Sample
Overview
This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan.
The data is stored in nuScenes format and can be loaded with the nuscenes-devkit.
In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.STRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.RIO-Bench
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
Real-world VLMs must decide when to read text and when to ignore it, e.g., reading traffic signs but not being fooled by text-based attacks on objects.
We propose a unified benchmark, RIO-Bench, to evaluate both typographic-attack robustness and text recognition in VLMs through a novel task called RIO-VQA.
Problem Settings: VLMs Must Adaptively Read… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/RIO-Bench.STRIDE-QA-Dataset-Mini
STRIDE-QA-Dataset-Mini
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
⚠️ Note: STRIDE-QA-Dataset-Mini is provided as a preliminary version and does not fully match the format of the… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset-Mini.motor-medicina-legala-v2-storagemotor-nameplate
Motor Nameplate
A small image dataset of electric motor nameplates collected from public Google Images results, intended for tasks such as:
Training OCR / document-understanding models to extract nameplate text fields (manufacturer, HP, RPM, voltage, frame, etc.)
Image classification by manufacturer (ABB, Siemens, Baldor-Reliance, WEG, Hyundai, etc.)
Object detection (locating the plate on the motor body)
Few-shot learning / evaluation baselines for industrial vision tasks
The… See the full description on the dataset page: https://huggingface.co/datasets/sergiudanstan/motor-nameplate.STRIDE-QA-Bench
STRIDE-QA-Bench
STRIDE-QA-Bench provides a standardized benchmark for evaluating spatiotemporal reasoning of Vision-Language Models (VLMs) in autonomous driving.This HuggingFace repository provides the images and JSON files of the benchmark.
For detailed benchmark description and execution code, please refer to STRIDE-QA-Dataset (GitHub).
🗂️ Data Fields
The main data fields are as follows.
Field
Type
Description
question_id
str
Unique question ID.… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Bench.motor-nameplate-labels
Motor Nameplate — Auto-Labeled (OCR + Qwen2.5-1.5B)
Motor Nameplate — Auto-Labeled (OCR + Qwen2.5-1.5B)
A derived dataset that adds structured field annotations to the
sergiudanstan/motor-nameplate
image collection (286 motor-nameplate photos from Google Images).
The original upstream dataset ships images only. This repo adds an
auto-labeling pipeline output, ready for downstream fine-tuning of a
vision-language model for nameplate field extraction.… See the full description on the dataset page: https://huggingface.co/datasets/sergiudanstan/motor-nameplate-labels.Japanese-Heron-Bench
Japanese-Heron-Bench
Dataset Description
Japanese-Heron-Bench is a benchmark for evaluating Japanese VLMs (Vision-Language Models). We collected 21 images related to Japan. We then set up three categories for each image: Conversation, Detail, and Complex, and prepared one or two questions for each category. The final evaluation dataset consists of 102 questions. Furthermore, each image is assigned one of seven subcategories: anime, art, culture, food, landscape, landmark… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japanese-Heron-Bench.RACER-Mini
RACER-Mini
RACER (Rationale-Aware Captioning of Edge-Case Driving Scenarios) is a reasoning caption dataset designed for training vision-language-action (VLA) models in autonomous driving.
This repository provides approximately 1,000 samples, as a small subset of the RACER dataset. Each sample consists of a temporal sequence of front camera images, the ego vehicle’s future trajectory, and a corresponding reasoning caption.
For details, please refer to our techblog RACER:… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/RACER-Mini.motor-medicina-legala-v2-cartiSpatialRGPT-Bench-Extended
SpatialRGPT-Bench-Extended
SpatialRGPT-Bench-Extended is an extension of SpatialRGPT-Bench that incorporates driving scene images from Japan. It augments object-centric QA (questions about two objects within an image) with ego-centric QA (questions about the relationship between the ego and a single object in the image). Each QA category contains 466 QA pairs. For further details, please refer to our paper: https://arxiv.org/abs/2508.10427.
🔗 Related Links
Project… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/SpatialRGPT-Bench-Extended.africa-synth-motor-insurance-claims-all
African Motor Insurance Claims | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: not declared - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-motor-insurance-claims-all.motorcycle-dataset
Motorcycle Dataset
This repository contains a collection of 10,000 motorcycle images sourced from a larger motorbike dataset collection.
Dataset Structure
The dataset is organized into training and validation splits. The training split is further partitioned into three folders (train_part1, train_part2, and train_part3) to keep download sizes and commit structures manageable:
Total Images: 10,000
Folder Splits:
train_part1/: 3,000 images (unannotated, JPEG… See the full description on the dataset page: https://huggingface.co/datasets/shravya11/motorcycle-dataset.Wikipedia-Vision-JA
Dataset Card for Wikipedia-Vision-JA
Dataset description
The Wikipedia-Vision-JA is a Vision Language Model dataset generated from Japanese Wikipedia, containing 1.6M pairs of images, captions, and descriptions.
This dataset itself does not contain raw image data. Instead, an image_url is provided for each item.
Format
Wikipedia_Vision_JA.jsonl contains JSON-formatted rows with the following keys:
key: Unique JSON ID
caption: Short caption for the image… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Wikipedia-Vision-JA.SDv2-Count70-details-appended-motorcyclewisconsin-motorists-handbook-dataset
Wisconsin Motorists Handbook Dataset
Generated by DocParserEngine.
Field
Value
Documents
1
Records
1
Schema
full
Usage
from datasets import load_dataset
ds = load_dataset("Remixonwin/wisconsin-motorists-handbook-dataset")
aloha_balanced_v5_motor3imnet1k_motor_scooter_scootermotorcycle-datasets
