Chenrongxin/STRIDE-QA-Dataset
STRIDE-QA Dataset ๐ฆ Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks. Category Description Object-centric Spatial QA Spatial relations between twoโฆ See the full description on the dataset page: https://huggingface.co/datasets/Chenrongxin/STRIDE-QA-Dataset.
STRIDE-QA Dataset
   
๐ฆ Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.


Data Format
This repository provides front camera images and JSONL annotations.
Directory Structure
STRIDE-QA-Dataset/
โโโ train/
โ โโโ images/
โ โโโ annotations/
โ โโโ ego_centric_spatial_qa/
โ โโโ ego_centric_spatiotemporal_qa/
โ โโโ object_centric_spatial_qa/
โโโ val/
โโโ images/
โโโ annotations/
โโโ ego_centric_spatial_qa/
โโโ ego_centric_spatiotemporal_qa/
โโโ object_centric_spatial_qa/Data Fields
The dataset is provided in .jsonl.gz. Each line in the JSONL file represents a single sample with the following fields:
Example Records
Object-centric Spatial QA
{
"id": "0004143dce0d1445",
"split": "val",
"category": "object_centric_spatial_qa",
"image": "0004143dce0d1445.jpg",
"conversations": [
{
"from": "human",
"value": "Is Region [3] taller than Region [0]?"
},
{
"from": "gpt",
"value": "No, Region [3] is not taller than Region [0]."
}
// ... more QA pairs ...
],
"image_info": {
"height": 1860,
"width": 2880,
"dataset": "STRIDE-QA",
"landmark": "outdoor",
"file_path": "0004143dce0d1445.jpg"
},
"token_info": {
"sample_token": "a2950b1d20375897e99b633a981e4dfe",
"sample_data_token": "4aba804904f3a1a2d3271b1fe8a6fc92"
},
"qa_info": [
{
"type": "qualitative",
"category": "tall_predicate",
"class": ["car", "large_vehicle"],
"token": {
"obj_A": {
"instance_token": "c7238672f1bfcc08929f99729f842948",
"sample_annotation_token": "a8189cdb160a525e2f798e739849fe8b"
},
"obj_B": {
"instance_token": "4941cec428bd732ff3e54ee43fbcc0e1",
"sample_annotation_token": "bac4d2c0ae8d029e406961655bf82ee9"
}
}
}
// ... more qa_info entries ...
],
"bbox": [
[0.0, 101.39, 881.81, 1657.32]
// ... more bounding boxes ...
],
"rle": [
{
"size": [1860, 2880],
"counts": "o7\\W12jhN^b0VW1b]OjhN^b0VW1..."
}
// ... more RLE masks ...
],
"region": [
[3, 0]
// ... more region tags ...
]
}Ego-centric Spatial QA
{
"id": "0004143dce0d1445",
"split": "val",
"category": "ego_centric_spatial_qa",
"image": "0004143dce0d1445.jpg",
"conversations": [
{
"from": "human",
"value": "Is the ego vehicle bigger than Region [0]?"
},
{
"from": "gpt",
"value": "Incorrect, the ego vehicle is not larger than Region [0]."
}
// ... more QA pairs ...
],
"image_info": {
"height": 1860,
"width": 2880,
"dataset": "STRIDE-QA",
"landmark": "outdoor",
"file_path": "0004143dce0d1445.jpg"
},
"token_info": {
"sample_token": "a2950b1d20375897e99b633a981e4dfe",
"sample_data_token": "4aba804904f3a1a2d3271b1fe8a6fc92"
},
"qa_info": [
{
"type": "qualitative",
"category": "ego_big_predicate",
"class": ["ego_vehicle", "large_vehicle"],
"token": {
"ego": {
"instance_token": null,
"sample_annotation_token": null
},
"other": {
"instance_token": "4941cec428bd732ff3e54ee43fbcc0e1",
"sample_annotation_token": "bac4d2c0ae8d029e406961655bf82ee9"
}
}
}
// ... more qa_info entries ...
],
"bbox": [
[0.0, 101.39, 881.81, 1657.32]
// ... more bounding boxes ...
],
"rle": [
{
"size": [1860, 2880],
"counts": "o7\\W12jhN^b0VW1b]OjhN^b0VW1..."
}
// ... more RLE masks ...
],
"region": [
[99, 0]
// ... more region tags ...
]
}Ego-centric Spatio-temporal QA
{
"id": "0004143dce0d1445",
"split": "val",
"category": "ego_centric_spatiotemporal_qa",
"images": [
"31ee9bcfc3ed9dec.jpg",
"f6db1dee9bddc641.jpg",
"c76385f109139ea8.jpg",
"0004143dce0d1445.jpg"
],
"conversations": [
{
"from": "human",
"value": "Can you give me an estimate of the distance between the ego vehicle and Region [0]?"
},
{
"from": "gpt",
"value": "The ego vehicle and Region [0] are 3.27 meters apart from each other."
}
// ... more QA pairs ...
],
"image_info": {
"height": 1860,
"width": 2880,
"dataset": "STRIDE-QA",
"landmark": "outdoor",
"file_path": "0004143dce0d1445.jpg"
},
"token_info": {
"current": {
"sample_token": "a2950b1d20375897e99b633a981e4dfe",
"sample_data_token": "4aba804904f3a1a2d3271b1fe8a6fc92"
},
"prev_1": {
"sample_token": "93f455351c376fb32dc6315c3d0c53e2",
"sample_data_token": "0ec22aef72d02da0f26763b22ec7afe3"
}
// ... prev_2, prev_3 ...
},
"qa_info": [
{
"category": "target_distance_t0",
"target": {
"instance_token": "4941cec428bd732ff3e54ee43fbcc0e1",
"sample_annotation_token": "bac4d2c0ae8d029e406961655bf82ee9",
"class": "large_vehicle",
"region_id": 0
},
"state": {
"ego": {
"speed_mps": 4.38
},
"target": {
"speed_mps": 6.42,
"distance_m": 3.27,
"bearing_deg": 18.58,
"clock_position": 11
},
"after_t_seconds": 0
}
}
// ... more qa_info entries ...
],
"bbox": {
"current": [
[0.0, 101.39, 881.81, 1657.32]
// ... more bounding boxes ...
],
"prev_1": [
[0.0, 0.0, 802.30, 1860.0]
// ... more bounding boxes ...
]
// ... prev_2, prev_3 ...
},
"rle": {
"current": [
{
"size": [1860, 2880],
"counts": "o7\\W12jhN^b0VW1b]OjhN^b0VW1..."
}
// ... more RLE masks ...
]
// ... prev_1, prev_2, prev_3 ...
}
}๐ License
STRIDE-QA-Dataset is released under the CC BY-NC-SA 4.0.
๐ Citation
@misc{strideqa2025,
title={STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes},
author={Keishi Ishihara and Kento Sasaki and Tsubasa Takahashi and Daiki Shiono and Yu Yamaguchi},
year={2025},
eprint={2508.10427},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2508.10427},
}๐ค Acknowledgements
This dataset was developed as part of the project JPNP20017, which is subsidized by the New Energy and Industrial Technology Development Organization (NEDO), Japan.
We would like to acknowledge the use of the following open-source repositories:
- SpatialRGPT for building dataset generation pipeline
- SAM 2.1 for segmentation mask generation
- dashcam-anonymizer for anonymization
๐ Privacy Protection
To ensure privacy protection, human faces and license plates in the images were anonymized using the Dashcam Anonymizer.
