matriegardiegia/MobilityBench
Note: This work is currently under review. The full dataset will be released progressively. MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios Paper | GitHub MobilityBench is a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios. It is built from large-scale, anonymized mobility queries from Amap, organized with a comprehensive task taxonomy, and provides structured ground truth (required tool… See the full description on the dataset page: https://huggingface.co/datasets/matriegardiegia/MobilityBench.
Note: This work is currently under review. The full dataset will be released progressively.
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
MobilityBench is a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios. It is built from large-scale, anonymized mobility queries from Amap, organized with a comprehensive task taxonomy, and provides structured ground truth (required tool calls + verifiable evidence). All tool calls are executed in a deterministic replay sandbox for reproducible, multi-dimensional evaluation.
Scale & Coverage: 100,000 episodes across 22 countries and 350+ cities (including metropolitan areas), with a long-tailed geographic distribution.
Scenario Distribution (11 intents)
- 36.6% Basic Information Retrieval
- 9.6% Route-Dependent Information Retrieval
- 42.5% Basic Route Planning
- 11.3% Preference-Constrained Route Planning
Data Format
Sample Data (5 Examples)
Citation
If you use this dataset in your research, please cite the following paper:
@article{song2026mobilitybench,
title={MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios},
author={Song, Zhiheng and Zhang, Jingshuai and Qin, Chuan and Wang, Chao and Chen, Chao and Xu, Longfei and Liu, Kaikui and Chu, Xiangxiang and Zhu, Hengshu},
journal={arXiv preprint arXiv:2602.22638},
year={2026}
}