datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MobilityBench
Note: This work is currently under review. The full dataset will be released progressively.
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
Paper | GitHub
MobilityBench is a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios. It is built from large-scale, anonymized mobility queries from Amap, organized with a comprehensive task taxonomy, and provides structured ground truth (required tool calls… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/MobilityBench.MobilityBench
Note: This work is currently under review. The full dataset will be released progressively.
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
Paper | GitHub
MobilityBench is a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios. It is built from large-scale, anonymized mobility queries from Amap, organized with a comprehensive task taxonomy, and provides structured ground truth (required tool calls… See the full description on the dataset page: https://huggingface.co/datasets/15985732665qqcom/MobilityBench.MobilityBench
Note: This work is currently under review. The full dataset will be released progressively.
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
Paper | GitHub
MobilityBench is a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios. It is built from large-scale, anonymized mobility queries from Amap, organized with a comprehensive task taxonomy, and provides structured ground truth (required tool calls… See the full description on the dataset page: https://huggingface.co/datasets/matriegardiegia/MobilityBench.
