datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPSBench
GPSBench: Do Large Language Models Understand GPS Coordinates?
GPSBench is a benchmark dataset of 57,800 samples across 17 tasks for evaluating geospatial reasoning in Large Language Models (LLMs).
Paper: arXiv:2602.16105
Code: github.com/joey234/gpsbench
Leaderboard: gpsbench.github.io
Benchmark Structure
GPSBench is organized into two complementary evaluation tracks:
Pure GPS Track (9 tasks)
Coordinate manipulation without geographic knowledge:… See the full description on the dataset page: https://huggingface.co/datasets/joey234/GPSBench.GPSBench-10pct
🛰️ GPSBench 10% Stratified Sample
A deterministic coordinate-reasoning subset for lightweight GPS and geographic QA experiments
GPSBench 10% is a reproducible subset of joey234/GPSBench for quick experiments on GPS coordinate manipulation and coordinate-grounded geographic reasoning.
Quick Start ·
At a Glance ·
Files ·
Sampling ·
Citation
[!IMPORTANT]
This is a 10% stratified sample of the public GPSBench release, not… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/GPSBench-10pct.GPSBench
GPSBench: Do Large Language Models Understand GPS Coordinates?
GPSBench is a benchmark dataset of 57,800 samples across 17 tasks for evaluating geospatial reasoning in Large Language Models (LLMs).
Benchmark Structure
GPSBench is organized into two complementary evaluation tracks:
Pure GPS Track (9 tasks)
Coordinate manipulation without geographic knowledge:
Representation: Format Conversion, Coordinate System Transformation
Measurement: Distance Calculation… See the full description on the dataset page: https://huggingface.co/datasets/anonuser1234/GPSBench.
