datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
actor
GPS-Bench Actor Layer
Who each AI-governance instrument reaches, what each reached actor did, and what followed.
Companion to GPS-bench/gps-bench-ai-bills,
which holds the instruments and their text.
Join on bill_key. This dataset keys instruments by bill_key (us-117-hr-4346);
the companion keys the same instrument by instrument_id (gps:us-117-hr-4346). The
mapping is exactly the gps: prefix and it is total in both directions.
instrument_id = "gps:" + bill_key… See the full description on the dataset page: https://huggingface.co/datasets/GPS-bench/actor.gps-bench-ai-bills
GPS-Bench AI Legal Instruments Dataset
An evidence-backed corpus of global AI-related laws, regulations, policies, guidelines, standards, and strategic instruments. Each canonical record retains its official title, issuing authority, normalized jurisdiction, legal status, dates, primary-source citations, and full verified text. Every non-English instrument keeps its source-language original as the authoritative text, with an English translation beside it produced by a single… See the full description on the dataset page: https://huggingface.co/datasets/GPS-bench/gps-bench-ai-bills.GPSBench
GPSBench: Do Large Language Models Understand GPS Coordinates?
GPSBench is a benchmark dataset of 57,800 samples across 17 tasks for evaluating geospatial reasoning in Large Language Models (LLMs).
Paper: arXiv:2602.16105
Code: github.com/joey234/gpsbench
Leaderboard: gpsbench.github.io
Benchmark Structure
GPSBench is organized into two complementary evaluation tracks:
Pure GPS Track (9 tasks)
Coordinate manipulation without geographic knowledge:… See the full description on the dataset page: https://huggingface.co/datasets/joey234/GPSBench.Compliance-to-CodeGPSBench
GPSBench: Do Large Language Models Understand GPS Coordinates?
GPSBench is a benchmark dataset of 57,800 samples across 17 tasks for evaluating geospatial reasoning in Large Language Models (LLMs).
Benchmark Structure
GPSBench is organized into two complementary evaluation tracks:
Pure GPS Track (9 tasks)
Coordinate manipulation without geographic knowledge:
Representation: Format Conversion, Coordinate System Transformation
Measurement: Distance Calculation… See the full description on the dataset page: https://huggingface.co/datasets/anonuser1234/GPSBench.gpsr-dataset
GPSR instructions
Overview
This dataset was created for research on domestic robotics, particularly in the context of planning and natural language understanding. It includes annotated natural language commands and corresponding structured representations used for robot task execution. The dataset was developed as part of the research presented in "End-to-End Robot Task Planning from Transcriptions of Voice Commands," focusing on evaluating various adaptation strategies… See the full description on the dataset page: https://huggingface.co/datasets/certafonso/gpsr-dataset.GPSBench-10pct
🛰️ GPSBench 10% Stratified Sample
A deterministic coordinate-reasoning subset for lightweight GPS and geographic QA experiments
GPSBench 10% is a reproducible subset of joey234/GPSBench for quick experiments on GPS coordinate manipulation and coordinate-grounded geographic reasoning.
Quick Start ·
At a Glance ·
Files ·
Sampling ·
Citation
[!IMPORTANT]
This is a 10% stratified sample of the public GPSBench release, not… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/GPSBench-10pct.DeKeyNLUGP-SimQ-LLaMa2-MistralGP-SimQ-LLaMa2-LLaMa2GP-SimQ-Mistral-LLaMa2GP-SimQ-Mistral-Mistral
