datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
osm-tokyo23-src-2026-08
osm-tokyo23-src-2026-08
A frozen cut of OpenStreetMap covering the 23 special wards of Tokyo, taken
from the planet file of 2026-08-31, together with everything needed to
rebuild the databases it was measured in.
The point is the freezing. A question about a city has an answer only against
a stated snapshot, and an answer computed today against the live API is not
reproducible tomorrow. Here the snapshot is one file with a checksum, and the
tools that read it are pinned by… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-src-2026-08.osm-tokyo23-qa-2026-08
osm-tokyo23-qa-2026-08
All 215 answers to osm-tokyo23-questions,
computed against the frozen extract in
osm-tokyo23-src-2026-08.
The questions ship without answers on purpose: an answer belongs to a
particular extract on a particular day. This is one such day.
Every answer carries the query that produced it. Not a citation of one, the
text of one, for each engine that was asked. An answer here is meant to be
recomputed rather than believed, and the thing that makes that possible… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-qa-2026-08.osm-tokyo23-questions
osm-tokyo23-questions
Questions a person would ask about the twenty-three special wards of Tokyo,
written to be answered from OpenStreetMap. 215 of them, in twenty kinds.
No answers. The set is questions and nothing else. Answers belong to a
particular extract on a particular day, and a question carrying its own answer
stops being a question. What is recorded instead is which data each question
was checked against, so that someone can compute the answers and say what they… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-questions.
