CoolFace
Datasetpublic

NoeFlandre/landuse-sentence-relevance-golden-human-set

Land-use sentence relevance golden human set This release contains the final 300-row V3 benchmark in English plus one parallel CSV for each of the 84 non-English project-provided sat-3l-sm language codes. There are 85 language files in total. Files Every file is at data/translations/<iso>/v3-final-<iso>.csv. The nine columns are: sentence, label, polygon_name, h3_cell, latitude, longitude, source, region, source_url. The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.

sourceHugging Faceotherupdated 7d agoView on Hugging Face
0likes210downloads
Dataset Card

Land-use sentence relevance golden human set

This release contains the final 300-row V3 benchmark in English plus one parallel CSV for each of the 84 non-English project-provided sat-3l-sm language codes. There are 85 language files in total.

Files

Every file is at data/translations/<iso>/v3-final-<iso>.csv. The nine columns are:

sentence, label, polygon_name, h3_cell, latitude, longitude, source, region, source_url.

The Dataset Viewer exposes these files as 85 selectable subsets, one per ISO language code (en is the default). Each subset has a single train split. Hugging Face caps a configuration at three splits, so languages are modelled as configs rather than splits, per the manual dataset configuration docs.

Only sentence is translated. Labels are the final human-reviewed English decisions and all geographic/source metadata is retained in the same row order in every language file. data/translations/en/v3-final-en.csv is an exact copy of the final English benchmark.

Geographical distribution

[image]

The map shows the coordinates of all 300 benchmark records, colored by source (100 Description, 100 Website, 100 Wikipedia). The same coordinates are retained across all language files. The background is the pinned Natural Earth 110m land layer in an equirectangular projection; this is a coverage visualization, not a population or source-density estimate.

This PNG is generated deterministically from the final English benchmark by `scripts/build_v3_world_map.py` with the pinned Matplotlib 3.11.1 Agg renderer. The generator validates the 300 benchmark coordinates, uses the committed Natural Earth 110m land vectors, performs no network access, and writes this card asset at assets/v3-world-distribution.png. The basemap is Natural Earth 1:110m physical land, which is in the public domain ("Made with Natural Earth"); the exact snapshot bytes are vendored in the project repository at data/benchmark/v3/assets/ne-110m-land.geojson, with the checksum documented in data/benchmark/v3/assets/ne-110m-land-manifest.json.

English benchmark provenance

The English benchmark was built from a completed human reference. An independent GPT assessment was used to identify 42 disagreements, those cases were reassessed by a human, and the final file was rebuilt in reference order. Only 16 labels changed after review (13 no to yes, 3 yes to no). The final distribution is 160 yes and 140 no, with 100 rows per source.

The repository documents the complete chain in `docs/v3-final-benchmark.md` and the multilingual release method in `docs/v3-translations.md`.

Translation provenance

Translation was performed by gpt-5.6-luna-max: the available gpt-5.6-luna model at max reasoning, executed through the Codex harness. The English label is deliberately not reclassified after translation. The translation manifest records the source hash, model/harness attribution, row counts, language inventory, and SHA-256 for every file.

Licensing and attribution

The project code and original benchmark curation are Apache-2.0. The included upstream-derived material retains its own terms: OpenStreetMap-derived fields are subject to the ODbL, and Wikipedia-derived text is subject to CC BY-SA 4.0. See the repository `NOTICE` and licensing documentation.