NoeFlandre/osm-polygon-website-tag-eunis
OSM Polygon Website Dataset OpenStreetMap closed ways and polygon relations carrying a non-empty website OR contact:website tag, with full main-page text extracted using Trafilatura. Every statistic below is regenerated from the current upload-acknowledged Parquet artifacts. Snapshot Metric Value What it means Snapshot status In progress Current published snapshot Regional PBFs included 386 / 386 Published source shards / expected source PBFs… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag-eunis.
OSM Polygon Website Dataset
OpenStreetMap closed ways and polygon relations carrying a non-empty website OR contact:website tag, with full main-page text extracted using Trafilatura. Every statistic below is regenerated from the current upload-acknowledged Parquet artifacts.
Snapshot
Website text
Website-text table counts are unique (osm_type, osm_id) identities across regional rows; regional overlap duplicates are removed globally.
Unique polygons with extracted text: 1,192,980 Counts unique (osm_type, osm_id) polygons across regional rows when any copy has successful, trimmed non-empty website or contact:website text; regional overlap duplicates removed globally. Combined extracted words: 407,685,655
Languages
Detected with GlotLID v3 on successfully extracted text; labels are exact script-aware language_Script codes with a top-1 probability column.
Top 10 labels across both tags:
Sentences
Extracted text is segmented with SaT for the 85 languages the segmenter covers; text in any other detected language records unsupported_language instead of sentences. Segments carry their own trailing spaces but not the line breaks that separated them, so joining them does not reproduce the source text; the full text stays in the *_text columns.
Mean sentences per segmented text: 38.2
Polygon geometry
Surface and shape statistics computed over every published polygon row from the area_m2, bbox, and geometry columns. Areas are geodesic on the WGS84 ellipsoid. Population scope: published polygon rows. The complete breakdown is published as `stats.json`.
Dataset bounding box: [-179.992396, -77.847199, 179.325994, 82.172551] (min lon, min lat, max lon, max lat).
Geographic distribution
H3 resolution 3 contains 5,850 occupied cells across 1,192,980 unique polygons with successfully extracted, non-empty website or contact:website text, globally deduplicated by (osm_type, osm_id); regional overlap duplicates removed globally. The color scale is logarithmic, counts are absolute, and a Natural Earth 1:110m land backdrop provides geographic context.
Links
Live metrics: Trackio dashboard; it shows this frozen dataset snapshot. Code and README: GitHub repository and README.
Top website hostnames
Top contact:website hostnames
Methodology and quality
Geometry is assembled with libosmium. Full main text is extracted independently for both website tags with Trafilatura and is not truncated. Word counts are Python Unicode \w+ matches.
Text statuses are absent, pending, success, empty, invalid_url, unsafe_url, fetch_error, or extract_error. A source is enriched only when every status is success or absent. Failed values retry on later resumptions; successful values are cached.
A URL is marked unsafe_url when its hostname, or any redirect target, does not resolve exclusively to globally routable public IP addresses. Localhost, private, reserved, multicast, and unspecified targets are blocked. Unsupported schemes and URLs containing credentials are classified as invalid_url; redirect limits, timeouts, oversized responses, and unsupported content types are recorded as fetch_error.
Dataset contents
polygons/*.parquet: the public polygon split, one shard per source PBF.analysis/*.parquet: detailed overlap, provenance, hostname, duplicate, conflict, and per-source statistics.deduplication_summary.json: counts and tag-conflict totals from the global canonicalization pass.manifests/: source inventory, upload checkpoints, and completion receipt.
Public polygon schema
Provenance and license
Source filename, byte size, and nanosecond modification time are recorded before processing. The completion receipt binds finalized artifacts by relative path, byte size, and SHA-256.
The map backdrop uses Natural Earth 1:110m Admin-0 country geography, distributed in the source tree under its public-domain terms.
© OpenStreetMap contributors. OpenStreetMap data is available under the Open Database License (ODbL) 1.0; see the OpenStreetMap copyright and attribution page. Regional PBF extracts are provided by Geofabrik.
Website text is third-party content, separate from the OSM data, and is not covered by the ODbL. This dataset asserts no license for that text and grants no additional reuse rights: copyright and licensing conditions remain with each source website. Check the source site's terms or license before using or redistributing extracted text.
Citation
If you use this dataset, please cite it using the machine-readable metadata in `CITATION.cff`. GitHub and the Hugging Face dataset page can then display the citation directly.
Flandre, Noé. OSM Polygon Website Tag Dataset. Hugging Face dataset
