NoeFlandre/osm-polygon-website-tag
OSM Polygon Website Dataset OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files. At a glance Polygons 1,726,474 With extracted text 1,192,980 Words of text 407,685,655 Languages 397 Regional sources 386 / 386 Duplicate objects removed 104,927 Candidates rejected 868,905,743 Status In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.
OSM Polygon Website Dataset
OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files.
At a glance
Website text
Counts are unique (osm_type, osm_id) polygons -- 407,685,655 words in total. Regional overlap duplicates are removed globally.
Languages
Detected with GlotLID v3. Labels are script-aware language_Script codes; each row carries its top-1 probability.
Top 10 labels across both tags:
Sentences
Segmented with SaT, which covers 85 languages; anything else records unsupported_language. Segments do not rejoin into the source text -- use the *_text columns for that.
Mean sentences per segmented text: 38.2
Most common unsupported languages: kiu_Latn (12,502), hrv_Latn (4,715), zxx_Zzzz (4,031), anp_Deva (3,014), bos_Latn (1,902).
Geographic distribution
5,850 occupied H3 cells at resolution 3, covering 1,192,980 unique polygons with extracted text. Log colour scale, Natural Earth 1:110m backdrop.
Polygon geometry
Geodesic areas on the WGS84 ellipsoid, over every published polygon row. Full breakdown in `stats.json`.
Dataset bounding box: [-179.992396, -77.847199, 179.325994, 82.172551] (min lon, min lat, max lon, max lat).
Top website hostnames
Top contact:website hostnames
Dataset contents
polygons/*.parquet-- the polygons and their extracted text, one shard per source.analysis/*.parquet-- languages, sentences, hostnames, duplicates and per-source counts.stats.json,deduplication_summary.json-- full geometry and dedup numbers.manifests/-- source inventory and completion receipt.
Public polygon schema
Method
- Geometry assembled with libosmium; text extracted with Trafilatura and never truncated. Word counts are Unicode
\w+matches. - Text status is one of
absent,pending,success,empty,invalid_url,unsafe_url,fetch_error,extract_error. A source is enriched only when every status issuccessorabsent. Failed values retry on later resumptions; successful values are cached. - URLs resolving to anything other than a public IP are refused as
unsafe_url, before and after redirects. - Live metrics · source code
Provenance and license
Every artifact is bound by relative path, byte size and SHA-256 in the completion receipt. Map backdrop: Natural Earth 1:110m (public domain).
© OpenStreetMap contributors, under the ODbL 1.0 -- see the copyright page. Extracts from Geofabrik.
The website text is not covered by the ODbL. It is third-party content; rights stay with each source site. Check a site's terms before reusing its text.
Citation
Machine-readable metadata: `CITATION.cff`.
Flandre, Noé. OSM Polygon Website Tag Dataset. Hugging Face dataset
