schema-org
schema-dot-org
Geolocated text from the Web Data Commons schema.org GeoCoordinates subset
12,427,530 geolocated text records, extracted from the class-specific
GeoCoordinates subset of the Web Data Commons schema.org data set series
(release 2024-12). Each record pairs one coordinate pair published on a web page
with the text published next to it on that same page.
Each source stream is deduplicated by host-local runs: a coordinate-and-name
pair is kept once per contiguous host run. A host… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/schema-dot-org.Schemaorg
Dataset Card for Schemaorg
This dataset is a collection of Mixed-hop Prediction datasets created from Schema.org's subsumption hierarchy (TBox) for evaluating hierarchy embedding models. It is an evaluation-only dataset consisting of just validation and test splits.
Mixed-hop Prediction: This task aims to evaluate the model’s capability in determining the existence of subsumption relationships between arbitrary entity pairs, where the entities are not necessarily seen during… See the full description on the dataset page: https://huggingface.co/datasets/Hierarchy-Transformers/Schemaorg.wim-schema-org-wiki-articles
Dutch Wikipedia Aligned Articles aligned with Schema.org Classes
Dataset Version: 1.0 (2025-06-04)Point of Contact: UWV Netherlands (UWV organization on Hugging Face)License: CC BY-SA 4.0Dataset: UWV/wim_schema_org_wiki_articles
Dataset Description
This dataset provides alignments between Schema.org classes and relevant Dutch Wikipedia articles. Each Schema.org class from a processed subset is linked to up to 20 distinct Wikipedia articles, including their full text, a… See the full description on the dataset page: https://huggingface.co/datasets/UWV/wim-schema-org-wiki-articles.schema-org
Schema.org
Dataset Description
Vocabulary schemas for structured data on the web
Original Source: https://schema.org/version/latest/schemaorg-current-https.ttl
Dataset Summary
This dataset contains RDF triples from Schema.org converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally turtle, converted to HuggingFace Dataset
Size: 0.01 GB (extracted)
Entities: ~2K types
Triples: ~15K
Original License:
CC BY-SA 3.0… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/schema-org.text_to_schema.org_json-ldschema-org-v1
Schema.org
Dataset Description
Vocabulary schemas for structured data on the web
Original Source: https://schema.org/version/latest/schemaorg-current-https.ttl
Dataset Summary
This dataset contains RDF triples from Schema.org converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally turtle, converted to HuggingFace Dataset
Size: 0.01 GB (extracted)
Entities: ~2K types
Triples: ~15K
Original License: CC BY-SA 3.0… See the full description on the dataset page: https://huggingface.co/datasets/Dabbu19/schema-org-v1.
