wikidata
Datasets
All datasets matching “wikidata”wikidata
Wikidata Entities Connected to Wikipedia
This dataset is a multilingual, JSON-formatted version of the Wikidata dump from May 7, 2026. It contains 73,769,737 entities after filtering out scholarly articles from the original 120,182,414 entity dump.
Curated by: Jonathan Fraine & Philippe Saadé, Wikimedia Deutschland
Funded by: Wikimedia Deutschland
Language(s) (NLP): All Wikidata Languages
License: CC0-1.0
Dataset Structure
Each row in this dataset represents a… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/wikidata.osm-polygon-wikidata-only
OSM Polygon Wikidata, Wikipedia and Wikivoyage
OSM polygons carrying wikidata=*, enriched with multilingual Wikipedia and Wikivoyage documents. The published tables preserve regional records and provenance.
Source code: GitHub repository.
Dataset snapshot
Metric
Value
Polygon rows across regional extracts
1,184,110
Unique polygon identities (osm_type, osm_id)
1,157,841
Polygons with successful non-empty text (unique OSM identities)
650,663… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-wikidata-only.WikidataLabels
Wikidata Labels
Large parallel corpus for machine translation
Entity label data extracted from Wikidata (2022-01-03), filtered for item entities only
Only download the languages you need with datasets>=2.14.0
Similar dataset: https://huggingface.co/datasets/wmt/wikititles (18 Wikipedia titles pairs instead of all Wikidata entities)
Dataset Details
Dataset Sources
Wikidata JSON dump (wikidata-20220103-all.json.gz)… See the full description on the dataset page: https://huggingface.co/datasets/rayliuca/WikidataLabels.Wikidata_Vectors_0.2
Wikidata Entity Embeddings 0.2
Dataset Summary
Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata.
The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.wikidata-extraction
Wikidata Extraction
This dataset contains all RDF triples extracted from the latest Wikidata, converted from the N-Triples format to Parquet.
The data originates from Wikidata, a free and open knowledge base that acts as central storage for structured data used by Wikipedia and other Wikimedia projects. The source file is the "truthy" N-Triples dump (latest-truthy.nt.bz2), which contains only the current, non-deprecated statements.
The code to extract this data is available at… See the full description on the dataset page: https://huggingface.co/datasets/piebro/wikidata-extraction.osm-polygon-wikidata-and-wikipedia
OSM Polygon Wikidata + Wikipedia, V2
V2 builds on the V1 Wikidata-only dataset with multilingual Wikipedia and Wikivoyage text. It also retains valid multilingual wikipedia=* references, including polygons without a Wikidata QID.
Source code: GitHub repository.
Dataset snapshot
Metric
Value
Polygon rows across regional extracts
1,259,424
Unique polygon identities (osm_type, osm_id)
1,188,854
Polygons with successful non-empty text (unique OSM… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-wikidata-and-wikipedia.
