dhruv-anand-aintech/en_wikidata_5M_entities
en_wikidata_5M_entities Hugging Face dataset card for a large, English-only Wikidata slice with optional Wikipedia links and Wikimedia Commons image URLs. One file, five million entities. Filename: en_wikidata_5M_entities.jsonl.gz TL;DR Format: JSON Lines, gzip-compressed (.jsonl.gz) Rows: 5,000,000 entities (one JSON object per line) Language: English labels/descriptions Fields: qid, label, description, enwiki_title, wikipedia_url, images (list of URLs)… See the full description on the dataset page: https://huggingface.co/datasets/dhruv-anand-aintech/en_wikidata_5M_entities.
en\wikidata\5M\_entities
Hugging Face dataset card for a large, English-only Wikidata slice with optional Wikipedia links and Wikimedia Commons image URLs.
One file, five million entities. Filename: en_wikidata_5M_entities.jsonl.gzTL;DR
- Format: JSON Lines, gzip-compressed (
.jsonl.gz) - Rows: 5,000,000 entities (one JSON object per line)
- Language: English labels/descriptions
- Fields:
qid,label,description,enwiki_title,wikipedia_url,images(list of URLs),has_image(bool),image_count(int) - License: CC0-1.0 for Wikidata-derived metadata. Wikipedia text is not included. Image URLs are provided; the images themselves remain under their respective licenses on Wikimedia Commons.
- Best way to load:
datasets.load_dataset(..., streaming=True)
Summary
This dataset provides a compact, streamable subset of Wikidata entities with English labels/descriptions, the (optional) matching English Wikipedia page title/URL, and zero or more Wikimedia Commons image URLs per entity. It is designed to be a simple foundation for knowledge indexing, entity linking, retrieval-augmented generation (RAG), and content moderation/research pipelines.
Files
en_wikidata_5M_entities.jsonl.gz— single shard containing exactly 5,000,000 lines (entities).
Each line is a standalone JSON object. The file is gzip-compressed for efficient storage and streaming.
Schema
Notes
- The dataset does not contain Wikipedia article text or image binaries—only metadata and links.
imagesare HTTP URLs served by Wikimedia; each media file has its own license and attribution requirements (see Licensing below).
Example rows
Below is an illustrative head (first 10 entities) from the file:
{'qid': 'Q31', 'label': 'Belgium', 'description': 'country in western Europe', 'enwiki_title': 'Belgium', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Belgium', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Belgique%20-%20Bruxelles%20-%20Grand-Place%20-%20C%C3%B4t%C3%A9%20nord-est.jpg?width=1024'], 'has_image': True, 'image_count': 1}
{'qid': 'Q8', 'label': 'happiness', 'description': 'positive emotional state', 'enwiki_title': 'Happiness', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Happiness', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Sweet%20Baby%20Kisses%20Family%20Love.jpg?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/Happiness%20.jpg?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/Felicidade%20A%20very%20happy%20boy.jpg?width=1024'], 'has_image': True, 'image_count': 3}
{'qid': 'Q24', 'label': 'Jack Bauer', 'description': 'character from the television series 24', 'enwiki_title': 'Jack_Bauer', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Jack_Bauer', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Kiefer%20Sutherland%20at%2024%20Redemption%20premiere%201%20%28headshot%29.jpg?width=1024'], 'has_image': True, 'image_count': 1}
{'qid': 'Q42', 'label': 'Douglas Adams', 'description': 'English science fiction writer and humorist (1952–2001)', 'enwiki_title': 'Douglas_Adams', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Douglas_Adams', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Douglas%20adams%20portrait.jpg?width=1024'], 'has_image': True, 'image_count': 1}
{'qid': 'Q1868', 'label': 'Paul Otlet', 'description': 'Belgian author, librarian and anti-colonial thinker', 'enwiki_title': 'Paul_Otlet', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Paul_Otlet', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Paul%20Otlet%20%C3%A0%20son%20bureau%20%28cropped%29.jpg?width=1024'], 'has_image': True, 'image_count': 1}
{'qid': 'Q2013', 'label': 'Wikidata', 'description': 'free multilingual online knowledge graph', 'enwiki_title': 'Wikidata', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Wikidata', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Wikidata-Homepage%2020220422%2021%2018%2054.png?width=1024'], 'has_image': True, 'image_count': 1}
{'qid': 'Q45', 'label': 'Portugal', 'description': 'country in Southwestern Europe', 'enwiki_title': 'Portugal', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Portugal', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Portugal%20-%20Location%20Map%20%282013%29%20-%20PRT%20-%20UNOCHA.svg?width=1024'], 'has_image': True, 'image_count': 1}
{'qid': 'Q51', 'label': 'Antarctica', 'description': 'polar continent', 'enwiki_title': 'Antarctica', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Antarctica', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Antarctica%206400px%20from%20Blue%20Marble.jpg?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/1%20Almirante%20Brown%20-%20Antarktische%20Halbinsel.jpg?width=1024'], 'has_image': True, 'image_count': 2}
{'qid': 'Q58', 'label': 'penis', 'description': 'primary sexual organ of male animals', 'enwiki_title': 'Penis', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Penis', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Penis%20asiatischer%20Elefant.JPG?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/Mallard%20with%20visible%20penis%20%28cropped%29.jpg?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/Callosobruchus%20analis%20penis.jpg?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/Penisstacheln.jpg?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/Mammalian%20penises.jpg?width=1024'], 'has_image': True, 'image_count': 5}
{'qid': 'Q68', 'label': 'computer', 'description': 'general-purpose device for performing arithmetic or logical operations', 'enwiki_title': 'Computer', 'wikipedia_url': 'https://en.wikipedia.org/wiki/Computer', 'images': ['https://commons.wikimedia.org/wiki/Special:FilePath/Summit%20%28supercomputer%29.jpg?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/Z3%20Deutsches%20Museum.JPG?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/Apple%20II%20Plus%2C%20Museum%20of%20the%20Moving%20Image.jpg?width=1024', 'https://commons.wikimedia.org/wiki/Special:FilePath/Lenovo%20G500s%20laptop-2903.jpg?width=1024'], 'has_image': True, 'image_count': 4}How to load (🤗 Datasets)
from datasets import load_dataset
# Stream the single shard without downloading it fully
# Replace <your-username> with your HF handle if hosted under your account.
ds = load_dataset(
"<your-username>/en_wikidata_5M_entities",
data_files="en_wikidata_5M_entities.jsonl.gz",
split="train",
streaming=True,
)
# Peek a few examples
for x in ds.take(5):
print(x)Filtering examples
# Only items that have at least one image
with_images = (ex for ex in ds if ex["has_image"]) # streaming-friendly
# Titles that look like persons (very naive heuristic)
person_like = (ex for ex in ds if ex["description"] and "(19" in ex["description"]) # birth years
# Convert a small slice to a list or pandas DataFrame
small = [x for _, x in zip(range(1000), ds)]Materialize to local Parquet (optional)
import json, gzip
import pandas as pd
rows = []
with gzip.open("en_wikidata_5M_entities.jsonl.gz", "rt", encoding="utf-8") as f:
for i, line in enumerate(f):
rows.append(json.loads(line))
if (i + 1) % 1_000_000 == 0:
print(f"Read {i+1:,} lines...")
pd.DataFrame(rows).to_parquet("en_wikidata_5M_entities.parquet")Tip: If you keep it on the Hub, prefer streaming=True for low-memory iteration.Quick local inspection
Use this snippet to confirm the file shape locally (size, head, and total count):
import gzip
import json
import os
file_path = "./en_wikidata_5M_entities.jsonl.gz"
# File size
file_size_bytes = os.path.getsize(file_path)
file_size_mb = file_size_bytes / (1024 * 1024)
print(f"File size: {file_size_mb:.2f} MB ({file_size_bytes:,} bytes)")
# Head
print("\n=== File Head (first 10 entities) ===")
with gzip.open(file_path, 'rt', encoding='utf-8') as f:
for i, line in enumerate(f):
if i >= 10:
break
try:
data = json.loads(line.strip())
print(data)
except json.JSONDecodeError:
print(f"Invalid JSON at line {i+1}: {line}")
# Count lines
print("\n=== Counting total entities ===")
count = 0
with gzip.open(file_path, 'rt', encoding='utf-8') as f:
for _ in f:
count += 1
print(f"Total entities in file: {count:,}")Intended use & examples
- Entity indexing / RAG: Build a vector index over
label + descriptionand optionally attach images downstream viaimages. - Entity linking / grounding: Map extracted Wikipedia URLs or QIDs back to structured metadata.
- Research / analytics: Explore distributions (e.g., which classes of entities have images).
- Moderation / safety: Because this mirrors real-world knowledge, mentions/URLs may include sensitive topics; see Safety below.
Minimal RAG sketch
from datasets import load_dataset
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
# Stream a subset (adjust N as needed)
ds = load_dataset("<your-username>/en_wikidata_5M_entities", data_files="en_wikidata_5M_entities.jsonl.gz", split="train", streaming=True)
model = SentenceTransformer("all-MiniLM-L6-v2")
texts, qids = [], []
for i, ex in enumerate(ds):
texts.append(f"{ex['label']}: {ex.get('description','')}")
qids.append(ex['qid'])
if i >= 50_000: # demo limit
break
emb = model.encode(texts, convert_to_numpy=True, show_progress_bar=True)
index = faiss.IndexFlatIP(emb.shape[1])
faiss.normalize_L2(emb)
index.add(emb)
# Query
q = "science fiction writer"
q_vec = model.encode([q], convert_to_numpy=True)
faiss.normalize_L2(q_vec)
D, I = index.search(q_vec, 5)
print([qids[i] for i in I[0]])Data quality & caveats
- Coverage: Only entities with English labels/descriptions are included. Non-English content is out of scope.
- Nulls:
enwiki_title/wikipedia_urlmay be missing (no linked English article). - Images: Many entities have zero images; some have multiple. URLs can change over time; always validate at use time.
- Duplicates:
qidis unique; however, labels/descriptions are not guaranteed unique. - Safety: Some entities/linked images cover sensitive topics (e.g., anatomy, violence, current events). Consumers should implement appropriate filtering for their use cases.
Licensing
- Wikidata metadata (labels, descriptions, QIDs, sitelinks) is released under CC0 1.0 (public domain dedication).
- Wikipedia: This dataset does not include article text; it may include titles/URLs to English Wikipedia, which are generally acceptable to reference.
- Wikimedia Commons images: The dataset only includes links to images. Do not mirror or redistribute image binaries without complying with their individual licenses (often CC BY/CC BY-SA/GFDL, sometimes with additional restrictions). Always provide attribution and share-alike when required. See each file’s page on Commons for precise terms.
If you repackage this dataset together with Wikipedia content or image binaries, you must comply with CC BY-SA 4.0 (for Wikipedia text) and the per-file licenses of images.
License for this packaging: CC0-1.0 (for the metadata arrangement only).
Versioning & changelog
- v1.0.0 — Initial release with 5,000,000 entities in a single gzip shard.
Citation
If you use this dataset, please cite Wikidata/Wikipedia and this packaging. A generic citation might be:
en\_wikidata\_5M\_entities: An English Wikidata slice with Wikipedia links and Commons image URLs (v1.0.0). Retrieved from the Hugging Face Hub. Underlying metadata from Wikidata (CC0 1.0). Wikipedia and Wikimedia Commons are trademarks of the Wikimedia Foundation.
You may also cite Wikidata:
Vrandečić, D., & Krötzsch, M. (2014). Wikidata: A free collaborative knowledgebase. Communications of the ACM, 57(10), 78–85.
Maintainers & contact
- Maintainer: Vijay Venkatesh Murugan / Institut Polytechnique de Paris
- Issues: Please open a ticket on the dataset’s Hugging Face page.
Disclaimer
This dataset is provided as is, without warranties. The maintainer is not affiliated with the Wikimedia Foundation. Use at your own risk and ensure downstream compliance with all applicable licenses and laws.
