CoolFace
Datasetpublic

chuckreynolds/wikimedia-enterprise-structured-contents-enwiki

enwiki_namespace_0 Structured Contents snapshot of enwiki_namespace_0 from the Wikimedia Enterprise API, converted to Parquet. Source Upstream: Wikimedia Enterprise Structured Contents API Snapshot identifier: enwiki_namespace_0 Format at source: .tar.gz containing sharded .ndjson Shards in this release: 3 Processing Downloaded the snapshot tarball from the Wikimedia Enterprise API. Streamed each .ndjson shard through a normalization pass:… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.

sourceHugging Facecc-by-sa-4.0updated 5mo agoView on Hugging Face
0likes31downloads
Dataset Card

enwikinamespace0

Structured Contents snapshot of enwiki_namespace_0 from the Wikimedia Enterprise API, converted to Parquet.

Source

Processing

  1. 1.Downloaded the snapshot tarball from the Wikimedia Enterprise API.
  2. 2.Streamed each .ndjson shard through a normalization pass:
  3. 3.JSON-encoded fields: sections, infoboxes, tables, and references[].metadata are stored as JSON-encoded strings. These fields either have recursive nesting (depth > 50) that exceeds Apache Arrow's C Data Interface limit, or are open-dict structures whose keys vary across articles. Decode with json.loads on read.
  4. 4.Canonicalised struct field ordering (alphabetic, recursive) so schemas match byte-for-byte across shards.
  5. 5.Wrote one Parquet file per shard (zstd level 9 compression, rowgroupsize tuned for HF streaming).
  6. 6.Unified per-shard schemas with pa.unify_schemas; pinned result to schema.json; re-cast every shard so embedded schemas are identical.

Loading

python
from datasets import load_dataset
import json

ds = load_dataset("chuckreynolds/wikimedia-enterprise-structured-contents-enwiki", split="train", streaming=True)
row = next(iter(ds))
print(row["name"], row["url"])

# JSON-encoded columns
sections = json.loads(row["sections"])
infoboxes = json.loads(row["infoboxes"])

Notes

  • —Fields stored as JSON strings: sections, infoboxes, tables, and references[].metadata. All other fields (references[], license[], version, event, etc.) retain their native Arrow struct/list types and are queryable without decoding.
  • —License passes through the upstream license for article text.