chuckreynolds/wikimedia-enterprise-structured-contents-enwiki
enwiki_namespace_0 Structured Contents snapshot of enwiki_namespace_0 from the Wikimedia Enterprise API, converted to Parquet. Source Upstream: Wikimedia Enterprise Structured Contents API Snapshot identifier: enwiki_namespace_0 Format at source: .tar.gz containing sharded .ndjson Shards in this release: 3 Processing Downloaded the snapshot tarball from the Wikimedia Enterprise API. Streamed each .ndjson shard through a normalization pass:… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.
031
enwikinamespace0
Structured Contents snapshot of enwiki_namespace_0 from the Wikimedia Enterprise API, converted to Parquet.
Source
- Upstream: Wikimedia Enterprise Structured Contents API
- Snapshot identifier:
enwiki_namespace_0 - Format at source:
.tar.gzcontaining sharded.ndjson - Shards in this release: 3
Processing
- Downloaded the snapshot tarball from the Wikimedia Enterprise API.
- Streamed each
.ndjsonshard through a normalization pass: - JSON-encoded fields:
sections,infoboxes,tables, andreferences[].metadataare stored as JSON-encoded strings. These fields either have recursive nesting (depth > 50) that exceeds Apache Arrow's C Data Interface limit, or are open-dict structures whose keys vary across articles. Decode withjson.loadson read. - Canonicalised struct field ordering (alphabetic, recursive) so schemas match byte-for-byte across shards.
- Wrote one Parquet file per shard (zstd level 9 compression, rowgroupsize tuned for HF streaming).
- Unified per-shard schemas with
pa.unify_schemas; pinned result toschema.json; re-cast every shard so embedded schemas are identical.
Loading
from datasets import load_dataset
import json
ds = load_dataset("chuckreynolds/wikimedia-enterprise-structured-contents-enwiki", split="train", streaming=True)
row = next(iter(ds))
print(row["name"], row["url"])
# JSON-encoded columns
sections = json.loads(row["sections"])
infoboxes = json.loads(row["infoboxes"])Notes
- Fields stored as JSON strings:
sections,infoboxes,tables, andreferences[].metadata. All other fields (references[],license[],version,event, etc.) retain their native Arrow struct/list types and are queryable without decoding. - License passes through the upstream license for article text.
