CoolFace
Datasetpublic

answerdotai/simplewiki

Simple English Wikipedia as clean md This dataset is a cleaned, structurally faithful approximation of the Simple English Wikipedia article corpus in Answer.AI's canonical md dialect. It was produced from the Wikimedia dump dated 20260901 by Answer.AI's wiki2dataset pipeline. It is designed for language-model training and for agent/RAG systems. The articles configuration provides complete documents for continued pretraining, corpus analysis, rechunking, and task-specific dataset… See the full description on the dataset page: https://huggingface.co/datasets/answerdotai/simplewiki.

sourceHugging Facecc-by-sa-4.0updated 7d agoView on Hugging Face
3likes366downloads
14 commits on main
5cdac2b7d ago

Update simplewiki to 20260901 (#9)

jph00
4c57ea37d ago

Update simplewiki to 20260901 (#8)

jph00
4a1fd3028d ago

Update simplewiki to 20260801 (#7)

jph00
bf7e05728d ago

Update simplewiki to 20260801 (#6)

jph00
b41fb6928d ago

Update simplewiki to 20260801 (#5)

jph00
ceabdbb28d ago

Update simplewiki to 20260801 (#4)

jph00
c4eee7828d ago

Update simplewiki to 20260801 (#3)

jph00
340bd0828d ago

Update simplewiki to 20260801 (#2)

jph00
86b625728d ago

Record clean source provenance

jph00
7d38bea29d ago

Simplify dataset processing description

jph00
a9da6a529d ago

Refresh Simple English Wikipedia dataset

jph00
dbbc8df29d ago

Upload folder using huggingface_hub

jph00
c620e0929d ago

Upload folder using huggingface_hub

jph00
f93fa9d29d ago

initial commit

jph00