CoolFace
Datasetpublic

answerdotai/simplewiki

Simple English Wikipedia as clean md This dataset is a cleaned, structurally faithful approximation of the Simple English Wikipedia article corpus in Answer.AI's canonical md dialect. It was produced from the Wikimedia dump dated 20260901 by Answer.AI's wiki2dataset pipeline. It is designed for language-model training and for agent/RAG systems. The articles configuration provides complete documents for continued pretraining, corpus analysis, rechunking, and task-specific dataset… See the full description on the dataset page: https://huggingface.co/datasets/answerdotai/simplewiki.

sourceHugging Facecc-by-sa-4.0updated 7d agoView on Hugging Face
3likes366downloads

answerdotai/simplewiki · main · files are served by the source, never re-hosted here