CoolFace
Datasetpublic

answerdotai/enwiki

English Wikipedia as clean md This dataset is a cleaned, structurally faithful approximation of the English Wikipedia article corpus in Answer.AI's canonical md dialect. It was produced from the Wikimedia dump dated 20260901 by Answer.AI's wiki2dataset pipeline. It is designed for language-model training and for agent/RAG systems. The articles configuration provides complete documents for continued pretraining, corpus analysis, rechunking, and task-specific dataset creation. The… See the full description on the dataset page: https://huggingface.co/datasets/answerdotai/enwiki.

sourceHugging Facecc-by-sa-4.0updated 8d agoView on Hugging Face
5likes509downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
answerdotai/enwiki · CoolFace