CoolFace
Datasetpublic

chris951027/wikipedia

Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/chris951027/wikipedia.

sourceHugging Facecc-by-sa-3.0updated 9mo agoView on Hugging Face
0likes405downloads
settings

This repository belongs to chris951027 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namewikipedia
visibilitypublic
licencecc-by-sa-3.0
gatedno
ownerchris951027
Account settings
chris951027/wikipedia · CoolFace