CoolFace
Datasetpublic

yaakov/wikipedia-de-splits

Dataset Card for yaakov/wikipedia-de-splits Dataset Description The only goal of this dataset is to have random German Wikipedia articles at various dataset sizes: Small datasets for fast development and large datasets for statistically relevant measurements. For this purpose, I loaded the 2665357 articles in the test set of the pre-processed German Wikipedia dump from 2022-03-01, randomly permuted the articles and created splits of sizes 2**n: 1, 2, 4, 8, ....… See the full description on the dataset page: https://huggingface.co/datasets/yaakov/wikipedia-de-splits.

sourceHugging Facecc-by-sa-3.0updated 4y agoView on Hugging Face
0likes667downloads
Dataset Card

Dataset Card for yaakov/wikipedia-de-splits

Dataset Description

The only goal of this dataset is to have random German Wikipedia articles at various dataset sizes: Small datasets for fast development and large datasets for statistically relevant measurements.

For this purpose, I loaded the 2665357 articles in the test set of the pre-processed German Wikipedia dump from 2022-03-01, randomly permuted the articles and created splits of sizes 2**n: 1, 2, 4, 8, .... The split names are strings. The split 'all' contains all 2665357 available articles.

Dataset creation

This dataset has been created with the following script:

!apt install git-lfs !pip install -q transformers datasets

from huggingfacehub import notebooklogin notebook_login()

from datasets import loaddataset wikipediade = load_dataset("wikipedia", "20220301.de")['train']

shuffled = wikipedia_de.shuffle(seed=42)

from datasets import DatasetDict res = DatasetDict()

k, n = 0, 1 while n <= shuffled.num_rows: res[str(k)] = shuffled.select(range(n)) k += 1; n *= 2 res['all'] = shuffled

res.pushtohub('yaakov/wikipedia-de-splits')