yaakov/wikipedia-de-splits
Dataset Card for yaakov/wikipedia-de-splits Dataset Description The only goal of this dataset is to have random German Wikipedia articles at various dataset sizes: Small datasets for fast development and large datasets for statistically relevant measurements. For this purpose, I loaded the 2665357 articles in the test set of the pre-processed German Wikipedia dump from 2022-03-01, randomly permuted the articles and created splits of sizes 2**n: 1, 2, 4, 8, ....… See the full description on the dataset page: https://huggingface.co/datasets/yaakov/wikipedia-de-splits.
Dataset Card for yaakov/wikipedia-de-splits
Dataset Description
The only goal of this dataset is to have random German Wikipedia articles at various dataset sizes: Small datasets for fast development and large datasets for statistically relevant measurements.
For this purpose, I loaded the 2665357 articles in the test set of the pre-processed German Wikipedia dump from 2022-03-01, randomly permuted the articles and created splits of sizes 2**n: 1, 2, 4, 8, .... The split names are strings. The split 'all' contains all 2665357 available articles.
Dataset creation
This dataset has been created with the following script:
!apt install git-lfs !pip install -q transformers datasets
from huggingfacehub import notebooklogin notebook_login()
from datasets import loaddataset wikipediade = load_dataset("wikipedia", "20220301.de")['train']
shuffled = wikipedia_de.shuffle(seed=42)
from datasets import DatasetDict res = DatasetDict()
k, n = 0, 1 while n <= shuffled.num_rows: res[str(k)] = shuffled.select(range(n)) k += 1; n *= 2 res['all'] = shuffled
res.pushtohub('yaakov/wikipedia-de-splits')
