CoolFace
Datasetpublic

yaakov/wikipedia-de-splits

Dataset Card for yaakov/wikipedia-de-splits Dataset Description The only goal of this dataset is to have random German Wikipedia articles at various dataset sizes: Small datasets for fast development and large datasets for statistically relevant measurements. For this purpose, I loaded the 2665357 articles in the test set of the pre-processed German Wikipedia dump from 2022-03-01, randomly permuted the articles and created splits of sizes 2**n: 1, 2, 4, 8, ....… See the full description on the dataset page: https://huggingface.co/datasets/yaakov/wikipedia-de-splits.

sourceHugging Facecc-by-sa-3.0updated 4y agoView on Hugging Face
0likes667downloads

No commit history came back for main. The revision may not exist, or the source declined the request.