alexandrainst/wiki40b-da
Dataset Card for "wiki40b-da" Dataset Summary This dataset is an upload of the Danish part of the Wiki40b dataset, being a cleaned version of a dump of Wikipedia. The dataset is identical in content to this dataset on the Hugging Face Hub, but that one requires both apache_beam, tensorflow and mwparserfromhell, which can lead to dependency issues since these are not compatible with several newer packages. The training, validation and test splits are the original… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/wiki40b-da.
Dataset Card for "wiki40b-da"
Dataset Description
- Point of Contact: Dan Saattrup Nielsen
- Size of downloaded dataset files: 150.57 MB
- Size of the generated dataset: 246.09 MB
- Total amount of disk used: 396.66 MB
Dataset Summary
This dataset is an upload of the Danish part of the Wiki40b dataset, being a cleaned version of a dump of Wikipedia.
The dataset is identical in content to this dataset on the Hugging Face Hub, but that one requires both apache_beam, tensorflow and mwparserfromhell, which can lead to dependency issues since these are not compatible with several newer packages.
The training, validation and test splits are the original ones.
Languages
The dataset is available in Danish (da).
Dataset Structure
Data Instances
- Size of downloaded dataset files: 150.57 MB
- Size of the generated dataset: 246.09 MB
- Total amount of disk used: 396.66 MB
An example from the dataset looks as follows.
{
'wikidata_id': 'Q17341862',
'text': "\n_START_ARTICLE_\nÆgyptiske tekstiler\n_START_PARAGRAPH_\nTekstiler havde mange (...)",
'version_id': '9018011197452276273'
}Data Fields
The data fields are the same among all splits.
wikidata_id: astringfeature.text: astringfeature.version_id: astringfeature.
Dataset Statistics
There are 109,486 samples in the training split, 6,173 samples in the validation split and 6,219 in the test split.
Document Length Distribution

Additional Information
Dataset Curators
Dan Saattrup Nielsen from the The Alexandra Institute uploaded it to the Hugging Face Hub.
Licensing Information
The dataset is licensed under the CC-BY-SA license.
