LTCB/enwik8
The dataset is based on the Hutter Prize (http://prize.hutter1.net) and contains the first 10^8 bytes of English Wikipedia in 2006 in XML
Dataset Card for enwik8
Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
Dataset Description
- Homepage: http://mattmahoney.net/dc/textdata.html
- Repository: [Needs More Information]
- Paper: [Needs More Information]
- Leaderboard: https://paperswithcode.com/sota/language-modelling-on-enwiki8
- Point of Contact: [Needs More Information]
- Size of downloaded dataset files: 36.45 MB
- Size of the generated dataset: 102.38 MB
- Total amount of disk used: 138.83 MB
Dataset Summary
The enwik8 dataset is the first 100,000,000 (100M) bytes of the English Wikipedia XML dump on Mar. 3, 2006 and is typically used to measure a model's ability to compress data.
Supported Tasks and Leaderboards
A leaderboard for byte-level causal language modelling can be found on paperswithcode
Languages
en
Dataset Structure
Data Instances
- Size of downloaded dataset files: 36.45 MB
- Size of the generated dataset: 102.38 MB
- Total amount of disk used: 138.83 MB
{
"text": "In [[Denmark]], the [[Freetown Christiania]] was created in downtown [[Copenhagen]]....",
}Data Fields
The data fields are the same among all sets.
enwik8
text: astringfeature.
enwik8-raw
text: astringfeature.
Data Splits
Dataset Creation
Curation Rationale
[Needs More Information]
Source Data
Initial Data Collection and Normalization
The data is just English Wikipedia XML dump on Mar. 3, 2006 split by line for enwik8 and not split by line for enwik8-raw.
Who are the source language producers?
[Needs More Information]
Annotations
Annotation process
[Needs More Information]
Who are the annotators?
[Needs More Information]
Personal and Sensitive Information
[Needs More Information]
Considerations for Using the Data
Social Impact of Dataset
[Needs More Information]
Discussion of Biases
[Needs More Information]
Other Known Limitations
[Needs More Information]
Additional Information
Dataset Curators
[Needs More Information]
Licensing Information
[Needs More Information]
Citation Information
Dataset is not part of a publication, and can therefore not be cited.
Contributions
Thanks to @HallerPatrick for adding this dataset and @mtanghu for updating it.
