Skylion007/openwebtext
Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/Skylion007/openwebtext.
Update size metadata in readme (#28)
Convert dataset to Parquet (part 00001-of-00002) (#23)
Convert dataset to Parquet (part 00000-of-00002) (#22)
Fix bibtex
Make citation match website
Convert dataset sizes from base 2 to base 10 in the dataset card (#5)
Make the dataset streamable (#3)
add dataset_info in dataset metadata
remove dummmy data
Align more metadata with other repo types (models,spaces) (#4607)
Remove config names as yaml keys (#4367)
Remove a copy-paste sentence in dataset cards (#4281)
Update datasets task tags to align tags with models (#4067)
Fix typo in train split name (#3751)
Update files from the datasets library (from 1.18.0)
Update files from the datasets library (from 1.12.0)
Update files from the datasets library (from 1.8.0)
Update files from the datasets library (from 1.7.0)
Update files from the datasets library (from 1.6.1)
Update files from the datasets library (from 1.6.0)
Update files from the datasets library (from 1.4.0)
Update files from the datasets library (from 1.3.0)
Update files from the datasets library (from 1.1.0)
