CoolFace
Datasetpublic

ThomasKendrick/openwebtext

Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/ThomasKendrick/openwebtext.

sourceHugging Facecc0-1.0updated 21d agoView on Hugging Face
0likes279downloads
1 commits on main
fb4a00721d ago

Duplicate from Skylion007/openwebtext

ThomasKendrick, parquet-converter