CoolFace
Datasetpublic

segyges/OpenWebText2

Dataset Card for OpenWebText2 OpenWebText2 is a reasonably large corpus of scraped natural language data. Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate… See the full description on the dataset page: https://huggingface.co/datasets/segyges/OpenWebText2.

sourceHugging Facemitupdated 2y agoView on Hugging Face
18likes135downloads
Dataset Card

Dataset Card for OpenWebText2

OpenWebText2 is a reasonably large corpus of scraped natural language data.

Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate those uses.

I am not acting on behalf of the original authors of the dataset.

More: https://openwebtext2.readthedocs.io/en/latest/

Dataset Description

  • Language(s) (NLP): English
  • License: MIT

Dataset Sources [optional]

  • Repository: https://github.com/EleutherAI/openwebtext2
  • Paper: https://arxiv.org/abs/2101.00027

Dataset Card Authors

SE Gyges

Dataset Card Contact

segyges on github or gmail.