CoolFace
Datasetpublic

sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen

Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled" num_examples: 33.5 million download_size: 15.3 GB dataset_size: 26.1 GB This dataset combines wikipedia20220301.en and bookcorpusopen, and splits the data into smaller chunks, of size ~820 chars (such that each item will be at least ~128 tokens for the average tokenizer). The order of the items in this dataset has been shuffled, meaning you don't have to use dataset.shuffle, which is slower to iterate… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.

sourceHugging Faceupdated 3y agoView on Hugging Face
4likes412downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen · CoolFace