CoolFace
Datasetpublic

eduagarcia/cc100-pt

C100-PT CC100-PT is the is the portuguese subset from C100. C100 was created for training the multilingual Transformer XLM-R, containing two terabytes of cleaned data from 2018 snapshots of the Common Crawl project in 100 languages.

sourceHugging Faceupdated 3y agoView on Hugging Face
1likes512downloads
Dataset Card

C100-PT

CC100-PT is the is the portuguese subset from C100. C100 was created for training the multilingual Transformer XLM-R, containing two terabytes of cleaned data from 2018 snapshots of the Common Crawl project in 100 languages.