eduagarcia/cc100-pt
C100-PT CC100-PT is the is the portuguese subset from C100. C100 was created for training the multilingual Transformer XLM-R, containing two terabytes of cleaned data from 2018 snapshots of the Common Crawl project in 100 languages.
1512
