LeoCordoba/CC-NEWS-ES
Dataset Card for CC-NEWS-ES Dataset Summary CC-NEWS-ES is a Spanish-language dataset of news. The corpus was generated by extracting the Spanish articles from CC-NEWS (news index of Common Crawl) of 2019. For doing that FastText model was used for language prediction. It contains a total of 7,473,286 texts and 1,812,009,283 words distributed as follows: domain texts words ar 532703 1.45127e+08 bo 29557 7.28996e+06 br 107 14207 cl 116661… See the full description on the dataset page: https://huggingface.co/datasets/LeoCordoba/CC-NEWS-ES.
12657
