CoolFace
Datasetpublic

LeoCordoba/CC-NEWS-ES

Dataset Card for CC-NEWS-ES Dataset Summary CC-NEWS-ES is a Spanish-language dataset of news. The corpus was generated by extracting the Spanish articles from CC-NEWS (news index of Common Crawl) of 2019. For doing that FastText model was used for language prediction. It contains a total of 7,473,286 texts and 1,812,009,283 words distributed as follows: domain texts words ar 532703 1.45127e+08 bo 29557 7.28996e+06 br 107 14207 cl 116661… See the full description on the dataset page: https://huggingface.co/datasets/LeoCordoba/CC-NEWS-ES.

sourceHugging Facemitupdated 4y agoView on Hugging Face
12likes657downloads
filear.zip314.2 MBdownload
filebo.zip15.8 MBdownload
filebr.zip36 KBdownload
filecl.zip64.2 MBdownload
fileco.zip37.8 MBdownload
filecom.zip1.53 GBdownload
filecr.zip7.2 MBdownload
filees.zip922.0 MBdownload
filegt.zip1.8 MBdownload
filehn.zip11.5 MBdownload
filemx.zip299.9 MBdownload
fileni.zip24.6 MBdownload
filepa.zip9.1 MBdownload
filepe.zip77.6 MBdownload
filepr.zip3.8 MBdownload
filepy.zip32.2 MBdownload
filesv.zip307 KBdownload
fileuy.zip61.3 MBdownload
fileve.zip14.8 MBdownload

LeoCordoba/CC-NEWS-ES · main · files are served by the source, never re-hosted here