CoolFace
Datasetpublic

bertin-project/mc4-es-sampled

50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.

sourceHugging Faceodc-byupdated 4y agoView on Hugging Face
2likes992downloads
settings

This repository belongs to bertin-project on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemc4-es-sampled
visibilitypublic
licenceodc-by
gatedno
ownerbertin-project
Account settings
bertin-project/mc4-es-sampled · CoolFace