CoolFace
Datasetpublic

bertin-project/mc4-es-sampled

50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.

sourceHugging Faceodc-byupdated 4y agoView on Hugging Face
2likes717downloads
29 commits on main
b03a8384y ago

Update README.md

versae
007295a4y ago

Update README.md

versae
b7cd2264y ago

Fix task tags (#2)

versae, albertvillanova
78e98e34y ago

Update README.md

versae
89952044y ago

Update README.md

versae
5dc758b4y ago

Update README.md

versae
492ea9f4y ago

Update README.md

versae
76427294y ago

Update README.md

versae
f5f8cbd4y ago

Update README.md

versae
7a98abb5y ago

Update mc4-es-sampled.py

versae
d0737435y ago

Update mc4-es-sampled.py

versae
75122325y ago

Update mc4-es-sampled.py

versae
4930a285y ago

Update dataset info

versae
636f5335y ago

Update README.md

versae
9b1ad675y ago

Update README.md

versae
0e0595b5y ago

Update README.md

versae
82dca2e5y ago

Fix indices

versae
82210545y ago

Update README.md

Pablogps
fb4e7a05y ago

Update README.md

versae
c42aca65y ago

Update README.md

versae
18d73e85y ago

Update README.md

versae
2e1ba115y ago

Create README.md

versae
bfeaed55y ago

Updating URL

versae
b91c10e5y ago

Testing new dataset

versae
890390d5y ago

Adding 50M samples from Spanish mC4 sampled using a gaussian perplexity function

versae
683bab65y ago

Adding 50M samples from Spanish mC4 sampled using a stepwise perplexity function

versae
0ad177c5y ago

Adding 50M random samples from Spanish mC4

versae
14939a35y ago

Adding .json.gz and .jsonl.gz to LFS

versae
1a5d3965y ago

initial commit

system