bertin-project/mc4-es-sampled
50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.
Update README.md
Update README.md
Fix task tags (#2)
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update mc4-es-sampled.py
Update mc4-es-sampled.py
Update mc4-es-sampled.py
Update dataset info
Update README.md
Update README.md
Update README.md
Fix indices
Update README.md
Update README.md
Update README.md
Update README.md
Create README.md
Updating URL
Testing new dataset
Adding 50M samples from Spanish mC4 sampled using a gaussian perplexity function
Adding 50M samples from Spanish mC4 sampled using a stepwise perplexity function
Adding 50M random samples from Spanish mC4
Adding .json.gz and .jsonl.gz to LFS
initial commit
