CoolFace
3 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face02LVSTCK /macedonian-corpus-cleaned-dedup Macedonian Corpus - Cleaned and Deduplicated Paper 🌟 Key Highlights Size: 16.78 GB, Word Count: 1.47 billion Deduplicated using MinHash to remove redundant documents. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.texttext-generation1M<n<10M1 likes56 downloads1y agoHugging Face03LVSTCK /macedonian-corpus-cleaned Macedonian Corpus - Cleaned raw version here Paper 🌟 Key Highlights Size: 35.5 GB, Word Count: 3.31 billion Filtered for irrelevant and low-quality content using C4 and Gopher filtering. Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.texttext-generation1M<n<10M0 likes45 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.