CoolFace
20 results

c4

allenai /c4 C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.texttext-generation10B<n<100B671 likes1.3m downloads3y agoHugging Facelegacy-datasets /c4A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset by AllenAI.text-generation100M<n<1B242 likes19k downloads3y agoHugging Facechanind /c4-10k-mini-tokenized-16-ctx-gelu-1l-tests1K<n<10K0 likes6.9k downloads2y agoHugging Facecarolina-c4ai /corpus-carolinaCarolina is an Open Corpus for Linguistics and Artificial Intelligence with a robust volume of texts of varied typology in contemporary Brazilian Portuguese (1970-).fill-mask1B<n<10B32 likes5.3k downloads1y agoHugging FaceNeelNanda /c4-10k Dataset Card for "c4-10k" More Information needed text10K<n<100K0 likes5.1k downloads4y agoHugging FaceNeelNanda /c4-tokenized-2b Dataset Card for "c4-tokenized-2b" More Information needed 1M<n<10M0 likes5k downloads4y agoHugging Face