CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /fineweb-edu 📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper: https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu.tabulartext-generation1B<n<10B1.3k likes430k downloads1y agoHugging Face02Helsinki-NLP /fineweb-edu-translated Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.texttranslation1B<n<10B16 likes212k downloads5mo agoHugging Face03airtrain-ai /fineweb-edu-fortified Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it? Fineweb-Edu-Fortified is a dataset derived from Fineweb-Edu by applying exact-match deduplication across the whole dataset and producing an embedding for each row. The number of times the text from each row appears is also included as a count column. The embeddings were produced using TaylorAI/bge-micro Fineweb and… See the full description on the dataset page: https://huggingface.co/datasets/airtrain-ai/fineweb-edu-fortified.tabulartext-generation100M<n<1B65 likes60k downloads2y agoHugging Face04HuggingFaceFW /fineweb-edu-score-2 📚 FineWeb-Edu-score-2 1.3 trillion tokens of the finest educational data the 🌐 web has to offer What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens (FineWeb-Edu) and 5.4T tokens of educational web pages filtered from 🍷 FineWeb dataset. This is the 5.4 trillion version. Note: this version uses a lower educational score threshold = 2, which results in more documents, but lower quality compared to the 1.3T version. For more details check the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu-score-2.tabulartext-generation10B<n<100B89 likes20k downloads1y agoHugging Face05opencsg /chinese-fineweb-edu This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu dataset is a meticulously constructed high-quality Chinese pre-training corpus, specifically designed for natural language processing tasks in the education domain. This dataset undergoes a rigorous selection and… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu.texttext-generation10M<n<100M117 likes17k downloads10mo agoHugging Face06karpathy /fineweb-edu-100b-shuffletext10M<n<100M171 likes11k downloads1y agoHugging Face07PrimeIntellect /fineweb-edu Pre-shuffled fineweb-edu dataset text1B<n<10B2 likes6.9k downloads2y agoHugging Face08ReliableAI /irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B. tabular100K<n<1M1 likes6.7k downloads2y agoHugging Face09kushalt /fineweb-edu-gpt2tabular10M<n<100M0 likes6.5k downloads8mo agoHugging Face10lance-format /fineweb-edu FineWeb-Edu (Lance Format) A Lance-formatted version of FineWeb-Edu — over 1.5 billion educational web passages with cleaned text, source metadata, language detection signals, and 384-dim text embeddings — available directly from the Hub at hf://datasets/lance-format/fineweb-edu/data/train.lance. Key features Cleaned passage text in the text column with the source url and title carried alongside. Language detection signals (language, language_probability) for filtered… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fineweb-edu.tabulartext-retrieval1B<n<10B8 likes5.8k downloads4mo agoHugging Face11HuggingFaceFW /fineweb_edu_100BT-shuffled FineWeb-Edu 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.tabular100M<n<1B6 likes4.3k downloads7mo agoHugging Face12chilomax /fineweb-edu 📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper: https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/fineweb-edu.tabulartext-generation1B<n<10B0 likes4.3k downloads3mo agoHugging Face13KathirKs /fineweb-edu-hindi Fineweb-edu-hindi Fineweb-edu-hindi is a synthetic dataset generated by translating the Fineweb-edu to Hindi Language using IndicTrans2. The model variant used is IndicTrans2-en-indic-dist-200M. It contains about 300 Billion tokens in the Gemma-2-2b Tokenizer. Hardware Resources: The Google Cloud TPUs and the Google Cloud Platform was utilized for the dataset creation process. Code: Github: fineweb-translation Contact: If any queries or issues… See the full description on the dataset page: https://huggingface.co/datasets/KathirKs/fineweb-edu-hindi.text100M<n<1B8 likes4.3k downloads2y agoHugging Face14opencsg /chinese-fineweb-edu-v2 This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset V2 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu Dataset V2 is a comprehensive upgrade of the original Chinese Fineweb Edu, designed and optimized for natural language processing (NLP) tasks in the education sector. This high-quality Chinese pretraining dataset has… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu-v2.tabulartext-generation100M<n<1B75 likes3.4k downloads10mo agoHugging Face15minpeter /fineweb-2-edu-korean-rawHuggingFaceFW/fineweb-2 (v2.1.0) It took about 9 hours on A100 80gbx4 to process the dataset. tabular10M<n<100M1 likes3.3k downloads1y agoHugging Face16hotchpotch /fineweb-2-edu-japanese 🍷 FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset This dataset consists of 120 million texts (approximately 89.3B tokens) filtered from the 376 million Japanese texts in FineWeb2 that were deemed educational. The following subsets are also provided: default: Approximately 120M texts (120 million texts) totaling around 89.3B tokens sample_10BT: A random sample of about 10B tokens from the default dataset small_tokens: Data composed solely of texts with 512 tokens… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-2-edu-japanese.tabular100M<n<1B34 likes3.3k downloads1y agoHugging Face17HuggingFaceFW /finepdfs_50BT-dclm_30BT-fineweb_edu_20BT FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing. Dataset Description Component Source Tokens FinePDFs 100BT FinePDFs ~50B DCLM 100BT DCLM-Baseline 1.0 ~30B FineWeb-Edu 100BT FineWeb-Edu ~20B The schema is reduced to the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.text10M<n<100M3 likes2.9k downloads7mo agoHugging Face18ByteDance-Seed /mga-fineweb-edu Massive Genre-Audience Augment Fineweb-Edu Corpus This dataset is a synthetic pretraining corpus described in paper Reformulation for Pretraining Data Augmentation. Overview of synthesis framework. Our method expands the original corpus through a two-stage synthesis process. Each document is reformulated to 5 new documents, achieving 3.9× token number expansion while maintaining diversity through massive (genre, audience) pairs. We build MGACorpus based on SmolLM Corpus… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/mga-fineweb-edu.texttext-generation100M<n<1B44 likes2.6k downloads1y agoHugging Face19DanielGallagherIRE /FineWeb-Edu-10B-PMI-Filteredtext1M<n<10M0 likes2.6k downloads3mo agoHugging Face20HuggingFaceFW /finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio, using the educational subset of FinePDFs. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing. Dataset Description Component Source Tokens FinePDFs-Edu 100BT FinePDFs-Edu ~50B DCLM 100BT DCLM-Baseline 1.0 ~30B FineWeb-Edu 100BT… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.text10M<n<100M0 likes2.5k downloads7mo agoHugging Face21DanielGallagherIRE /FineWeb-Edu-10B-Shuffledtext1M<n<10M0 likes2k downloads3mo agoHugging Face22willRD /Fineweb-Edu-Chinese-V2.1 Chinese Fineweb Edu Dataset V2.1 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different folders… See the full description on the dataset page: https://huggingface.co/datasets/willRD/Fineweb-Edu-Chinese-V2.1.texttext-generation100M<n<1B0 likes1.9k downloads10mo agoHugging Face23jiviteshjn /fineweb-edu-zh-chengyu-cpt Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.tabulartext-generation1M<n<10M1 likes1.8k downloads2mo agoHugging Face24swiss-ai /fineweb-edu-compliant-tagtext1B<n<10B1 likes1.8k downloads1y agoHugging Face25ByteSpanTokenisers /finewebedu-20B FineWebEDU 20B A copy of FineWebEDU-20B used for out tokenizer experiments. The subsets are as follows: bytelevel: the full dataset tokenized using our bytelevel tokenizer bytelevel-subset_1: a 100k-row subset of the bytelevel subset, used to train bytelevel models. bytelevel-subset_2: a 100k-row subset of the bytelevel subset, used to extract llm predictions. bytelevel-llm-data: a copy of bytelevel-subset_2 with lm predictions, used to train bytespan tokenizers… See the full description on the dataset page: https://huggingface.co/datasets/ByteSpanTokenisers/finewebedu-20B.text100M<n<1B0 likes1.7k downloads1y agoHugging Face26bhavnicksm /fineweb-edu-micro FineWeb-Edu Micro This dataset is a subset of the FineWeb-Edu Sample-10BT, which contains passages that are at least 1000 tokens long, totalling about 1 Million tokens . This dataset was primarily made to evaluate different RAG Chunking mechanisms in Chonkie tabularn<1K0 likes1.6k downloads2y agoHugging Face27wissamantoun /fineweb-edu-format-topic FineWeb-Edu w/ Topic and Format Annotations FineWeb-Edu dataset consists of 1.3T tokens annotated for Topic and Format using wissamantoun/WebOrganizer-TopicClassifier-ModernBERT and wissamantoun/WebOrganizer-FormatClassifier-ModernBERT classifiers. Similar to WebOrganizer/Corpus-200B but using FineEdu instead of DCLM. Topic Labels: Adult Art & Design Software Dev. Crime & Law Education & Jobs Hardware Entertainment Social Life Fashion & Beauty Finance & Business Food & Dining… See the full description on the dataset page: https://huggingface.co/datasets/wissamantoun/fineweb-edu-format-topic.texttext-generation1B<n<10B5 likes1.6k downloads1y agoHugging Face28DanielGallagherIRE /FineWeb-Edu-10B-Nouns-Onlytext1M<n<10M0 likes1.6k downloads3mo agoHugging Face29mmarone /fineweb-edu-full-metadata[WIP] FineWeb-Edu with Metadata This repo contains 3 versions of the FineWeb-Edu v1 dataset: fwedu1-metaonly/ fwedu1-text-content-zstd/ fineweb-edu-1.0.0-meta-and-text/ These are all joinable via the hash column, which is xxhash64 in pyspark, calculated on the text column. This hash is unique for all instances in the dataset. For convenience, this join is done for you in the third table fwedu1-metaonly is just the metadata of the data exactly as it comes from the FineWeb-Edu v1… See the full description on the dataset page: https://huggingface.co/datasets/mmarone/fineweb-edu-full-metadata.tabular100M<n<1B0 likes1.5k downloads1y agoHugging Face30SultanR /fineweb-edu-arabic fineweb-edu-arabic Arabic translation of FineWeb-Edu (sample/350BT subset, filtered to language_score > 0.9), translated with Seed-X-PPO-7B using greedy decoding. Documents were split into ~490-token chunks, translated, and reassembled. Each row is one complete document. A companion corpus translated with the same pipeline is available at dclm-pro-arabic. Details Documents: 82,840,410 (27.9% of the source subset, uniformly sampled) Arabic tokens: ~170B (Seed-X… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/fineweb-edu-arabic.texttext-generation10M<n<100M1 likes1.5k downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.