CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /fineweb-2 🥂 FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages. The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments. In particular, on the set of 9 diverse languages we used to guide our processing decisions, 🥂… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.tabulartext-generation1B<n<10B888 likes97k downloads11mo agoHugging Face02epfml /FineWeb2-HQ FineWeb2-HQ Dataset summary FineWeb2-HQ is a high-quality, model-filtered pretraining dataset derived as a subset of FineWeb2, spanning 20 languages. It enables around 6x faster pretraining compared to the base dataset. FineWeb2-HQ was created by selecting the top 10% quality documents of FineWeb2 in each language, based on scores assigned by a deep learning classifier trained to identify structured and knowledge-rich samples using XLM-RoBERTa embeddings. Validation… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-HQ.tabulartext-generation100M<n<1B81 likes29k downloads2y agoHugging Face03epfml /FineWeb2-embedded FineWeb2-embedded Dataset summary FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research. Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer texts… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.tabulartext-generation1B<n<10B6 likes18k downloads2y agoHugging Face04ssmits /fineweb-2-dutchtabulartext-generation10M<n<100M3 likes1.8k downloads2y agoHugging Face05Yahoo-Finance-News /FineWeb2024 FineWeb-Edu 2024 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2024. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2024 Rows 162,500,784… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2024.tabulartext-generation100M<n<1B0 likes1k downloads7d agoHugging Face06duarteocarmo /fineweb2-bagaco Bagaço 🍷🇵🇹 Bagaço is a pretraining dataset for European Portuguese. It filters the Fineweb2 dataset to URLs from Portuguese domains (e.g., .pt/). Each document is classified into one of 9 categories and scored for educational quality. Filtering Source: HuggingFaceFW/fineweb-2, subset por_Latn, split train Filter: URLs containing .pt/ (Portuguese top-level domain) Document classification Each document is classified into one of 9 categories: Society, Arts… See the full description on the dataset page: https://huggingface.co/datasets/duarteocarmo/fineweb2-bagaco.tabulartext-generation10M<n<100M2 likes956 downloads7mo agoHugging Face07Yahoo-Finance-News /FineWeb2025 FineWeb-Edu 2025 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2025 Rows 99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.tabulartext-generation10M<n<100M1 likes745 downloads7d agoHugging Face08Yahoo-Finance-News /FineWeb-2023 FineWeb-Edu 2023 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2023. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2023 Rows 104,280,950… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb-2023.tabulartext-generation100M<n<1B0 likes620 downloads7d agoHugging Face09DataHound26 /FineWeb-2020 FineWeb-Edu 2020 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2020. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2020 Rows 123,382,457… See the full description on the dataset page: https://huggingface.co/datasets/DataHound26/FineWeb-2020.tabulartext-generation100M<n<1B0 likes560 downloads7d agoHugging Face10DataHound26 /FineWeb-2021 FineWeb-Edu 2021 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2021. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2021 Rows 139,636,993… See the full description on the dataset page: https://huggingface.co/datasets/DataHound26/FineWeb-2021.tabulartext-generation100M<n<1B0 likes526 downloads7d agoHugging Face11DataHound26 /FineWeb-2022 FineWeb-Edu 2022 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2022. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2022 Rows 106,753,442… See the full description on the dataset page: https://huggingface.co/datasets/DataHound26/FineWeb-2022.tabulartext-generation100M<n<1B0 likes519 downloads7d agoHugging Face12Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes518 downloads1y agoHugging Face13tartuNLP /fineweb-2-et FineWeb2-et The Estonian subset of HuggingFaceFW/fineweb-2 reuploaded for ease of access. Licensing Information The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. The use of this dataset is also subject to CommonCrawl's Terms of Use. Citation Information @misc{penedo2025fineweb2pipelinescale, title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language}… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/fineweb-2-et.tabulartext-generation1M<n<10M0 likes449 downloads9mo agoHugging Face14bowang0911 /fineweb-2-autocurate fineweb-2-autocurate Autonomously curated subsets of HuggingFaceFW/fineweb-2. An LLM agent iteratively samples documents, identifies quality problems, and proposes heuristic fixes. Each fix is validated by training a small language model for 5 minutes and measuring BPB improvement on a Wikipedia eval set. Only fixes that improve BPB are kept. Built with autocurate. Subsets Language Subset Original Docs Kept Docs Kept % BPB Before → After Improvement Danish… See the full description on the dataset page: https://huggingface.co/datasets/bowang0911/fineweb-2-autocurate.tabulartext-generation100M<n<1B0 likes330 downloads6mo agoHugging Face15Ba2han /fineweb-2-turkish-categorized-long altaidevorg/fineweb-2-turkish-categorized long filtered Turkish texts Source: altaidevorg/fineweb-2-turkish-categorized (config: default). The script streamed 10,000,000 raw source rows before stopping. Categories ads, adult content, sports, tabloid were rejected before length and quality filtering. Retained rows contain 3,000–16,500 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/fineweb-2-turkish-categorized-long.tabulartext-generation100K<n<1M0 likes135 downloads2mo agoHugging Face16ReactiveAI /fineweb-2-pol-latest ReactiveAI - FineWeb2 PL subset This dataset is derived from polish subset of FineWeb2 by HuggingFace. Includes latest ~8.5M examples. Original dataset description below 🥂 FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages. The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/fineweb-2-pol-latest.tabulartext-generation1M<n<10M0 likes112 downloads10mo agoHugging Face17lianghsun /fineweb-2-zhtwgated Dataset Card for fineweb-2-zhtw fineweb-2-zhtw 是以 Hugging Face FineWeb-2 的 cmn_Hani(官話/漢字)子集為來源,經多層產地與字形/用語過濾後,取出以台灣繁體中文為主的網頁語料,保留 FineWeb-2 的完整 metadata,可作為繁中持續預訓練語料。 與單純「挑出繁體字」的做法不同:本資料集把「繁體字」和「台灣繁體內容」當成兩件事處理。中國網站的 BIG5 轉碼閘道、香港媒體、以及訂房網站的機器翻譯頁面都會產出繁體字,但用語與語感並非台灣中文,這些都在過濾流程中被排除。 Dataset Details 來源 cmn_Hani/train 共 370 個 parquet 分片、約 636,058,984 列(約 1.6 TB)。經過濾後保留 11,416,204 列,留存率 **1.81%**。 Dataset Sources Repository: lianghsun/fineweb-2-zhtw… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/fineweb-2-zhtw.tabulartext-generation10M<n<100M0 likes27 downloads2mo agoHugging Face18lianghsun /fineweb-2-edu-zhtwgated Dataset Card for fineweb-2-edu-zhtw UltraX 清洗欄位 原有欄位(text 與 zhtw_*、edu_* 等共 22 欄)維持不變,另新增以下欄位,記錄以 openbmb/UltraX-0.6B-Preview 清洗的結果。 UltraX 不做端到端改寫,而是預測結構化編輯操作(keep_all / remove_all / remove_lines / replace_str / add_line),再由程式確定性套用,因此每個決策都可被檢驗與否決。 欄位 說明 cleaned_text 清洗後文字,緊接 text 之後;needs_review 為 true 時等同原文 word_count / token_count 原文的詞數與 token 數(本資料集原本沒有,一併補上) cleaned_word_count / cleaned_token_count 清洗後對應數值,同一套算法 word_reduction_pct /… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/fineweb-2-edu-zhtw.tabulartext-generation100K<n<1M0 likes20 downloads1mo agoHugging Face19tampakwilll /fineweb2-id-filtered-10k FineWeb2-ID Filtered (min 10.000 karakter) Dataset ini adalah subset turunan dari HuggingFaceFW/fineweb-2, subset bahasa Indonesia (ind_Latn), yang disaring dengan syarat: Panjang teks (setelah strip whitespace) minimal 10.000 karakter per dokumen Diambil secara streaming dari split train, urutan asli (tanpa shuffle) Jumlah sampel: 2,097,152 dokumen, tersebar di 80 file JSONL. Sumber & Lisensi Seluruh isi teks berasal dari FineWeb-2 (lisensi ODC-By), yang pada… See the full description on the dataset page: https://huggingface.co/datasets/tampakwilll/fineweb2-id-filtered-10k.tabulartext-generation1M<n<10M1 likes20h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.