CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HPLT /HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. The Cleaned variant of HPLT Datasets v2.0 This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.tabularfill-mask1B<n<10B45 likes175k downloads4mo agoHugging Face02argilla /ultrafeedback-binarized-preferences-cleaned UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md. Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.tabulartext-generation10K<n<100K165 likes27k downloads3y agoHugging Face03yahma /alpaca-cleaned Dataset Card for Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer. "instruction":"Summarize… See the full description on the dataset page: https://huggingface.co/datasets/yahma/alpaca-cleaned.texttext-generation10K<n<100K888 likes25k downloads3y agoHugging Face04argilla /ultrafeedback-binarized-preferences-cleaned-kto UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.texttext-generation100K<n<1M10 likes16k downloads3y agoHugging Face05Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face06AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.8k downloads1y agoHugging Face07theelderemo /genius-lyrics-cleaned ◎ Genius Lyrics Dataset Cleaned & Deduplicated 🤗 Hugging Face 🤗 Hugging Face DOI: 10.57967/hf/7978 DOI: 10.57967/hf/7978 revision: 9742989 revision: 9742989 A heavily cleaned, English-only, genre-filtered subset of the Genius Song… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/genius-lyrics-cleaned.texttext-generation1M<n<10M19 likes7k downloads7mo agoHugging Face08sapientinc /HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.texttext-generation100M<n<1B17 likes3.7k downloads4mo agoHugging Face09unsloth /alpaca-cleaned Dataset Card for Alpaca-Cleaned Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.texttext-generation10K<n<100K25 likes3.4k downloads9mo agoHugging Face10moganai /turkishfineweb2-cleaned TurkishFineweb2-Cleaned A Turkish web corpus derived from the Turkish (tur_Latn) subset of FineWeb-2, augmented with an additional quality-classification layer and a near-duplicate removal pass. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Source FineWeb-2 is a large-scale, multilingual web corpus built from Common Crawl. This dataset covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.tabulartext-generation10M<n<100M5 likes1.6k downloads27d agoHugging Face11Jackrong /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M297 likes1k downloads5mo agoHugging Face12Crownelius /Creative-Writing-Sonnet4.6-Cleaned Creative-Writing-Sonnet4.6-Cleaned Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.texttext-generationn<1K3 likes853 downloads2mo agoHugging Face13oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes620 downloads24d agoHugging Face14Jackrong /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M36 likes604 downloads5mo agoHugging Face15ubaada /booksum-complete-cleaned Description: This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization . This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.textsummarization1K<n<10K23 likes452 downloads2y agoHugging Face16fhswf /TinyStoriesV2_cleaned License: CDLA-Sharing-1.0 Dataset containing synthetically generated (GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. This is a cleaned up Version of the original TinyStories Dataset: https://huggingface.co/datasets/roneneldan/TinyStories. We thank the authors for their contribution. This Version only contains cleaned-up stories generated by GPT4. Stories were deleted that contained spelling and… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/TinyStoriesV2_cleaned.texttext-generation1M<n<10M13 likes391 downloads2y agoHugging Face17KeisukeMiyamoto /CleanedFineWeb2Edu-jp CleanedFineWeb2Edu-jp CleanedFineWeb2Edu-jp is a cleaned Japanese web text dataset. This dataset was created from the sample_10BT subset of hotchpotch/fineweb-2-edu-japanese. The source text was refined with MK0727/corpus-refiner-jp. Purpose The main purpose of this dataset is to provide cleaner Japanese web text for language model pretraining and continued pretraining. This dataset keeps Japanese web documents from FineWeb2-Edu while reducing boilerplate… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedFineWeb2Edu-jp.tabulartext-generation10M<n<100M1 likes363 downloads1mo agoHugging Face18YCWTG /MoeGirlPedia_zh_cleaned_latest 🌐Language 中文|English 本数据集由2025年10月萌娘百科的快照经过清洗得来,专用于预训练等文本生成相关的模型训练。 特色 ⚡体积优势 🧠文本易理解 💬更符合中文语境 仅经过基础清洗的数据集 1.06GB 存在复杂的网址链接残留的html标记正文内容被清除后残存的标题牛皮癣一样的引文注脚 暴力抹除非中文文字,导致信息缺失严重 本数据集 0.74GB(30.2%↓) 通过多重工序清洗基本不存在难以理解的文本内容保留部分英文以及少量其他语言文字(如日语) 仅经过基础清洗的数据集 size=66px|color=#8230FF|她已经不是我所认识的那个-{zh-hans:茜;zh-hant:仓式茜}-了。 '''仓式 茜'''(Kurashiki Akane)是由Spike Chunsoft所创作的系列游戏'''《极限脱出》'''及其衍生作品的主要角色之一。{{ZETOP}} url=akanejunpei.jpg|position=up 图片说明=999中的茜(2027,21岁) |本名=仓式 茜(くらしき… See the full description on the dataset page: https://huggingface.co/datasets/YCWTG/MoeGirlPedia_zh_cleaned_latest.texttext-generation100K<n<1M4 likes344 downloads11mo agoHugging Face19thegoodfellas /mc4-pt-cleaned Description This is a clenned version of AllenAI mC4 PtBR section. The original dataset can be found here https://huggingface.co/datasets/allenai/c4 Clean procedure We applied the same clenning procedure as explained here: https://gitlab.com/yhavinga/c4nlpreproc.git The repository offers two strategies. The first one, found in the main.py file, uses pyspark to create a dataframe that can both clean the text and create a pseudo mix on the entire dataset. We found this… See the full description on the dataset page: https://huggingface.co/datasets/thegoodfellas/mc4-pt-cleaned.textfill-mask100M<n<1B4 likes340 downloads3y agoHugging Face20Crownelius /Creative-Writing-KimiK2.5-Cleaned Creative-Writing-KimiK2.5-Cleaned Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens 80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.texttext-generationn<1K8 likes338 downloads2mo agoHugging Face21RyokoExtra /SuperWIKI-Cleaned Dataset Card for SuperWIKI Cleaned Dataset Summary If you show most of those to people and ask them to form an opinion,the answer isn't just going to be "I don't know": it'll be "I don't care." Tom Scott SuperWIKI Cleaned is a focused dataset on wikipedia articles. This dataset is derived from raw files provided in SuperWIKI. Supported Tasks and Leaderboards The dataset is generally used for Language Modeling. Languages English… See the full description on the dataset page: https://huggingface.co/datasets/RyokoExtra/SuperWIKI-Cleaned.texttext-generation100K<n<1M4 likes336 downloads3y agoHugging Face22blueapple8259 /c4-ko-cleaned-2이전 데이터셋에서 아쉬운 점이 많이 보여 조금 개선한 데이터셋 입니다. 원본 데이터셋: c4 파일 크기: 약 10gb 데이터 수: 2261464 texttext-generation1M<n<10M3 likes311 downloads2y agoHugging Face23DanielTobi0 /openresearcher-sft-deep-research-cleaned OpenResearcher SFT DeepResearch — Parquet Mirror This is a re-hosted copy of the tool-reasoning SFT deep-research dataset by Aman Priyanshu, itself a cleaned/restructured version of the OpenResearcher Dataset from TIGER-AI-Lab. Why this repo exists: the source wasn't laid out as ready-to-download Parquet files. This mirror simply stores the data as plain seed_*.parquet files so you can grab the whole dataset or a single segment easily. No changes were made to the content — all… See the full description on the dataset page: https://huggingface.co/datasets/DanielTobi0/openresearcher-sft-deep-research-cleaned.tabulartext-generation10K<n<100K0 likes286 downloads2mo agoHugging Face24d0rj /alpaca-cleaned-ru alpaca-cleaned-ru Translated version of yahma/alpaca-cleaned into Russian. texttext-generation10K<n<100K22 likes285 downloads3y agoHugging Face25yhavinga /mc4_nl_cleanedgated Dataset Card for Clean Dutch mC4 Dataset Summary A cleaned version (151GB) of the Dutch part (277GB) of the C4 multilingual dataset (mC4). Based on the Common Crawl dataset. The original version was prepared by AllenAI, hosted at the address https://huggingface.co/datasets/allenai/c4. Preprocessing The Dutch portion of mC4 was cleaned in a similar fashion as the English cleaned C4 version. See GitLab for details. In summary, the preprocessing procedure… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/mc4_nl_cleaned.texttext-generation100M<n<1B16 likes257 downloads1y agoHugging Face26KeisukeMiyamoto /CleanedWiki-jp CleanedWiki-jp CleanedWiki-jp is a cleaned Japanese Wikipedia dataset prepared for LLM pre-training. It is built from Japanese Wikipedia article HTML, converted into Markdown, filtered for trainability. The dataset keeps useful article structure instead of flattening everything into plain text. Suitable body tables are preserved as Markdown tables, and mathematical expressions are preserved in TeX form. Each row also includes a predicted Nippon Decimal Classification (NDC)… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedWiki-jp.tabulartext-generation1M<n<10M0 likes240 downloads1mo agoHugging Face27GitBag /Reviewer2_PGE_cleaned Cleaned Review Dataset for Reviewer2 This is a cleaned version of our dataset and can be directly used for fine-tuning. The raw data files including metadata for each paper is in this directory. venue: venue of the paper; paper_content: content of the paper divided into sections prompt: prompt generated for the review based on our PGE pipeline format: the format of the review review: human-written review for the paper Dataset Sources We incorporate parts of the… See the full description on the dataset page: https://huggingface.co/datasets/GitBag/Reviewer2_PGE_cleaned.texttext-generation10K<n<100K2 likes230 downloads3y agoHugging Face28BramVanroy /ultra_feedback_dutch_cleaned Ultra Feedback Dutch Cleaned This is a cleaned version of BramVanroy/ultra_feedback_dutch, based on the cleaning done by Argilla on the original Ultra Feedback dataset. Another difference is that we only include GEITje 7B Ultra and GPT-4-Turbo. GEITje chat, which was used in the original dataset, is not used. After cleaning I also generated replies for other models (like TowerInstruct, Mistral), but the results were too poor (in Dutch) to include so we only kept the GEITje Ultra and… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/ultra_feedback_dutch_cleaned.tabulartext-generation100K<n<1M6 likes230 downloads2y agoHugging Face29SlayerLab /hplt-v3-pl-cleaned HPLT v3 Polish — Cleaned & PII-Gated Oczyszczony, polski podzbiór korpusu web HPLT v3, przygotowany jako materiał pretreningowy. Autor / kurator zbioru: Arkadiusz Słota (SlayerLab). Wartość dodana względem surowego HPLT: wieloetapowa bramka PII (usuwanie numerów telefonów, identyfikatorów, adresów) z niezależną weryfikacją na pełnych danych + lekkie czyszczenie boilerplate. Wersja: v1.0 — floor (bins 8_5 + 8_6). Track B (bins 9_1 + 8_1, po deduplikacji względem bazy dynaword)… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/hplt-v3-pl-cleaned.texttext-generation10M<n<100M0 likes221 downloads1mo agoHugging Face30Deltarunefan /Deltarune-Complete-Transcript-Cleaned Deltarune Chapters 1–4 Dataset Fan-made transcript dataset covering Deltarune Chapters 1 through 4. Processed from video playthroughs and cross-referenced with game data. Intended to provide LLMs with structured narrative context for a game whose content is underrepresented in training corpora. Why This Exists As of early 2026, major LLMs (including models with training cutoffs past July 2025) fail to recall basic plot details of Deltarune Chapters 3 and 4 despite their… See the full description on the dataset page: https://huggingface.co/datasets/Deltarunefan/Deltarune-Complete-Transcript-Cleaned.texttext-generation10K<n<100K4 likes220 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.