datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release:
HPLT3.0
We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0.
This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project.
The source of the data is mostly Internet Archive with some additions from Common Crawl.
For a detailed description of the dataset, please refer to our website and our pre-print.
The Cleaned variant of HPLT Datasets v2.0
This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.ultrafeedback-binarized-preferences-cleaned
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md.
Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize… See the full description on the dataset page: https://huggingface.co/datasets/yahma/alpaca-cleaned.ultrafeedback-binarized-preferences-cleaned-kto
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO
A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.smollm-corpus-cleaned
SmolLM-Corpus: Now shuffled and sharded (and Cleaned)!
This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming!
The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo.
Dataset Structure
The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.SCPWiki-Cleaned-PDF-Archivesgenius-lyrics-cleaned
◎
Genius Lyrics Dataset
Cleaned & Deduplicated
🤗 Hugging Face
🤗 Hugging Face
DOI: 10.57967/hf/7978
DOI: 10.57967/hf/7978
revision: 9742989
revision: 9742989
A heavily cleaned, English-only, genre-filtered subset of the Genius Song… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/genius-lyrics-cleaned.HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.turkishfineweb2-cleaned
TurkishFineweb2-Cleaned
A Turkish web corpus derived from the Turkish (tur_Latn) subset of
FineWeb-2, augmented with an additional quality-classification layer
and a near-duplicate removal pass.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Source
FineWeb-2 is a
large-scale, multilingual web corpus built from Common Crawl. This dataset
covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GLM-5.1-Reasoning-1M-Cleaned.Creative-Writing-Sonnet4.6-Cleaned
Creative-Writing-Sonnet4.6-Cleaned
Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved.
Format
Each line is a JSON object with:
messages: list of message dicts with roles (system, user, assistant)
System: writing quality instructions
User: cleaned creative writing prompt
Assistant: creative writing response (may include <think> traces)
Stats
Metric
Value
Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.UDM_cleaned_docs
UDM cleaned docs
6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6.
Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want.
How it was built
step
pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.Kimi-K2.5-Reasoning-1M-Cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.booksum-complete-cleaned
Description:
This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization
.
This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.TinyStoriesV2_cleaned
License: CDLA-Sharing-1.0
Dataset containing synthetically generated (GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
This is a cleaned up Version of the original TinyStories Dataset: https://huggingface.co/datasets/roneneldan/TinyStories.
We thank the authors for their contribution.
This Version only contains cleaned-up stories generated by GPT4.
Stories were deleted that contained spelling and… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/TinyStoriesV2_cleaned.CleanedFineWeb2Edu-jp
CleanedFineWeb2Edu-jp
CleanedFineWeb2Edu-jp is a cleaned Japanese web text dataset.
This dataset was created from the sample_10BT subset of
hotchpotch/fineweb-2-edu-japanese.
The source text was refined with
MK0727/corpus-refiner-jp.
Purpose
The main purpose of this dataset is to provide cleaner Japanese web text for
language model pretraining and continued pretraining.
This dataset keeps Japanese web documents from FineWeb2-Edu while reducing
boilerplate… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedFineWeb2Edu-jp.MoeGirlPedia_zh_cleaned_latest
🌐Language 中文|English
本数据集由2025年10月萌娘百科的快照经过清洗得来,专用于预训练等文本生成相关的模型训练。
特色
⚡体积优势
🧠文本易理解
💬更符合中文语境
仅经过基础清洗的数据集
1.06GB
存在复杂的网址链接残留的html标记正文内容被清除后残存的标题牛皮癣一样的引文注脚
暴力抹除非中文文字,导致信息缺失严重
本数据集
0.74GB(30.2%↓)
通过多重工序清洗基本不存在难以理解的文本内容保留部分英文以及少量其他语言文字(如日语)
仅经过基础清洗的数据集
size=66px|color=#8230FF|她已经不是我所认识的那个-{zh-hans:茜;zh-hant:仓式茜}-了。
'''仓式 茜'''(Kurashiki Akane)是由Spike Chunsoft所创作的系列游戏'''《极限脱出》'''及其衍生作品的主要角色之一。{{ZETOP}}
url=akanejunpei.jpg|position=up
图片说明=999中的茜(2027,21岁)
|本名=仓式 茜(くらしき… See the full description on the dataset page: https://huggingface.co/datasets/YCWTG/MoeGirlPedia_zh_cleaned_latest.mc4-pt-cleaned
Description
This is a clenned version of AllenAI mC4 PtBR section. The original dataset can be found here https://huggingface.co/datasets/allenai/c4
Clean procedure
We applied the same clenning procedure as explained here: https://gitlab.com/yhavinga/c4nlpreproc.git
The repository offers two strategies. The first one, found in the main.py file, uses pyspark to create a dataframe that can both clean the text and create a
pseudo mix on the entire dataset. We found this… See the full description on the dataset page: https://huggingface.co/datasets/thegoodfellas/mc4-pt-cleaned.Creative-Writing-KimiK2.5-Cleaned
Creative-Writing-KimiK2.5-Cleaned
Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved.
Format
Each line is a JSON object with:
messages: list of message dicts with roles (system, user, assistant)
System: writing quality instructions
User: cleaned creative writing prompt
Assistant: creative writing response (may include <think> traces)
Stats
Metric
Value
Total prompt tokens
80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.SuperWIKI-Cleaned
Dataset Card for SuperWIKI Cleaned
Dataset Summary
If you show most of those to people and ask them to form an opinion,the answer isn't just going to be "I don't know": it'll be "I don't care."
Tom Scott
SuperWIKI Cleaned is a focused dataset on wikipedia articles.
This dataset is derived from raw files provided in SuperWIKI.
Supported Tasks and Leaderboards
The dataset is generally used for Language Modeling.
Languages
English… See the full description on the dataset page: https://huggingface.co/datasets/RyokoExtra/SuperWIKI-Cleaned.c4-ko-cleaned-2이전 데이터셋에서 아쉬운 점이 많이 보여 조금 개선한 데이터셋 입니다.
원본 데이터셋: c4
파일 크기: 약 10gb
데이터 수: 2261464
openresearcher-sft-deep-research-cleaned
OpenResearcher SFT DeepResearch — Parquet Mirror
This is a re-hosted copy of the tool-reasoning SFT deep-research dataset by Aman Priyanshu, itself a cleaned/restructured version of the OpenResearcher Dataset from TIGER-AI-Lab.
Why this repo exists: the source wasn't laid out as ready-to-download Parquet files. This mirror simply stores the data as plain seed_*.parquet files so you can grab the whole dataset or a single segment easily. No changes were made to the content — all… See the full description on the dataset page: https://huggingface.co/datasets/DanielTobi0/openresearcher-sft-deep-research-cleaned.alpaca-cleaned-ru
alpaca-cleaned-ru
Translated version of yahma/alpaca-cleaned into Russian.
mc4_nl_cleaned
Dataset Card for Clean Dutch mC4
Dataset Summary
A cleaned version (151GB) of the Dutch part (277GB) of the C4 multilingual dataset (mC4).
Based on the Common Crawl dataset.
The original version was prepared by AllenAI, hosted at the address https://huggingface.co/datasets/allenai/c4.
Preprocessing
The Dutch portion of mC4 was cleaned in a similar fashion as the English cleaned C4 version.
See GitLab for details.
In summary, the preprocessing procedure… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/mc4_nl_cleaned.CleanedWiki-jp
CleanedWiki-jp
CleanedWiki-jp is a cleaned Japanese Wikipedia dataset prepared for LLM pre-training. It is built from Japanese Wikipedia article HTML, converted into Markdown, filtered for trainability.
The dataset keeps useful article structure instead of flattening everything into plain text. Suitable body tables are preserved as Markdown tables, and mathematical expressions are preserved in TeX form. Each row also includes a predicted Nippon Decimal Classification (NDC)… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedWiki-jp.Reviewer2_PGE_cleaned
Cleaned Review Dataset for Reviewer2
This is a cleaned version of our dataset and can be directly used for fine-tuning. The raw data files including metadata for each paper is in this directory.
venue: venue of the paper;
paper_content: content of the paper divided into sections
prompt: prompt generated for the review based on our PGE pipeline
format: the format of the review
review: human-written review for the paper
Dataset Sources
We incorporate parts of the… See the full description on the dataset page: https://huggingface.co/datasets/GitBag/Reviewer2_PGE_cleaned.ultra_feedback_dutch_cleaned
Ultra Feedback Dutch Cleaned
This is a cleaned version of BramVanroy/ultra_feedback_dutch, based on the cleaning done by Argilla on the original Ultra Feedback dataset. Another difference is that we only include GEITje 7B Ultra and GPT-4-Turbo. GEITje chat, which was used in the original dataset, is not used.
After cleaning I also generated replies for other models (like TowerInstruct, Mistral), but the results were too poor (in Dutch) to include so we only kept the GEITje Ultra and… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/ultra_feedback_dutch_cleaned.hplt-v3-pl-cleaned
HPLT v3 Polish — Cleaned & PII-Gated
Oczyszczony, polski podzbiór korpusu web HPLT v3, przygotowany jako materiał pretreningowy. Autor / kurator zbioru: Arkadiusz Słota (SlayerLab). Wartość dodana względem surowego HPLT: wieloetapowa bramka PII (usuwanie numerów telefonów, identyfikatorów, adresów) z niezależną weryfikacją na pełnych danych + lekkie czyszczenie boilerplate.
Wersja: v1.0 — floor (bins 8_5 + 8_6). Track B (bins 9_1 + 8_1, po deduplikacji względem bazy dynaword)… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/hplt-v3-pl-cleaned.Deltarune-Complete-Transcript-Cleaned
Deltarune Chapters 1–4 Dataset
Fan-made transcript dataset covering Deltarune Chapters 1 through 4. Processed from video playthroughs and cross-referenced with game data. Intended to provide LLMs with structured narrative context for a game whose content is underrepresented in training corpora.
Why This Exists
As of early 2026, major LLMs (including models with training cutoffs past July 2025) fail to recall basic plot details of Deltarune Chapters 3 and 4 despite their… See the full description on the dataset page: https://huggingface.co/datasets/Deltarunefan/Deltarune-Complete-Transcript-Cleaned.
