datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release:
HPLT3.0
We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0.
This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project.
The source of the data is mostly Internet Archive with some additions from Common Crawl.
For a detailed description of the dataset, please refer to our website and our pre-print.
The Cleaned variant of HPLT Datasets v2.0
This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.ultrafeedback-binarized-preferences-cleaned
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md.
Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.cad-corpus-cleanedGenius-song-lyrics-cleaned
🎵 Genius Song Lyrics cleaned Dataset
Dataset Description
This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis.
The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content.
Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.codeparrot-train-more-filter-3.3b-cleanedultrafeedback_binarized_cleaned
Dataset Card for "ultrafeedback_binarized_cleaned"
Update 1/12/2023: I've removed examples identified as faulty by Argilla - see their awesome work for more details.
This is a version of the UltraFeedback binarized dataset but with TruthfulQA prompts removed and source annotations added (so you can filter out samples from different sources yourself if you want!).
Please see the binarized dataset card for more information, or the original UltraFeedback dataset card.
turkishfineweb2-cleaned
TurkishFineweb2-Cleaned
A Turkish web corpus derived from the Turkish (tur_Latn) subset of
FineWeb-2, augmented with an additional quality-classification layer
and a near-duplicate removal pass.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Source
FineWeb-2 is a
large-scale, multilingual web corpus built from Common Crawl. This dataset
covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.funes-xiaowu0162-longmemeval-cleaned-s
Funes recall store — LongMemEval_s cleaned corpus
A funes recall store built by indexing the
longmemeval_s_cleaned.json haystack of
xiaowu0162/longmemeval-cleaned
(LongMemEval, arXiv:2410.10813) — every unique
chat session across all 500 questions' haystacks, in one corpus-wide store.
What this is
This is not a raw trace dataset — it is a pre-built funes index: the source
sessions chunked into content blocks and embedded, stored as a
Lance table (chunks.lance).… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-xiaowu0162-longmemeval-cleaned-s.mc4_3.1.0_fi_cleaned
Dataset Card for "mc4_3.1.0_fi_cleaned"
More Information needed
CrawlPT_dedup_Cleaned📚 CrawlPT Clean — High-Quality Portuguese Corpus
Versão limpa, filtrada e refinada do dataset CrawlPT_dedup
🧼 Visão Geral
Este repositório fornece uma versão limpa, filtrada e padronizada do dataset:
➡️ eduagarcia/CrawlPT_dedup
https://huggingface.co/datasets/eduagarcia/CrawlPT_dedup
A limpeza tem como objetivo criar um corpus de alta qualidade para:
pré-treino contínuo de modelos LLM (Qwen, Mistral, LLaMA, Phi etc.)
melhora de fluência e coerência em português
pesquisas em NLP
geração de… See the full description on the dataset page: https://huggingface.co/datasets/tiagoloeblein/CrawlPT_dedup_Cleaned.oscar_2301_fi_cleaned
Dataset Card for "oscar_2301_fi_cleaned"
More Information needed
python-edu-cleaned
SmolLM-Corpus: Python-Edu (Cleaned)
This dataset contains the python-edu subset of SmolLM-Corpus with the contents of the files stored in a new text field. All files were downloaded from the S3 bucket on January the 8th 2025, using the blob IDs from the original dataset with revision 3ba9d605774198c5868892d7a8deda78031a781f. Only 1 file was marked as not found and the corresponding row removed from the dataset (content/39c3e5b85cc678d1d54b4d93a55271c51d54126c which I suspect is… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/python-edu-cleaned.UDM_cleaned_docs
UDM cleaned docs
6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6.
Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want.
How it was built
step
pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.jetson1-060926-subtask-place-full-cleanedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060926-subtask-place-full-cleaned.jetson1-060826-subtask-grab2-full-cleanedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060826-subtask-grab2-full-cleaned.uci-drug-review-cleanedEarnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.CleanedFineWeb2Edu-jp
CleanedFineWeb2Edu-jp
CleanedFineWeb2Edu-jp is a cleaned Japanese web text dataset.
This dataset was created from the sample_10BT subset of
hotchpotch/fineweb-2-edu-japanese.
The source text was refined with
MK0727/corpus-refiner-jp.
Purpose
The main purpose of this dataset is to provide cleaner Japanese web text for
language model pretraining and continued pretraining.
This dataset keeps Japanese web documents from FineWeb2-Edu while reducing
boilerplate… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedFineWeb2Edu-jp.qmsum-cleaned
qmsum-cleaned
prefixes
It's worth noting that each "document" in input is prefixed by a question/prompt on what the model is supposed to do. You may want to explicitly handle this in some way, or prefix your models trained on this dataset.
Most frequent "prefixes" separated via sentence-splitter in the train split:
Sentence
Count
0
Summarize the whole meeting.
121
1
Summarize the meeting
25
2
What did the team discuss about the product cost?
4
3
How did… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/qmsum-cleaned.openresearcher-sft-deep-research-cleaned
OpenResearcher SFT DeepResearch — Parquet Mirror
This is a re-hosted copy of the tool-reasoning SFT deep-research dataset by Aman Priyanshu, itself a cleaned/restructured version of the OpenResearcher Dataset from TIGER-AI-Lab.
Why this repo exists: the source wasn't laid out as ready-to-download Parquet files. This mirror simply stores the data as plain seed_*.parquet files so you can grab the whole dataset or a single segment easily. No changes were made to the content — all… See the full description on the dataset page: https://huggingface.co/datasets/DanielTobi0/openresearcher-sft-deep-research-cleaned.sft-one_shot-cleanedultra_feedback_dutch_cleaned
Ultra Feedback Dutch Cleaned
This is a cleaned version of BramVanroy/ultra_feedback_dutch, based on the cleaning done by Argilla on the original Ultra Feedback dataset. Another difference is that we only include GEITje 7B Ultra and GPT-4-Turbo. GEITje chat, which was used in the original dataset, is not used.
After cleaning I also generated replies for other models (like TowerInstruct, Mistral), but the results were too poor (in Dutch) to include so we only kept the GEITje Ultra and… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/ultra_feedback_dutch_cleaned.CleanedWiki-jp
CleanedWiki-jp
CleanedWiki-jp is a cleaned Japanese Wikipedia dataset prepared for LLM pre-training. It is built from Japanese Wikipedia article HTML, converted into Markdown, filtered for trainability.
The dataset keeps useful article structure instead of flattening everything into plain text. Suitable body tables are preserved as Markdown tables, and mathematical expressions are preserved in TeX form. Each row also includes a predicted Nippon Decimal Classification (NDC)… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedWiki-jp.prompt_injection_cleaned_dataset
Dataset Card for "prompt_injection_cleaned_dataset"
More Information needed
EMBER_cleaned
EMBER Cleaned
EMBER Cleaned is a cleaned and AI-ready version of the original EMBER (Endgame Malware Benchmark for Research) dataset, a widely used benchmark for static malware detection on Windows Portable Executable (PE) files.
The original EMBER dataset was introduced by Endgame / Elastic as an open benchmark for machine-learning-based malware detection using only static PE-derived features, without executing binaries. This cleaned release preserves that purpose while making the… See the full description on the dataset page: https://huggingface.co/datasets/it4lia/EMBER_cleaned.fineweb-2-vie-2022-cleanedmy_dataset_cleaned
my_dataset
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
数据集信息
总episodes数: 47 (原48个,已移除episode_000000)
总帧数: 8,918
任务: 抓取立方体并放入盒子
机器人: so-100
帧率: 30 FPS
数据质量说明
注意: 原始数据集中的第0个episode (episode_000000) 由于视频质量问题已被移除。当前数据集从episode_000001开始,包含47个高质量的episode。… See the full description on the dataset page: https://huggingface.co/datasets/myzxyz/my_dataset_cleaned.fineweb-edu-cleaned-simplifiedusda-fdc-foods-cleaned
Comprehensive & Cleaned USDA Foods Nutrition Dataset
Dataset Summary
This dataset is a cleaned, de-duplicated, and enhanced version of the USDA's FoodData Central (FDC) database, combining Branded Foods, Foundation Foods (generic), and SR Legacy data into a single, analysis-ready file. It is designed to be a robust resource for nutritional analysis, machine learning, and food-related applications.
The raw USDA data is spread across dozens of CSV files, contains numerous… See the full description on the dataset page: https://huggingface.co/datasets/omid5/usda-fdc-foods-cleaned.green-only-200-cleanedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jayp132/green-only-200-cleaned.
