datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
high_quality_foldingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1200,
"total_frames": 3254196,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/high_quality_folding.quality
Dataset Card for "quality"
More Information needed
truthfulness_high_quality
Dataset Card for "truthfulness_high_quality"
More Information needed
TM-DATA_quality_score_v1
Dataset Card for "TM-DATA_quality_score_v1"
Adding quality score v1 to Locutusque/TM-DATA
More Information needed
cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia
low_quality_call_voice_preprocessed
Dataset Card for "low_quality_call_voice_preprocessed"
More Information needed
openwebtext_quality_score_v1
Dataset Card for "openwebtext_quality_score_v1"
Adding quality score v1 to Skylion007/openwebtext
More Information needed
soc139-quality-sidecars
soc139-quality-sidecars
Mirror of the R2 prefix soc139-quality-sidecars/ from soc127-dedup (Cloudflare R2) into a private
HF Dataset. Generated by scripts/handoff/mirror_r2_sidecar_prefix.py on
2026-05-23.
Each row in the source R2 parquets is preserved, with one additional column appended:
source_shard_path — the R2 object key the row came from. This lets you reconstruct
the per-shard view if needed.
Counts
Source (R2)
Mirror (this dataset)
Files
58… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/soc139-quality-sidecars.fineweb-edu-highest-quality-2025
FineWeb-Edu Highest Quality Dataset (2025 Collection)
Dataset Summary
This dataset contains 4.17 billion tokens of the highest quality educational content, carefully filtered from the FineWeb-Edu dataset's 2025 Common Crawl snapshots. This represents the cream of the crop - only the top ~2% of documents that meet strict quality criteria.
Key Statistics
Total Tokens: 4,176,738,951
Total Documents: 1,477,151
Average Tokens per Document: 2,827
Storage Size: ~11… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/fineweb-edu-highest-quality-2025.air-quality
Synthetic Kolkata Air Quality & Meteorology 100M
A reproducible fully synthetic spatiotemporal benchmark inspired by broad Kolkata, West Bengal climatological and air-quality behavior. The release contains exactly 100,000,000 station-hour rows from 1,000 explicitly synthetic sensor sites.
This is not an official CPCB/WBPCB/IMD monitoring archive and is not a reconstruction of historical measurements. Synthetic coordinates, event effects and station classes are benchmark… See the full description on the dataset page: https://huggingface.co/datasets/neuralsorcerer/air-quality.level2_final_quality2_augmentedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 2366,
"total_frames": 6205242,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2366"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality2_augmented.QuALITY
Dataset Card for "QuALITY"
@article{bowman2022quality,
title={QuALITY: Question Answering with Long Input Texts, Yes!},
author={Bowman, Samuel R and Chen, Angelica and He, He and Joshi, Nitish and Ma, Johnny and Nangia, Nikita and Padmakumar, Vishakh and Pang, Richard Yuanzhe and Parrish, Alicia and Phang, Jason and others},
journal={NAACL 2022},
year={2022}
}
wine-qualityFineWeb-Edu-Quality4plus
📘 FineWeb-Edu-Quality4plus
Overview
FineWeb-Edu-Quality4plus is a high-quality filtered subset of the original
HuggingFaceFW/fineweb-edu dataset (ODC-By License).
This subset retains only samples with:
quality_score ≥ 4
The goal is to provide a cleaner and more reliable dataset suitable for
language model pre-training, instruction tuning, education-related NLP,
and quality-sensitive downstream tasks.
This work is independent and not affiliated with the official FineWeb… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/FineWeb-Edu-Quality4plus.refinedweb-3m_quality_score_v1
Dataset Card for "refinedweb-3m_quality_score_v1"
Adding quality score v1 to mattymchen/refinedweb-3m
More Information needed
dolma_20bn_cc_high_qualitylevel2_final_qualityThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1171,
"total_frames": 2940342,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1171"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality.Recursive-Task-Synthesis-Quality-1K
Recursive Task Synthesis Quality 1K
This dataset contains 1,000 quality-selected, validated command-line task
instances. It is a curated subset of the
Recursive Task Synthesis dataset.
Public task and group identifiers are opaque and stable across both datasets.
Selection
The subset was selected from 37,484 validated tasks using structural and safety
checks, two-pass semantic review, strict gates for instruction clarity,
instruction-verifier alignment, verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Quality-1K.scientific-quality-score-predictionDatasets related to the task of Scholarly Document Quality Prediction (SDQP).
Each sample is an academic paper for which either the citation count or the review score can be predicted (depending on availability).
ACL-OCL Extended
A dataset for citation count prediction only, based on the ACL-OCL dataset.
Extended with updated citation counts, references and annotated research hypothesis.
OpenReview (Last Update: 1.1.2025)
A dataset for review score and citation count… See the full description on the dataset page: https://huggingface.co/datasets/nhop/scientific-quality-score-prediction.textbook_quality_programming
Dataset Card for "textbook_quality_programming"
Synthetic programming textbooks generated with GPT-3.5 and retrieval. Very high quality, aimed at being used in a phi replication. Currently 115M tokens. Covers many languages and technologies, with a bias towards python.
~10k of the books (65M tokens) use an older generation method, and average 6k tokens in length. ~1.5k books (50M tokens) use a newer generation method, with a more detailed outline, and average 33k tokens in… See the full description on the dataset page: https://huggingface.co/datasets/vikp/textbook_quality_programming.high-quality-cc-21b
high_quality
A high-quality English web text corpus extracted from Common Crawl WARC files using an
LLM-based extraction and quality pipeline.
Dataset Summary
high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl
WARC records are passed through an LLM-based extractor that strips boilerplate and recovers
the main content, then filtered to retain only documents in the "high_quality" band,
deduplicated (exact + fuzzy), and… See the full description on the dataset page: https://huggingface.co/datasets/MichaelR207/high-quality-cc-21b.DUE-v2-quality-filtered
removed markdown formatting, links, emojis, multiple white spaces
replaced user and channel tags with @user and #channel
deduplicated (not fuzzy only exact)
basic filtering:
word count filter(min_words=3, max_words=200)
mean word length filter(min_mean_word_length=2.0, max_mean_word_length=12.0)
long word filter(max_word_length=200)
whitespace filter (max_white_space_ratio=0.25)
open-license-corpus_quality_score_v1AIGI-Detection-Quality-Paradox
AIGI-Detection-Quality-Paradox Dataset
The dataset was created for the paper:
Are High-Quality AI-Generated Images More Difficult for Models to Detect?Authors: Yao Xiao, Binbin Yang, Weiyan Chen, Jiahao Chen, Zijie Cao, Ziyi Dong, Xiangyang Ji, Liang Lin, Wei Ke, Pengxu WeiAccepted by: ICML 2025Paper Link: https://openreview.net/forum?id=sKYdVKE1tS
Overview
This dataset contains diverse realistic AI-generated images from multiple generators, each with detailed… See the full description on the dataset page: https://huggingface.co/datasets/Coxy7/AIGI-Detection-Quality-Paradox.level12_quality0_2026-02-08This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 258,
"total_frames": 595888,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:258"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_quality0_2026-02-08.low_quality_call_voice
Dataset Card for "low_quality_call_voice"
More Information needed
rice-quality-assessment-qwen-vlhigh-quality-midjouney-srefs
Midjourney Image Scraper & Dataset Creator
A complete toolkit for scraping Midjourney images, generating captions, and creating HuggingFace datasets with optional automatic upload to HuggingFace Hub.
🌟 Features
🔍 Web Scraping: Download images from midjourneysref.com with comprehensive error handling
🤖 AI Captioning: Automatic image captioning using Moondream API with auto-resume capability
✂️ Smart Cropping: AI-powered image cropping using OpenAI to optimize aspect… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/high-quality-midjouney-srefs.task1283_hrngo_quality_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1283_hrngo_quality_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1283_hrngo_quality_classification.uzbek-high-quality-10h
