datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
high_quality_foldingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1200,
"total_frames": 3254196,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1200"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/high_quality_folding.video-quality-scored
Image-to-Video Quality-Scored Clips
A collection of prompted image-to-video samples with quality-evaluation metadata.
Each sample pairs a first frame (the I2V conditioning image) with one or both
of:
a generated video produced by a video model from the first frame + prompt
an original clip (the reference/source video the prompt was authored around)
A subset of the samples also carry per-clip quality scores: an overall
quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.quality
Dataset Card for "quality"
More Information needed
high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/high-quality-invoice-images-for-ocr.truthfulness_high_quality
Dataset Card for "truthfulness_high_quality"
More Information needed
synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.Creative-Writing-High-Quality-1300x
Creative Writing - Part One (Shadow & Skeleton)
This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology.
Methodology: Shadow & Skeleton
Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach:
Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.TM-DATA_quality_score_v1
Dataset Card for "TM-DATA_quality_score_v1"
Adding quality score v1 to Locutusque/TM-DATA
More Information needed
sea-commoncrawl-high-qualityargument_quality_ranking_30k
Dataset Card for Argument-Quality-Ranking-30k Dataset
Dataset Summary
Argument Quality Ranking
The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets.
The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis.
Argument Topic
This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.PACE-Water-Qualitylow_quality_call_voice_preprocessed
Dataset Card for "low_quality_call_voice_preprocessed"
More Information needed
soc139-quality-sidecars
soc139-quality-sidecars
Mirror of the R2 prefix soc139-quality-sidecars/ from soc127-dedup (Cloudflare R2) into a private
HF Dataset. Generated by scripts/handoff/mirror_r2_sidecar_prefix.py on
2026-05-23.
Each row in the source R2 parquets is preserved, with one additional column appended:
source_shard_path — the R2 object key the row came from. This lets you reconstruct
the per-shard view if needed.
Counts
Source (R2)
Mirror (this dataset)
Files
58… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/soc139-quality-sidecars.cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia
openwebtext_quality_score_v1
Dataset Card for "openwebtext_quality_score_v1"
Adding quality score v1 to Skylion007/openwebtext
More Information needed
fineweb-edu-highest-quality-2025
FineWeb-Edu Highest Quality Dataset (2025 Collection)
Dataset Summary
This dataset contains 4.17 billion tokens of the highest quality educational content, carefully filtered from the FineWeb-Edu dataset's 2025 Common Crawl snapshots. This represents the cream of the crop - only the top ~2% of documents that meet strict quality criteria.
Key Statistics
Total Tokens: 4,176,738,951
Total Documents: 1,477,151
Average Tokens per Document: 2,827
Storage Size: ~11… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/fineweb-edu-highest-quality-2025.High-Quality-Code
High-Quality-Code: Synthetic + Real (MAXIMUM CODE)
A massive, high-quality code dataset built with maximum code philosophy – as much code as possible.
Components
Synthetic syntax-correction dataset – 5M+ examples across 33 languages (original code_syntax_dataset_1GB.csv)
Real high-quality code from GitHub – 500 top-starred repositories – BOTH zips and extracted source
Current Status: IN PROGRESS
Target: 500 repos
Currently uploaded: 12 extracted… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/High-Quality-Code.high-quality-english-sentences
High-Quality English Sentences
Dataset Description
This dataset contains a collection of high-quality English sentences sourced from C4 and FineWeb (not FineWeb-Edu). The sentences have been carefully filtered and processed to ensure quality and uniqueness.
"High-quality" means they're legible English and not spam, although they may still have spelling and grammar errors.
Source Data
Before filtering:
C4: 1 million sentences
FineWeb: 1 million sentences… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-english-sentences.air-quality
Synthetic Kolkata Air Quality & Meteorology 100M
A reproducible fully synthetic spatiotemporal benchmark inspired by broad Kolkata, West Bengal climatological and air-quality behavior. The release contains exactly 100,000,000 station-hour rows from 1,000 explicitly synthetic sensor sites.
This is not an official CPCB/WBPCB/IMD monitoring archive and is not a reconstruction of historical measurements. Synthetic coordinates, event effects and station classes are benchmark… See the full description on the dataset page: https://huggingface.co/datasets/neuralsorcerer/air-quality.wine-qualitylevel2_final_quality2_augmentedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 2366,
"total_frames": 6205242,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2366"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality2_augmented.high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/JuanfelipeX123/high-quality-invoice-images-for-ocr.QuALITY
Dataset Card for "QuALITY"
@article{bowman2022quality,
title={QuALITY: Question Answering with Long Input Texts, Yes!},
author={Bowman, Samuel R and Chen, Angelica and He, He and Joshi, Nitish and Ma, Johnny and Nangia, Nikita and Padmakumar, Vishakh and Pang, Richard Yuanzhe and Parrish, Alicia and Phang, Jason and others},
journal={NAACL 2022},
year={2022}
}
alma-avatar-quality-pilot-v1FineWeb-Edu-Quality4plus
📘 FineWeb-Edu-Quality4plus
Overview
FineWeb-Edu-Quality4plus is a high-quality filtered subset of the original
HuggingFaceFW/fineweb-edu dataset (ODC-By License).
This subset retains only samples with:
quality_score ≥ 4
The goal is to provide a cleaner and more reliable dataset suitable for
language model pre-training, instruction tuning, education-related NLP,
and quality-sensitive downstream tasks.
This work is independent and not affiliated with the official FineWeb… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/FineWeb-Edu-Quality4plus.high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/Shubhal829/high-quality-invoice-images-for-ocr.droid_success_high_quality
DROID Success (High-Quality Extrinsics)
Subset of ~17k successful episodes in DROID-COMMUNITY filtered for high-quality camera extrinsics.
Ported from raw 1.0.1 data at full resolution to LeRobotDataset v3.0 format (0.33 TiB | 3.6k inodes) with extra annotation from KarlP/droid.
Your browser does not support the video tag.
Dataset Structure
The external cameras are assigned to left and right views depending on the episode. For their extrinsics… See the full description on the dataset page: https://huggingface.co/datasets/jnogga/droid_success_high_quality.Quality-Control-App-Amazon-Big-Data-2023refinedweb-3m_quality_score_v1
Dataset Card for "refinedweb-3m_quality_score_v1"
Adding quality score v1 to mattymchen/refinedweb-3m
More Information needed
METRAQ-Air-Quality
The METRAQ air quality dataset
This is the official dataset repository for the METRAQ air quality dataset.
METRAQ air quality is an air quality dataset comprising hourly measurements of up to 14 pollutants from January 1, 2001, to December 31, 2024. In addition, the dataset has been spatially and temporally aligned with up to seven meteorological parameters (available since January 1, 2019) and three traffic monitoring metrics, aggregated using five different interpolation methods… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/METRAQ-Air-Quality.
