datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
high_quality_foldingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1200,
"total_frames": 3254196,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1200"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/high_quality_folding.video-quality-scored
Image-to-Video Quality-Scored Clips
A collection of prompted image-to-video samples with quality-evaluation metadata.
Each sample pairs a first frame (the I2V conditioning image) with one or both
of:
a generated video produced by a video model from the first frame + prompt
an original clip (the reference/source video the prompt was authored around)
A subset of the samples also carry per-clip quality scores: an overall
quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.truthfulness_high_quality
Dataset Card for "truthfulness_high_quality"
More Information needed
argument_quality_ranking_30k
Dataset Card for Argument-Quality-Ranking-30k Dataset
Dataset Summary
Argument Quality Ranking
The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets.
The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis.
Argument Topic
This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.soc139-quality-sidecars
soc139-quality-sidecars
Mirror of the R2 prefix soc139-quality-sidecars/ from soc127-dedup (Cloudflare R2) into a private
HF Dataset. Generated by scripts/handoff/mirror_r2_sidecar_prefix.py on
2026-05-23.
Each row in the source R2 parquets is preserved, with one additional column appended:
source_shard_path — the R2 object key the row came from. This lets you reconstruct
the per-shard view if needed.
Counts
Source (R2)
Mirror (this dataset)
Files
58… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/soc139-quality-sidecars.cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia
fineweb-edu-highest-quality-2025
FineWeb-Edu Highest Quality Dataset (2025 Collection)
Dataset Summary
This dataset contains 4.17 billion tokens of the highest quality educational content, carefully filtered from the FineWeb-Edu dataset's 2025 Common Crawl snapshots. This represents the cream of the crop - only the top ~2% of documents that meet strict quality criteria.
Key Statistics
Total Tokens: 4,176,738,951
Total Documents: 1,477,151
Average Tokens per Document: 2,827
Storage Size: ~11… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/fineweb-edu-highest-quality-2025.air-quality
Synthetic Kolkata Air Quality & Meteorology 100M
A reproducible fully synthetic spatiotemporal benchmark inspired by broad Kolkata, West Bengal climatological and air-quality behavior. The release contains exactly 100,000,000 station-hour rows from 1,000 explicitly synthetic sensor sites.
This is not an official CPCB/WBPCB/IMD monitoring archive and is not a reconstruction of historical measurements. Synthetic coordinates, event effects and station classes are benchmark… See the full description on the dataset page: https://huggingface.co/datasets/neuralsorcerer/air-quality.wine-qualitylevel2_final_quality2_augmentedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 2366,
"total_frames": 6205242,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2366"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality2_augmented.QuALITY
Dataset Card for "QuALITY"
@article{bowman2022quality,
title={QuALITY: Question Answering with Long Input Texts, Yes!},
author={Bowman, Samuel R and Chen, Angelica and He, He and Joshi, Nitish and Ma, Johnny and Nangia, Nikita and Padmakumar, Vishakh and Pang, Richard Yuanzhe and Parrish, Alicia and Phang, Jason and others},
journal={NAACL 2022},
year={2022}
}
FineWeb-Edu-Quality4plus
📘 FineWeb-Edu-Quality4plus
Overview
FineWeb-Edu-Quality4plus is a high-quality filtered subset of the original
HuggingFaceFW/fineweb-edu dataset (ODC-By License).
This subset retains only samples with:
quality_score ≥ 4
The goal is to provide a cleaner and more reliable dataset suitable for
language model pre-training, instruction tuning, education-related NLP,
and quality-sensitive downstream tasks.
This work is independent and not affiliated with the official FineWeb… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/FineWeb-Edu-Quality4plus.Quality-Control-App-Amazon-Big-Data-2023METRAQ-Air-Quality
The METRAQ air quality dataset
This is the official dataset repository for the METRAQ air quality dataset.
METRAQ air quality is an air quality dataset comprising hourly measurements of up to 14 pollutants from January 1, 2001, to December 31, 2024. In addition, the dataset has been spatially and temporally aligned with up to seven meteorological parameters (available since January 1, 2019) and three traffic monitoring metrics, aggregated using five different interpolation methods… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/METRAQ-Air-Quality.level2_final_qualityThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 1171,
"total_frames": 2940342,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1171"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality.dolma_20bn_cc_high_qualityRecursive-Task-Synthesis-Quality-1K
Recursive Task Synthesis Quality 1K
This dataset contains 1,000 quality-selected, validated command-line task
instances. It is a curated subset of the
Recursive Task Synthesis dataset.
Public task and group identifiers are opaque and stable across both datasets.
Selection
The subset was selected from 37,484 validated tasks using structural and safety
checks, two-pass semantic review, strict gates for instruction clarity,
instruction-verifier alignment, verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Quality-1K.scientific-quality-score-predictionDatasets related to the task of Scholarly Document Quality Prediction (SDQP).
Each sample is an academic paper for which either the citation count or the review score can be predicted (depending on availability).
ACL-OCL Extended
A dataset for citation count prediction only, based on the ACL-OCL dataset.
Extended with updated citation counts, references and annotated research hypothesis.
OpenReview (Last Update: 1.1.2025)
A dataset for review score and citation count… See the full description on the dataset page: https://huggingface.co/datasets/nhop/scientific-quality-score-prediction.africa-air-quality-all
African Air Pollution Source Apportionment | Africa (World Health Organization)
Size category: 100K<n<1M - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-air-quality-all.level12_quality0_2026-02-08This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 258,
"total_frames": 595888,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:258"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_quality0_2026-02-08.wikipedia_quality_wikirankDatasets with WikiRank quality score as of 1 August 2024 for 47 million Wikipedia articles in 55 language versions (also simplified version for each language in separate files).
The WikiRank quality score is a metric designed to assess the overall quality of a Wikipedia article. Although its specific algorithm can vary depending on the implementation, the score typically combines several key features of the Wikipedia article.
Why It’s Important
Enhances Trust: For readers and… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/wikipedia_quality_wikirank.cosmopedia_quality_score_v1v2pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.prompt-quality
Prompt Quality Assessment
Prompt quality strongly affects how well large language models (LLMs) perform, especially when user inputs are vague or incomplete. A good prompt is clear, specific, and complete, giving the model enough relevant context to produce accurate and useful responses.
This report describes a dataset created by evaluating prompts with several different LLMs. These evaluations can be used to train prompt-quality classifiers and to improve methods for prompt… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-quality.robotics-quality-leaderboard
Robotics Dataset Quality Leaderboard
What This Is
This repository hosts an automatically-updated quality leaderboard for robotics
imitation-learning datasets on HuggingFace. Each dataset is scored by the
HaptalAI quality scorer, an
open-source tool that streams a sample of episodes from a dataset, detects the
available sensor schema, runs a suite of failure-detection checks, and computes
a single 0–100 quality score. The leaderboard is intended as a first-pass… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/robotics-quality-leaderboard.french-administrative-hierarchy-data-quality
French Administrative Hierarchy Data Quality Benchmark
A reproducible benchmark for evaluating the validation, classification, and repair of French administrative geographic records.
The dataset is derived from the INSEE Code officiel géographique (COG) 2026 and focuses on the hierarchical relationship:
Region → Department → Commune
Important: commune_code is an INSEE/COG administrative identifier, not a postal code.This dataset is not intended for postal-address validation or… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/french-administrative-hierarchy-data-quality.polymarket-settlement-quality-register
Polymarket settlement-quality register
Frozen summaries of 123,499 settled UMA requests, window 2023-12-05 to 2026-08-12. Among settled disputes, 7.12% changed the proposal. Group summaries cover category and rule-text features.
Files and viewer
The viewer loads the canonical aggregate snapshot only. The dated files preserve export history and are not independent observations.
Method and source
See the embedded metadata and repository inventory.… See the full description on the dataset page: https://huggingface.co/datasets/ailinsun/polymarket-settlement-quality-register.where-quality-breaks-results
Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization — reported result summary
This repository contains an author-maintained, machine-readable summary of the key quantitative values reported in Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization.
Scope: this is a small table-level result summary. It is not the underlying training corpus, evaluation corpus, model code, checkpoint, benchmark release, or a… See the full description on the dataset page: https://huggingface.co/datasets/aogavrilov/where-quality-breaks-results.Jerusalem-Air-Quality-Shabbat
Jerusalem Air Quality — Weekday vs Shabbat
Twelve months of 5-minute-resolution air-quality readings from 12 monitoring
stations across Jerusalem, formatted for analysing the weekday vs Friday vs
Shabbat (Saturday) pattern in urban pollution.
📊 Companion analysis repository:
danielrosehill/JLM-Air-Quality-Analysis
— exact halachic relabeling, cross-city control analysis (London + New
York), and the figure set summarised below. Reproducible scripts included.
The motivation:… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Jerusalem-Air-Quality-Shabbat.swe-bench-trajectory-quality-subsets
SWE-bench Trajectory Quality Subsets
Curated subsets of nebius/SWE-rebench-openhands-trajectories constructed using the v3 quality scoring framework for fine-tuning evaluation.
Subsets Overview
Subset
Size
Selection
Mean Score
Resolved Rate
Purpose
Ablation-NoB2-500
500
Top 500 with Efficiency = B3 alone (drop B2 error_retry)
0.6410
100%
Ablation study
Ablation-NoB3-500
500
Top 500 with Efficiency = B2 alone (drop B3 step_count_ratio)
0.7253
100%
Ablation… See the full description on the dataset page: https://huggingface.co/datasets/davongluck/swe-bench-trajectory-quality-subsets.
