CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lerobot /high_quality_foldingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 1200, "total_frames": 3254196, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1200"}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/high_quality_folding.tabularrobotics1M<n<10M6 likes7.1k downloads7mo agoHugging Face02mohantesting /video-quality-scored Image-to-Video Quality-Scored Clips A collection of prompted image-to-video samples with quality-evaluation metadata. Each sample pairs a first frame (the I2V conditioning image) with one or both of: a generated video produced by a video model from the first frame + prompt an original clip (the reference/source video the prompt was authored around) A subset of the samples also carry per-clip quality scores: an overall quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.imagetext-to-video1K<n<10K0 likes6.5k downloads3mo agoHugging Face03notrichardren /truthfulness_high_quality Dataset Card for "truthfulness_high_quality" More Information needed tabular100K<n<1M2 likes3k downloads3y agoHugging Face04ibm-research /argument_quality_ranking_30k Dataset Card for Argument-Quality-Ranking-30k Dataset Dataset Summary Argument Quality Ranking The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets. The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis. Argument Topic This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.tabulartext-classification10K<n<100K13 likes1.7k downloads3y agoHugging Face05HCAI-Lab-GT /soc139-quality-sidecars soc139-quality-sidecars Mirror of the R2 prefix soc139-quality-sidecars/ from soc127-dedup (Cloudflare R2) into a private HF Dataset. Generated by scripts/handoff/mirror_r2_sidecar_prefix.py on 2026-05-23. Each row in the source R2 parquets is preserved, with one additional column appended: source_shard_path — the R2 object key the row came from. This lets you reconstruct the per-shard view if needed. Counts Source (R2) Mirror (this dataset) Files 58… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/soc139-quality-sidecars.tabular1B<n<10B0 likes1.6k downloads4mo agoHugging Face06kenhktsui /cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia tabular10M<n<100M0 likes1.4k downloads2y agoHugging Face07Yxanul /fineweb-edu-highest-quality-2025 FineWeb-Edu Highest Quality Dataset (2025 Collection) Dataset Summary This dataset contains 4.17 billion tokens of the highest quality educational content, carefully filtered from the FineWeb-Edu dataset's 2025 Common Crawl snapshots. This represents the cream of the crop - only the top ~2% of documents that meet strict quality criteria. Key Statistics Total Tokens: 4,176,738,951 Total Documents: 1,477,151 Average Tokens per Document: 2,827 Storage Size: ~11… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/fineweb-edu-highest-quality-2025.tabular1M<n<10M0 likes1.1k downloads1y agoHugging Face08neuralsorcerer /air-quality Synthetic Kolkata Air Quality & Meteorology 100M A reproducible fully synthetic spatiotemporal benchmark inspired by broad Kolkata, West Bengal climatological and air-quality behavior. The release contains exactly 100,000,000 station-hour rows from 1,000 explicitly synthetic sensor sites. This is not an official CPCB/WBPCB/IMD monitoring archive and is not a reconstruction of historical measurements. Synthetic coordinates, event effects and station classes are benchmark… See the full description on the dataset page: https://huggingface.co/datasets/neuralsorcerer/air-quality.tabulartime-series-forecasting100M<n<1B0 likes898 downloads5d agoHugging Face09codesignal /wine-qualitytabular1K<n<10K2 likes891 downloads11mo agoHugging Face10lerobot-data-collection /level2_final_quality2_augmentedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 2366, "total_frames": 6205242, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:2366" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality2_augmented.tabularrobotics1M<n<10M0 likes891 downloads7mo agoHugging Face11tasksource /QuALITY Dataset Card for "QuALITY" @article{bowman2022quality, title={QuALITY: Question Answering with Long Input Texts, Yes!}, author={Bowman, Samuel R and Chen, Angelica and He, He and Joshi, Nitish and Ma, Johnny and Nangia, Nikita and Padmakumar, Vishakh and Pang, Richard Yuanzhe and Parrish, Alicia and Phang, Jason and others}, journal={NAACL 2022}, year={2022} } tabular1K<n<10K1 likes788 downloads2y agoHugging Face12Morton-Li /FineWeb-Edu-Quality4plus 📘 FineWeb-Edu-Quality4plus Overview FineWeb-Edu-Quality4plus is a high-quality filtered subset of the original HuggingFaceFW/fineweb-edu dataset (ODC-By License). This subset retains only samples with: quality_score ≥ 4 The goal is to provide a cleaner and more reliable dataset suitable for language model pre-training, instruction tuning, education-related NLP, and quality-sensitive downstream tasks. This work is independent and not affiliated with the official FineWeb… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/FineWeb-Edu-Quality4plus.tabulartext-generation10M<n<100M1 likes583 downloads9mo agoHugging Face13gamusa /Quality-Control-App-Amazon-Big-Data-2023tabularn<1K0 likes525 downloads4mo agoHugging Face14dmariaa70 /METRAQ-Air-Quality The METRAQ air quality dataset This is the official dataset repository for the METRAQ air quality dataset. METRAQ air quality is an air quality dataset comprising hourly measurements of up to 14 pollutants from January 1, 2001, to December 31, 2024. In addition, the dataset has been spatially and temporally aligned with up to seven meteorological parameters (available since January 1, 2019) and three traffic monitoring metrics, aggregated using five different interpolation methods… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/METRAQ-Air-Quality.tabulartime-series-forecasting10M<n<100M2 likes476 downloads4mo agoHugging Face15lerobot-data-collection /level2_final_qualityThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 1171, "total_frames": 2940342, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1171" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality.tabularrobotics1M<n<10M0 likes472 downloads8mo agoHugging Face16orionweller /dolma_20bn_cc_high_qualitytabular10M<n<100M0 likes448 downloads2y agoHugging Face17Zhongzhi1228 /Recursive-Task-Synthesis-Quality-1K Recursive Task Synthesis Quality 1K This dataset contains 1,000 quality-selected, validated command-line task instances. It is a curated subset of the Recursive Task Synthesis dataset. Public task and group identifiers are opaque and stable across both datasets. Selection The subset was selected from 37,484 validated tasks using structural and safety checks, two-pass semantic review, strict gates for instruction clarity, instruction-verifier alignment, verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Quality-1K.tabularreinforcement-learning1K<n<10K0 likes385 downloads1mo agoHugging Face18nhop /scientific-quality-score-predictionDatasets related to the task of Scholarly Document Quality Prediction (SDQP). Each sample is an academic paper for which either the citation count or the review score can be predicted (depending on availability). ACL-OCL Extended A dataset for citation count prediction only, based on the ACL-OCL dataset. Extended with updated citation counts, references and annotated research hypothesis. OpenReview (Last Update: 1.1.2025) A dataset for review score and citation count… See the full description on the dataset page: https://huggingface.co/datasets/nhop/scientific-quality-score-prediction.tabulartext-classification100K<n<1M0 likes299 downloads1y agoHugging Face19electricsheepafrica /africa-air-quality-all African Air Pollution Source Apportionment | Africa (World Health Organization) Size category: 100K<n<1M - Formats: csv - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health datasets help researchers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-air-quality-all.texttabular-classification100K<n<1M0 likes287 downloads1mo agoHugging Face20lerobot-data-collection /level12_quality0_2026-02-08This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 258, "total_frames": 595888, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:258" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_quality0_2026-02-08.tabularrobotics100K<n<1M0 likes256 downloads8mo agoHugging Face21lewoniewski /wikipedia_quality_wikirankDatasets with WikiRank quality score as of 1 August 2024 for 47 million Wikipedia articles in 55 language versions (also simplified version for each language in separate files). The WikiRank quality score is a metric designed to assess the overall quality of a Wikipedia article. Although its specific algorithm can vary depending on the implementation, the score typically combines several key features of the Wikipedia article. Why It’s Important Enhances Trust: For readers and… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/wikipedia_quality_wikirank.tabular10M<n<100M4 likes223 downloads2y agoHugging Face22kenhktsui /cosmopedia_quality_score_v1v2tabular10M<n<100M0 likes159 downloads2y agoHugging Face23lapa-llm /pretraining-high-quality Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness of the text lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.tabulartext-generation10M<n<100M0 likes139 downloads10mo agoHugging Face24agentlans /prompt-quality Prompt Quality Assessment Prompt quality strongly affects how well large language models (LLMs) perform, especially when user inputs are vague or incomplete. A good prompt is clear, specific, and complete, giving the model enough relevant context to produce accurate and useful responses. This report describes a dataset created by evaluating prompts with several different LLMs. These evaluations can be used to train prompt-quality classifiers and to improve methods for prompt… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-quality.tabulartext-classification10K<n<100K2 likes134 downloads6mo agoHugging Face25HaptalAI /robotics-quality-leaderboard Robotics Dataset Quality Leaderboard What This Is This repository hosts an automatically-updated quality leaderboard for robotics imitation-learning datasets on HuggingFace. Each dataset is scored by the HaptalAI quality scorer, an open-source tool that streams a sample of episodes from a dataset, detects the available sensor schema, runs a suite of failure-detection checks, and computes a single 0–100 quality score. The leaderboard is intended as a first-pass… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/robotics-quality-leaderboard.tabularroboticsn<1K0 likes132 downloads15d agoHugging Face26Jaymerry /french-administrative-hierarchy-data-quality French Administrative Hierarchy Data Quality Benchmark A reproducible benchmark for evaluating the validation, classification, and repair of French administrative geographic records. The dataset is derived from the INSEE Code officiel géographique (COG) 2026 and focuses on the hierarchical relationship: Region → Department → Commune Important: commune_code is an INSEE/COG administrative identifier, not a postal code.This dataset is not intended for postal-address validation or… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/french-administrative-hierarchy-data-quality.tabulartabular-classification100K<n<1M0 likes124 downloads2mo agoHugging Face27ailinsun /polymarket-settlement-quality-register Polymarket settlement-quality register Frozen summaries of 123,499 settled UMA requests, window 2023-12-05 to 2026-08-12. Among settled disputes, 7.12% changed the proposal. Group summaries cover category and rule-text features. Files and viewer The viewer loads the canonical aggregate snapshot only. The dated files preserve export history and are not independent observations. Method and source See the embedded metadata and repository inventory.… See the full description on the dataset page: https://huggingface.co/datasets/ailinsun/polymarket-settlement-quality-register.tabularn<1K0 likes123 downloads9d agoHugging Face28aogavrilov /where-quality-breaks-results Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization — reported result summary This repository contains an author-maintained, machine-readable summary of the key quantitative values reported in Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization. Scope: this is a small table-level result summary. It is not the underlying training corpus, evaluation corpus, model code, checkpoint, benchmark release, or a… See the full description on the dataset page: https://huggingface.co/datasets/aogavrilov/where-quality-breaks-results.tabularn<1K0 likes122 downloads2mo agoHugging Face29danielrosehill /Jerusalem-Air-Quality-Shabbat Jerusalem Air Quality — Weekday vs Shabbat Twelve months of 5-minute-resolution air-quality readings from 12 monitoring stations across Jerusalem, formatted for analysing the weekday vs Friday vs Shabbat (Saturday) pattern in urban pollution. 📊 Companion analysis repository: danielrosehill/JLM-Air-Quality-Analysis — exact halachic relabeling, cross-city control analysis (London + New York), and the figure set summarised below. Reproducible scripts included. The motivation:… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Jerusalem-Air-Quality-Shabbat.image1M<n<10M0 likes114 downloads5mo agoHugging Face30davongluck /swe-bench-trajectory-quality-subsets SWE-bench Trajectory Quality Subsets Curated subsets of nebius/SWE-rebench-openhands-trajectories constructed using the v3 quality scoring framework for fine-tuning evaluation. Subsets Overview Subset Size Selection Mean Score Resolved Rate Purpose Ablation-NoB2-500 500 Top 500 with Efficiency = B3 alone (drop B2 error_retry) 0.6410 100% Ablation study Ablation-NoB3-500 500 Top 500 with Efficiency = B2 alone (drop B3 step_count_ratio) 0.7253 100% Ablation… See the full description on the dataset page: https://huggingface.co/datasets/davongluck/swe-bench-trajectory-quality-subsets.tabulartext-generation10K<n<100K2 likes113 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.