CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lerobot /high_quality_foldingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 1200, "total_frames": 3254196, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1200"}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/high_quality_folding.tabularrobotics1M<n<10M6 likes7.1k downloads7mo agoHugging Face02mohantesting /video-quality-scored Image-to-Video Quality-Scored Clips A collection of prompted image-to-video samples with quality-evaluation metadata. Each sample pairs a first frame (the I2V conditioning image) with one or both of: a generated video produced by a video model from the first frame + prompt an original clip (the reference/source video the prompt was authored around) A subset of the samples also carry per-clip quality scores: an overall quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.imagetext-to-video1K<n<10K0 likes6.5k downloads3mo agoHugging Face03emozilla /quality Dataset Card for "quality" More Information needed text1K<n<10K7 likes3.7k downloads3y agoHugging Face04Voxel51 /high-quality-invoice-images-for-ocr Dataset Card for high_quality_invoice_images_ocr This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/high-quality-invoice-images-for-ocr.image1K<n<10K6 likes3.5k downloads8mo agoHugging Face05notrichardren /truthfulness_high_quality Dataset Card for "truthfulness_high_quality" More Information needed tabular100K<n<1M2 likes3k downloads3y agoHugging Face06ProGamerGov /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M154 likes2.3k downloads2y agoHugging Face07Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.2k downloads2mo agoHugging Face08kenhktsui /TM-DATA_quality_score_v1 Dataset Card for "TM-DATA_quality_score_v1" Adding quality score v1 to Locutusque/TM-DATA More Information needed text1M<n<10M0 likes2k downloads3y agoHugging Face09sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes1.9k downloads2y agoHugging Face10ibm-research /argument_quality_ranking_30k Dataset Card for Argument-Quality-Ranking-30k Dataset Dataset Summary Argument Quality Ranking The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets. The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis. Argument Topic This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.tabulartext-classification10K<n<100K13 likes1.7k downloads3y agoHugging Face11giswqs /PACE-Water-Qualityimage1K<n<10K2 likes1.6k downloads5d agoHugging Face12INo0121 /low_quality_call_voice_preprocessed Dataset Card for "low_quality_call_voice_preprocessed" More Information needed 10K<n<100K1 likes1.6k downloads3y agoHugging Face13HCAI-Lab-GT /soc139-quality-sidecars soc139-quality-sidecars Mirror of the R2 prefix soc139-quality-sidecars/ from soc127-dedup (Cloudflare R2) into a private HF Dataset. Generated by scripts/handoff/mirror_r2_sidecar_prefix.py on 2026-05-23. Each row in the source R2 parquets is preserved, with one additional column appended: source_shard_path — the R2 object key the row came from. This lets you reconstruct the per-shard view if needed. Counts Source (R2) Mirror (this dataset) Files 58… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/soc139-quality-sidecars.tabular1B<n<10B0 likes1.6k downloads4mo agoHugging Face14kenhktsui /cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia tabular10M<n<100M0 likes1.4k downloads2y agoHugging Face15kenhktsui /openwebtext_quality_score_v1 Dataset Card for "openwebtext_quality_score_v1" Adding quality score v1 to Skylion007/openwebtext More Information needed texttext-generation1M<n<10M0 likes1.2k downloads3y agoHugging Face16Yxanul /fineweb-edu-highest-quality-2025 FineWeb-Edu Highest Quality Dataset (2025 Collection) Dataset Summary This dataset contains 4.17 billion tokens of the highest quality educational content, carefully filtered from the FineWeb-Edu dataset's 2025 Common Crawl snapshots. This represents the cream of the crop - only the top ~2% of documents that meet strict quality criteria. Key Statistics Total Tokens: 4,176,738,951 Total Documents: 1,477,151 Average Tokens per Document: 2,827 Storage Size: ~11… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/fineweb-edu-highest-quality-2025.tabular1M<n<10M0 likes1.1k downloads1y agoHugging Face17Corpus-NZ /High-Quality-Code High-Quality-Code: Synthetic + Real (MAXIMUM CODE) A massive, high-quality code dataset built with maximum code philosophy – as much code as possible. Components Synthetic syntax-correction dataset – 5M+ examples across 33 languages (original code_syntax_dataset_1GB.csv) Real high-quality code from GitHub – 500 top-starred repositories – BOTH zips and extracted source Current Status: IN PROGRESS Target: 500 repos Currently uploaded: 12 extracted… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/High-Quality-Code.100M<n<1B0 likes1.1k downloads13d agoHugging Face18agentlans /high-quality-english-sentences High-Quality English Sentences Dataset Description This dataset contains a collection of high-quality English sentences sourced from C4 and FineWeb (not FineWeb-Edu). The sentences have been carefully filtered and processed to ensure quality and uniqueness. "High-quality" means they're legible English and not spam, although they may still have spelling and grammar errors. Source Data Before filtering: C4: 1 million sentences FineWeb: 1 million sentences… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-english-sentences.texttext-classification1M<n<10M38 likes918 downloads2y agoHugging Face19neuralsorcerer /air-quality Synthetic Kolkata Air Quality & Meteorology 100M A reproducible fully synthetic spatiotemporal benchmark inspired by broad Kolkata, West Bengal climatological and air-quality behavior. The release contains exactly 100,000,000 station-hour rows from 1,000 explicitly synthetic sensor sites. This is not an official CPCB/WBPCB/IMD monitoring archive and is not a reconstruction of historical measurements. Synthetic coordinates, event effects and station classes are benchmark… See the full description on the dataset page: https://huggingface.co/datasets/neuralsorcerer/air-quality.tabulartime-series-forecasting100M<n<1B0 likes898 downloads5d agoHugging Face20codesignal /wine-qualitytabular1K<n<10K2 likes891 downloads11mo agoHugging Face21lerobot-data-collection /level2_final_quality2_augmentedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 2366, "total_frames": 6205242, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:2366" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality2_augmented.tabularrobotics1M<n<10M0 likes891 downloads7mo agoHugging Face22JuanfelipeX123 /high-quality-invoice-images-for-ocr Dataset Card for high_quality_invoice_images_ocr This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import… See the full description on the dataset page: https://huggingface.co/datasets/JuanfelipeX123/high-quality-invoice-images-for-ocr.image1K<n<10K0 likes801 downloads1mo agoHugging Face23tasksource /QuALITY Dataset Card for "QuALITY" @article{bowman2022quality, title={QuALITY: Question Answering with Long Input Texts, Yes!}, author={Bowman, Samuel R and Chen, Angelica and He, He and Joshi, Nitish and Ma, Johnny and Nangia, Nikita and Padmakumar, Vishakh and Pang, Richard Yuanzhe and Parrish, Alicia and Phang, Jason and others}, journal={NAACL 2022}, year={2022} } tabular1K<n<10K1 likes788 downloads2y agoHugging Face24igorcouto /alma-avatar-quality-pilot-v1video4 likes733 downloads7d agoHugging Face25Morton-Li /FineWeb-Edu-Quality4plus 📘 FineWeb-Edu-Quality4plus Overview FineWeb-Edu-Quality4plus is a high-quality filtered subset of the original HuggingFaceFW/fineweb-edu dataset (ODC-By License). This subset retains only samples with: quality_score ≥ 4 The goal is to provide a cleaner and more reliable dataset suitable for language model pre-training, instruction tuning, education-related NLP, and quality-sensitive downstream tasks. This work is independent and not affiliated with the official FineWeb… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/FineWeb-Edu-Quality4plus.tabulartext-generation10M<n<100M1 likes583 downloads9mo agoHugging Face26Shubhal829 /high-quality-invoice-images-for-ocr Dataset Card for high_quality_invoice_images_ocr This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import… See the full description on the dataset page: https://huggingface.co/datasets/Shubhal829/high-quality-invoice-images-for-ocr.image1K<n<10K0 likes563 downloads3mo agoHugging Face27jnogga /droid_success_high_quality DROID Success (High-Quality Extrinsics) Subset of ~17k successful episodes in DROID-COMMUNITY filtered for high-quality camera extrinsics. Ported from raw 1.0.1 data at full resolution to LeRobotDataset v3.0 format (0.33 TiB | 3.6k inodes) with extra annotation from KarlP/droid. Your browser does not support the video tag. Dataset Structure The external cameras are assigned to left and right views depending on the episode. For their extrinsics… See the full description on the dataset page: https://huggingface.co/datasets/jnogga/droid_success_high_quality.videorobotics1K<n<10K2 likes555 downloads7mo agoHugging Face28gamusa /Quality-Control-App-Amazon-Big-Data-2023tabularn<1K0 likes525 downloads4mo agoHugging Face29kenhktsui /refinedweb-3m_quality_score_v1 Dataset Card for "refinedweb-3m_quality_score_v1" Adding quality score v1 to mattymchen/refinedweb-3m More Information needed texttext-generation1M<n<10M0 likes510 downloads3y agoHugging Face30dmariaa70 /METRAQ-Air-Quality The METRAQ air quality dataset This is the official dataset repository for the METRAQ air quality dataset. METRAQ air quality is an air quality dataset comprising hourly measurements of up to 14 pollutants from January 1, 2001, to December 31, 2024. In addition, the dataset has been spatially and temporally aligned with up to seven meteorological parameters (available since January 1, 2019) and three traffic monitoring metrics, aggregated using five different interpolation methods… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/METRAQ-Air-Quality.tabulartime-series-forecasting10M<n<100M2 likes476 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.