CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lerobot /high_quality_foldingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 1200, "total_frames": 3254196, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1200" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/high_quality_folding.tabularrobotics1M<n<10M6 likes7.2k downloads7mo agoHugging Face02emozilla /quality Dataset Card for "quality" More Information needed text1K<n<10K7 likes3.8k downloads3y agoHugging Face03notrichardren /truthfulness_high_quality Dataset Card for "truthfulness_high_quality" More Information needed tabular100K<n<1M2 likes3k downloads3y agoHugging Face04kenhktsui /TM-DATA_quality_score_v1 Dataset Card for "TM-DATA_quality_score_v1" Adding quality score v1 to Locutusque/TM-DATA More Information needed text1M<n<10M0 likes2.3k downloads3y agoHugging Face05kenhktsui /cosmopedia_quality_score_v2Adding quality score v2 to HuggingFaceTB/cosmopedia tabular10M<n<100M0 likes1.7k downloads2y agoHugging Face06INo0121 /low_quality_call_voice_preprocessed Dataset Card for "low_quality_call_voice_preprocessed" More Information needed 10K<n<100K1 likes1.6k downloads3y agoHugging Face07kenhktsui /openwebtext_quality_score_v1 Dataset Card for "openwebtext_quality_score_v1" Adding quality score v1 to Skylion007/openwebtext More Information needed texttext-generation1M<n<10M0 likes1.3k downloads3y agoHugging Face08HCAI-Lab-GT /soc139-quality-sidecars soc139-quality-sidecars Mirror of the R2 prefix soc139-quality-sidecars/ from soc127-dedup (Cloudflare R2) into a private HF Dataset. Generated by scripts/handoff/mirror_r2_sidecar_prefix.py on 2026-05-23. Each row in the source R2 parquets is preserved, with one additional column appended: source_shard_path — the R2 object key the row came from. This lets you reconstruct the per-shard view if needed. Counts Source (R2) Mirror (this dataset) Files 58… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/soc139-quality-sidecars.tabular1B<n<10B0 likes1.2k downloads4mo agoHugging Face09Yxanul /fineweb-edu-highest-quality-2025 FineWeb-Edu Highest Quality Dataset (2025 Collection) Dataset Summary This dataset contains 4.17 billion tokens of the highest quality educational content, carefully filtered from the FineWeb-Edu dataset's 2025 Common Crawl snapshots. This represents the cream of the crop - only the top ~2% of documents that meet strict quality criteria. Key Statistics Total Tokens: 4,176,738,951 Total Documents: 1,477,151 Average Tokens per Document: 2,827 Storage Size: ~11… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/fineweb-edu-highest-quality-2025.tabular1M<n<10M0 likes1k downloads1y agoHugging Face10neuralsorcerer /air-quality Synthetic Kolkata Air Quality & Meteorology 100M A reproducible fully synthetic spatiotemporal benchmark inspired by broad Kolkata, West Bengal climatological and air-quality behavior. The release contains exactly 100,000,000 station-hour rows from 1,000 explicitly synthetic sensor sites. This is not an official CPCB/WBPCB/IMD monitoring archive and is not a reconstruction of historical measurements. Synthetic coordinates, event effects and station classes are benchmark… See the full description on the dataset page: https://huggingface.co/datasets/neuralsorcerer/air-quality.tabulartime-series-forecasting100M<n<1B0 likes933 downloads8d agoHugging Face11lerobot-data-collection /level2_final_quality2_augmentedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 2366, "total_frames": 6205242, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:2366" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality2_augmented.tabularrobotics1M<n<10M0 likes905 downloads8mo agoHugging Face12tasksource /QuALITY Dataset Card for "QuALITY" @article{bowman2022quality, title={QuALITY: Question Answering with Long Input Texts, Yes!}, author={Bowman, Samuel R and Chen, Angelica and He, He and Joshi, Nitish and Ma, Johnny and Nangia, Nikita and Padmakumar, Vishakh and Pang, Richard Yuanzhe and Parrish, Alicia and Phang, Jason and others}, journal={NAACL 2022}, year={2022} } tabular1K<n<10K1 likes821 downloads2y agoHugging Face13codesignal /wine-qualitytabular1K<n<10K2 likes785 downloads11mo agoHugging Face14Morton-Li /FineWeb-Edu-Quality4plus 📘 FineWeb-Edu-Quality4plus Overview FineWeb-Edu-Quality4plus is a high-quality filtered subset of the original HuggingFaceFW/fineweb-edu dataset (ODC-By License). This subset retains only samples with: quality_score ≥ 4 The goal is to provide a cleaner and more reliable dataset suitable for language model pre-training, instruction tuning, education-related NLP, and quality-sensitive downstream tasks. This work is independent and not affiliated with the official FineWeb… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/FineWeb-Edu-Quality4plus.tabulartext-generation10M<n<100M1 likes669 downloads9mo agoHugging Face15kenhktsui /refinedweb-3m_quality_score_v1 Dataset Card for "refinedweb-3m_quality_score_v1" Adding quality score v1 to mattymchen/refinedweb-3m More Information needed texttext-generation1M<n<10M0 likes591 downloads3y agoHugging Face16orionweller /dolma_20bn_cc_high_qualitytabular10M<n<100M0 likes518 downloads2y agoHugging Face17lerobot-data-collection /level2_final_qualityThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 1171, "total_frames": 2940342, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1171" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level2_final_quality.tabularrobotics1M<n<10M0 likes491 downloads8mo agoHugging Face18Zhongzhi1228 /Recursive-Task-Synthesis-Quality-1K Recursive Task Synthesis Quality 1K This dataset contains 1,000 quality-selected, validated command-line task instances. It is a curated subset of the Recursive Task Synthesis dataset. Public task and group identifiers are opaque and stable across both datasets. Selection The subset was selected from 37,484 validated tasks using structural and safety checks, two-pass semantic review, strict gates for instruction clarity, instruction-verifier alignment, verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Quality-1K.tabularreinforcement-learning1K<n<10K0 likes428 downloads1mo agoHugging Face19nhop /scientific-quality-score-predictionDatasets related to the task of Scholarly Document Quality Prediction (SDQP). Each sample is an academic paper for which either the citation count or the review score can be predicted (depending on availability). ACL-OCL Extended A dataset for citation count prediction only, based on the ACL-OCL dataset. Extended with updated citation counts, references and annotated research hypothesis. OpenReview (Last Update: 1.1.2025) A dataset for review score and citation count… See the full description on the dataset page: https://huggingface.co/datasets/nhop/scientific-quality-score-prediction.tabulartext-classification100K<n<1M0 likes351 downloads1y agoHugging Face20vikp /textbook_quality_programming Dataset Card for "textbook_quality_programming" Synthetic programming textbooks generated with GPT-3.5 and retrieval. Very high quality, aimed at being used in a phi replication. Currently 115M tokens. Covers many languages and technologies, with a bias towards python. ~10k of the books (65M tokens) use an older generation method, and average 6k tokens in length. ~1.5k books (50M tokens) use a newer generation method, with a more detailed outline, and average 33k tokens in… See the full description on the dataset page: https://huggingface.co/datasets/vikp/textbook_quality_programming.text10K<n<100K182 likes322 downloads3y agoHugging Face21MichaelR207 /high-quality-cc-21b high_quality A high-quality English web text corpus extracted from Common Crawl WARC files using an LLM-based extraction and quality pipeline. Dataset Summary high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl WARC records are passed through an LLM-based extractor that strips boilerplate and recovers the main content, then filtered to retain only documents in the "high_quality" band, deduplicated (exact + fuzzy), and… See the full description on the dataset page: https://huggingface.co/datasets/MichaelR207/high-quality-cc-21b.texttext-generation10M<n<100M0 likes321 downloads3mo agoHugging Face22ManBib /DUE-v2-quality-filtered removed markdown formatting, links, emojis, multiple white spaces replaced user and channel tags with @user and #channel deduplicated (not fuzzy only exact) basic filtering: word count filter(min_words=3, max_words=200) mean word length filter(min_mean_word_length=2.0, max_mean_word_length=12.0) long word filter(max_word_length=200) whitespace filter (max_white_space_ratio=0.25) text100M<n<1B0 likes305 downloads1y agoHugging Face23kenhktsui /open-license-corpus_quality_score_v1text1M<n<10M0 likes295 downloads2y agoHugging Face24Coxy7 /AIGI-Detection-Quality-Paradox AIGI-Detection-Quality-Paradox Dataset The dataset was created for the paper: Are High-Quality AI-Generated Images More Difficult for Models to Detect?Authors: Yao Xiao, Binbin Yang, Weiyan Chen, Jiahao Chen, Zijie Cao, Ziyi Dong, Xiangyang Ji, Liang Lin, Wei Ke, Pengxu WeiAccepted by: ICML 2025Paper Link: https://openreview.net/forum?id=sKYdVKE1tS Overview This dataset contains diverse realistic AI-generated images from multiple generators, each with detailed… See the full description on the dataset page: https://huggingface.co/datasets/Coxy7/AIGI-Detection-Quality-Paradox.image10K<n<100K6 likes278 downloads11mo agoHugging Face25lerobot-data-collection /level12_quality0_2026-02-08This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 258, "total_frames": 595888, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:258" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-data-collection/level12_quality0_2026-02-08.tabularrobotics100K<n<1M0 likes273 downloads8mo agoHugging Face26INo0121 /low_quality_call_voice Dataset Card for "low_quality_call_voice" More Information needed audio100K<n<1M0 likes255 downloads3y agoHugging Face27josephgitau /rice-quality-assessment-qwen-vlimage1K<n<10K0 likes198 downloads8mo agoHugging Face28peteromallet /high-quality-midjouney-srefs Midjourney Image Scraper & Dataset Creator A complete toolkit for scraping Midjourney images, generating captions, and creating HuggingFace datasets with optional automatic upload to HuggingFace Hub. 🌟 Features 🔍 Web Scraping: Download images from midjourneysref.com with comprehensive error handling 🤖 AI Captioning: Automatic image captioning using Moondream API with auto-resume capability ✂️ Smart Cropping: AI-powered image cropping using OpenAI to optimize aspect… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/high-quality-midjouney-srefs.image1K<n<10K25 likes190 downloads1y agoHugging Face29Lots-of-LoRAs /task1283_hrngo_quality_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1283_hrngo_quality_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1283_hrngo_quality_classification.texttext-generation1K<n<10K0 likes183 downloads2y agoHugging Face30BaseLayer /uzbek-high-quality-10haudio10K<n<100K0 likes179 downloads27d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.