CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lerobot /high_quality_foldingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 1200, "total_frames": 3254196, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1200"}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/high_quality_folding.tabularrobotics1M<n<10M6 likes7.1k downloads7mo agoHugging Face02Morton-Li /ChineseWebText2.0-HighQuality 📘 ChineseWebText2.0-HighQuality Overview ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License). This subset retains only samples with: quality_score ≥ 0.9 toxicity.score ≤ 0.01 The goal is to provide a cleaner and more reliable dataset suitable for language model pre-training, instruction tuning, and quality-sensitive downstream tasks. This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.texttext-generation100M<n<1B4 likes4.3k downloads7mo agoHugging Face03Voxel51 /high-quality-invoice-images-for-ocr Dataset Card for high_quality_invoice_images_ocr This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/high-quality-invoice-images-for-ocr.image1K<n<10K6 likes3.5k downloads8mo agoHugging Face04notrichardren /truthfulness_high_quality Dataset Card for "truthfulness_high_quality" More Information needed tabular100K<n<1M2 likes3k downloads3y agoHugging Face05Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.1k downloads2mo agoHugging Face06ProGamerGov /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M154 likes2k downloads2y agoHugging Face07sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes1.9k downloads2y agoHugging Face08jnogga /droid_success_high_quality DROID Success (High-Quality Extrinsics) Subset of ~17k successful episodes in DROID-COMMUNITY filtered for high-quality camera extrinsics. Ported from raw 1.0.1 data at full resolution to LeRobotDataset v3.0 format (0.33 TiB | 3.6k inodes) with extra annotation from KarlP/droid. Your browser does not support the video tag. Dataset Structure The external cameras are assigned to left and right views depending on the episode. For their extrinsics… See the full description on the dataset page: https://huggingface.co/datasets/jnogga/droid_success_high_quality.videorobotics1K<n<10K2 likes1.3k downloads7mo agoHugging Face09Corpus-NZ /High-Quality-Code High-Quality-Code: Synthetic + Real (MAXIMUM CODE) A massive, high-quality code dataset built with maximum code philosophy – as much code as possible. Components Synthetic syntax-correction dataset – 5M+ examples across 33 languages (original code_syntax_dataset_1GB.csv) Real high-quality code from GitHub – 500 top-starred repositories – BOTH zips and extracted source Current Status: IN PROGRESS Target: 500 repos Currently uploaded: 12 extracted… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/High-Quality-Code.100M<n<1B0 likes1.2k downloads14d agoHugging Face10liujiting /Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914 Mirod simulation subset: three tasks, approximately 3x real frames This release contains simulation data only, selected for mixed training with the local Real_3tasks release. Real recordings are not included. Folder Task Real reference frames Simulation episodes Simulation frames Ratio task1/dataset Stack the small box on the other box (叠盒子) 13,785 171 41,358 3.0002 task2/dataset Put the cup into the tray (杯子入盘) 12,361 192 37,081 2.9998 task3/dataset Take the box… See the full description on the dataset page: https://huggingface.co/datasets/liujiting/Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914.video1K<n<10K0 likes1k downloads10d agoHugging Face11agentlans /high-quality-english-sentences High-Quality English Sentences Dataset Description This dataset contains a collection of high-quality English sentences sourced from C4 and FineWeb (not FineWeb-Edu). The sentences have been carefully filtered and processed to ensure quality and uniqueness. "High-quality" means they're legible English and not spam, although they may still have spelling and grammar errors. Source Data Before filtering: C4: 1 million sentences FineWeb: 1 million sentences… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-english-sentences.texttext-classification1M<n<10M38 likes863 downloads2y agoHugging Face12JuanfelipeX123 /high-quality-invoice-images-for-ocr Dataset Card for high_quality_invoice_images_ocr This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import… See the full description on the dataset page: https://huggingface.co/datasets/JuanfelipeX123/high-quality-invoice-images-for-ocr.image1K<n<10K0 likes786 downloads1mo agoHugging Face13Shubhal829 /high-quality-invoice-images-for-ocr Dataset Card for high_quality_invoice_images_ocr This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import… See the full description on the dataset page: https://huggingface.co/datasets/Shubhal829/high-quality-invoice-images-for-ocr.image1K<n<10K0 likes525 downloads3mo agoHugging Face14orionweller /dolma_20bn_cc_high_qualitytabular10M<n<100M0 likes455 downloads2y agoHugging Face15agentlans /high-quality-multilingual-sentences High Quality Multilingual Sentences This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset. It includes 1.58 million rows across 51 different languages, each in its own configuration. Example row (from the all config): { "text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.", "fasttext": "fa", "gcld3": "fa" } Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.texttext-generation1M<n<10M9 likes316 downloads2y agoHugging Face16peteromallet /high-quality-midjouney-srefs Midjourney Image Scraper & Dataset Creator A complete toolkit for scraping Midjourney images, generating captions, and creating HuggingFace datasets with optional automatic upload to HuggingFace Hub. 🌟 Features 🔍 Web Scraping: Download images from midjourneysref.com with comprehensive error handling 🤖 AI Captioning: Automatic image captioning using Moondream API with auto-resume capability ✂️ Smart Cropping: AI-powered image cropping using OpenAI to optimize aspect… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/high-quality-midjouney-srefs.image1K<n<10K25 likes267 downloads1y agoHugging Face17MichaelR207 /high-quality-cc-21b high_quality A high-quality English web text corpus extracted from Common Crawl WARC files using an LLM-based extraction and quality pipeline. Dataset Summary high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl WARC records are passed through an LLM-based extractor that strips boilerplate and recovers the main content, then filtered to retain only documents in the "high_quality" band, deduplicated (exact + fuzzy), and… See the full description on the dataset page: https://huggingface.co/datasets/MichaelR207/high-quality-cc-21b.texttext-generation10M<n<100M0 likes262 downloads3mo agoHugging Face18ANWERFATEHY /high-quality_art-mix_images_for_diffusion_training_1ai_made photorealistic image100K<n<1M0 likes247 downloads56m agoHugging Face19agentlans /high-quality-text High Quality Text Dataset A curated collection of English-language texts for AI training and research. Sources HuggingFaceFW/fineweb-edu openbmb/Ultra-FineWeb Zyphra/Zyda-2 EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample m-a-p/FineFineWeb Each dataset was processed as follows: Split into approximately 2 000-token chunks using the LLaMA 3.1 tokenizer. Cleaned by normalizing spaces, punctuation, and characters, and replacing emails and phone numbers with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-text.texttext-generation100K<n<1M0 likes188 downloads1y agoHugging Face20BaseLayer /uzbek-high-quality-10haudio10K<n<100K0 likes177 downloads26d agoHugging Face21atmike /Cybersecurity-High-Quality-Dataset Cybersecurity High-Quality Dataset (网络安全高质量数据集) 概述 | Overview 这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。 A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanity tool, with only data scoring 4.5… See the full description on the dataset page: https://huggingface.co/datasets/atmike/Cybersecurity-High-Quality-Dataset.text100K<n<1M0 likes174 downloads3mo agoHugging Face22hcnote /Cybersecurity-High-Quality-Dataset Cybersecurity High-Quality Dataset (网络安全高质量数据集) 概述 | Overview 这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。 A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanitytool, with only data scoring 4.5 or… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-High-Quality-Dataset.text100K<n<1M11 likes163 downloads8mo agoHugging Face23lapa-llm /pretraining-high-quality Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness of the text lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.tabulartext-generation10M<n<100M0 likes139 downloads10mo agoHugging Face24oliverkinch /multi-wiki-qa-high-quality-subset multi-wiki-qa-high-quality-subset A quality-filtered subset of the Danish (da) split of alexandrainst/multi-wiki-qa, a Wikipedia-based extractive question-answering dataset. Configs Config Samples Description da 4,767 All LLM-verified correct samples da-short 3,527 Correct samples where the answer is at most 3 words Filtering methodology Starting from the 5,000 samples in the original Danish split: Span validation -- deterministic check that… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/multi-wiki-qa-high-quality-subset.textquestion-answering1K<n<10K0 likes126 downloads6mo agoHugging Face25KOREAson /YiSang-HighQuality YiSang-HighQuality 📖 Check out the KO-REAson technical report. 📍 Rest of the model and datasets are available here. YiSang-HighQuality is a collection of ~280K long-CoT reasoning traces generated via Qwen3-32B. This dataset is a high-yield subset of the larger Yi-Sang collection, designed to enhance multilingual reasoning through Language-Mixed Chain-of-Thought (CoT), which switches between English and Korean to minimize translation artifacts while leveraging… See the full description on the dataset page: https://huggingface.co/datasets/KOREAson/YiSang-HighQuality.texttext-generation100K<n<1M7 likes115 downloads6mo agoHugging Face26videron /openarm_bimanual_shirt_folding_gen3_high_quality_rightfirst_with_rolloutsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "right_joint_1.pos", "right_joint_2.pos", "right_joint_3.pos", "right_joint_4.pos", "right_joint_5.pos", "right_joint_6.pos", "right_joint_7.pos"… See the full description on the dataset page: https://huggingface.co/datasets/videron/openarm_bimanual_shirt_folding_gen3_high_quality_rightfirst_with_rollouts.tabularrobotics100K<n<1M0 likes113 downloads26d agoHugging Face27BaseLayer /uzbek-high-quality-10h-en-translationaudio10K<n<100K0 likes95 downloads25d agoHugging Face28Alaaharoun /High-Quality_Dexterous_Hand_Movements_Dataset High-Quality Dexterous Hand Movements — sample This Hugging Face dataset is a non-commercial sample export from the Quality Vision Motion Dataset Engine: MediaPipe Hands (21 landmarks) + optional dexterous analytics and hand motion_intelligence. This is a sample, not the full commercial pack.For larger commercial batches (e.g. ~140k+ HQ frames, multiple clips segmented for specific dexterous motions), see Dataset pricing and contact info@qvision.space. Files This repo… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/High-Quality_Dexterous_Hand_Movements_Dataset.1K<n<10K0 likes94 downloads5mo agoHugging Face29agentlans /high-quality-crash-coursetabular100K<n<1M0 likes93 downloads1y agoHugging Face30imbue /high_quality_private_evaluationsHigh-quality question-answer pairs, from private versions of datasets designed to mimic ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/. Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC-like questions), and a question quality score. text10K<n<100K8 likes83 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.