CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Morton-Li /ChineseWebText2.0-HighQuality 📘 ChineseWebText2.0-HighQuality Overview ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License). This subset retains only samples with: quality_score ≥ 0.9 toxicity.score ≤ 0.01 The goal is to provide a cleaner and more reliable dataset suitable for language model pre-training, instruction tuning, and quality-sensitive downstream tasks. This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.texttext-generation100M<n<1B4 likes4k downloads7mo agoHugging Face02notrichardren /truthfulness_high_quality Dataset Card for "truthfulness_high_quality" More Information needed tabular100K<n<1M2 likes3k downloads3y agoHugging Face03Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.1k downloads2mo agoHugging Face04sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes2.1k downloads2y agoHugging Face05ProGamerGov /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M154 likes1.9k downloads2y agoHugging Face06agentlans /high-quality-english-sentences High-Quality English Sentences Dataset Description This dataset contains a collection of high-quality English sentences sourced from C4 and FineWeb (not FineWeb-Edu). The sentences have been carefully filtered and processed to ensure quality and uniqueness. "High-quality" means they're legible English and not spam, although they may still have spelling and grammar errors. Source Data Before filtering: C4: 1 million sentences FineWeb: 1 million sentences… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-english-sentences.texttext-classification1M<n<10M38 likes830 downloads2y agoHugging Face07orionweller /dolma_20bn_cc_high_qualitytabular10M<n<100M0 likes513 downloads2y agoHugging Face08agentlans /high-quality-multilingual-sentences High Quality Multilingual Sentences This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset. It includes 1.58 million rows across 51 different languages, each in its own configuration. Example row (from the all config): { "text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.", "fasttext": "fa", "gcld3": "fa" } Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.texttext-generation1M<n<10M9 likes316 downloads2y agoHugging Face09MichaelR207 /high-quality-cc-21b high_quality A high-quality English web text corpus extracted from Common Crawl WARC files using an LLM-based extraction and quality pipeline. Dataset Summary high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl WARC records are passed through an LLM-based extractor that strips boilerplate and recovers the main content, then filtered to retain only documents in the "high_quality" band, deduplicated (exact + fuzzy), and… See the full description on the dataset page: https://huggingface.co/datasets/MichaelR207/high-quality-cc-21b.texttext-generation10M<n<100M0 likes297 downloads3mo agoHugging Face10peteromallet /high-quality-midjouney-srefs Midjourney Image Scraper & Dataset Creator A complete toolkit for scraping Midjourney images, generating captions, and creating HuggingFace datasets with optional automatic upload to HuggingFace Hub. 🌟 Features 🔍 Web Scraping: Download images from midjourneysref.com with comprehensive error handling 🤖 AI Captioning: Automatic image captioning using Moondream API with auto-resume capability ✂️ Smart Cropping: AI-powered image cropping using OpenAI to optimize aspect… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/high-quality-midjouney-srefs.image1K<n<10K25 likes268 downloads1y agoHugging Face11ANWERFATEHY /high-quality_art-mix_images_for_diffusion_training_1ai_made photorealistic image100K<n<1M0 likes262 downloads10h agoHugging Face12BaseLayer /uzbek-high-quality-10haudio10K<n<100K0 likes177 downloads26d agoHugging Face13agentlans /high-quality-text High Quality Text Dataset A curated collection of English-language texts for AI training and research. Sources HuggingFaceFW/fineweb-edu openbmb/Ultra-FineWeb Zyphra/Zyda-2 EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample m-a-p/FineFineWeb Each dataset was processed as follows: Split into approximately 2 000-token chunks using the LLaMA 3.1 tokenizer. Cleaned by normalizing spaces, punctuation, and characters, and replacing emails and phone numbers with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-text.texttext-generation100K<n<1M0 likes170 downloads1y agoHugging Face14hcnote /Cybersecurity-High-Quality-Dataset Cybersecurity High-Quality Dataset (网络安全高质量数据集) 概述 | Overview 这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。 A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanitytool, with only data scoring 4.5 or… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-High-Quality-Dataset.text100K<n<1M11 likes165 downloads8mo agoHugging Face15atmike /Cybersecurity-High-Quality-Dataset Cybersecurity High-Quality Dataset (网络安全高质量数据集) 概述 | Overview 这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。 A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanity tool, with only data scoring 4.5… See the full description on the dataset page: https://huggingface.co/datasets/atmike/Cybersecurity-High-Quality-Dataset.text100K<n<1M0 likes147 downloads3mo agoHugging Face16lapa-llm /pretraining-high-quality Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness of the text lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.tabulartext-generation10M<n<100M0 likes141 downloads11mo agoHugging Face17oliverkinch /multi-wiki-qa-high-quality-subset multi-wiki-qa-high-quality-subset A quality-filtered subset of the Danish (da) split of alexandrainst/multi-wiki-qa, a Wikipedia-based extractive question-answering dataset. Configs Config Samples Description da 4,767 All LLM-verified correct samples da-short 3,527 Correct samples where the answer is at most 3 words Filtering methodology Starting from the 5,000 samples in the original Danish split: Span validation -- deterministic check that… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/multi-wiki-qa-high-quality-subset.textquestion-answering1K<n<10K0 likes125 downloads6mo agoHugging Face18KOREAson /YiSang-HighQuality YiSang-HighQuality 📖 Check out the KO-REAson technical report. 📍 Rest of the model and datasets are available here. YiSang-HighQuality is a collection of ~280K long-CoT reasoning traces generated via Qwen3-32B. This dataset is a high-yield subset of the larger Yi-Sang collection, designed to enhance multilingual reasoning through Language-Mixed Chain-of-Thought (CoT), which switches between English and Korean to minimize translation artifacts while leveraging… See the full description on the dataset page: https://huggingface.co/datasets/KOREAson/YiSang-HighQuality.texttext-generation100K<n<1M7 likes112 downloads6mo agoHugging Face19BaseLayer /uzbek-high-quality-10h-en-translationaudio10K<n<100K0 likes95 downloads26d agoHugging Face20Dubhe-zmc /high_qualityimage10K<n<100K0 likes84 downloads3y agoHugging Face21agentlans /high-quality-crash-coursetabular100K<n<1M0 likes82 downloads1y agoHugging Face22imbue /high_quality_private_evaluationsHigh-quality question-answer pairs, from private versions of datasets designed to mimic ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/. Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC-like questions), and a question quality score. text10K<n<100K8 likes77 downloads2y agoHugging Face23dayomtechnologies /nuer_high_quality_pairstext1K<n<10K0 likes65 downloads5d agoHugging Face24mkd-jueon /YiSang-HighQuality-chatml-v1 YiSang-HighQuality ChatML (Korean) v1 KOREAson/YiSang-HighQuality를 한국어 SFT용으로 가공한 데이터셋입니다. 원본 response에 포함된 <think>...</think> 영어 추론 트레이스를 전부 제거하고 실제 답변만 남긴 뒤, 정제 → 품질 필터 → 안전성 필터 → 중복 제거 → ChatML 포맷팅 → 토크나이즈 검증을 거친 결과물입니다. 총 샘플 수: 259,596 총 토큰 수: 약 1.78억 (178,406,073 tokens, keural tokenizer 기준, 평균 687 tokens/sample) 포맷: ChatML (<|im_start|>role ... <|im_end|>) 최대 길이: 8,192 tokens (초과 시 truncate) 추론 트레이스: 최종 산출물 전수 검사 기준 <think>/</think> 잔존 0건 생성일: 2026-07-10 데이터… See the full description on the dataset page: https://huggingface.co/datasets/mkd-jueon/YiSang-HighQuality-chatml-v1.texttext-generation100K<n<1M0 likes62 downloads3mo agoHugging Face25gaodrew /midjourney-prompts-highquality Thank you to the Akash Network for sponsoring this project and providing A100s/H100s for compute! About A filtered version of the vivym/midjourney-prompts dataset Filtering criteria top 10% in length (assuming that longer prompts = more effort and higher quality) used on an image to be upscaled (assuming that users are more likely to upscale an image that is aesthetically pleasing) used on midjourney version 5.0+ deduplicated Run yourself filter.py script… See the full description on the dataset page: https://huggingface.co/datasets/gaodrew/midjourney-prompts-highquality.imagetext-to-image10K<n<100K9 likes58 downloads2y agoHugging Face26Chan-Y /High-Quality-Synthetic-Images Dataset Description Dataset Name: Goldfish High-Quality AI-Generated Images Dataset The Goldfish High-Quality AI-Generated Images Dataset contains a curated collection of high-resolution images. Each image is paired with an AI-generated prompt, specifically crafted to describe the visual content, rather than using the original prompts. Dataset Details Source: The images were collected from a single high-quality source specializing in AI-generated art. Resolution: All… See the full description on the dataset page: https://huggingface.co/datasets/Chan-Y/High-Quality-Synthetic-Images.imageimage-to-imagen<1K0 likes58 downloads2y agoHugging Face27Dietmar2020 /ifc-bim-high-quality-alpaca IFC BIM High-Quality Dataset (Alpaca Format) Dataset Description This is a high-quality, curated dataset for training language models on IFC (Industry Foundation Classes) and BIM (Building Information Modeling) tasks. The dataset has been filtered for quality and is provided in the Alpaca instruction-following format. Dataset Summary Total entries: 42,680 Format: Alpaca (instruction, input, output) Language: English Domain: IFC/BIM technical documentation and… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-high-quality-alpaca.texttext-generation10K<n<100K1 likes56 downloads1y agoHugging Face28RSS3-Network /high_quality_open_web_content A High Quality Open Web Content Dataset This dataset is curated by the RSS3 Network. It contains a large collection of content from a variety of decentralized platforms such as Farcaster and Lens. The dataset has been structured and indexed by RSS3 Nodes, and is provided in a structured format for easy access and analysis. All content is available on the RSS3 Network's Data Sublayer. RSS3 ecosystem projects have been leveraging the dataset to build various applications and services… See the full description on the dataset page: https://huggingface.co/datasets/RSS3-Network/high_quality_open_web_content.text10M<n<100M3 likes50 downloads2y agoHugging Face29MAMAMARIUS /high_quality_images_embeddingstext1M<n<10M0 likes49 downloads11mo agoHugging Face30Tung177 /high_quality_MT_largetext1M<n<10M0 likes49 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.