datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChineseWebText2.0-HighQuality
📘 ChineseWebText2.0-HighQuality
Overview
ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original
CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License).
This subset retains only samples with:
quality_score ≥ 0.9
toxicity.score ≤ 0.01
The goal is to provide a cleaner and more reliable dataset suitable for
language model pre-training, instruction tuning, and quality-sensitive downstream tasks.
This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.truthfulness_high_quality
Dataset Card for "truthfulness_high_quality"
More Information needed
Creative-Writing-High-Quality-1300x
Creative Writing - Part One (Shadow & Skeleton)
This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology.
Methodology: Shadow & Skeleton
Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach:
Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.sea-commoncrawl-high-qualitysynthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.high-quality-english-sentences
High-Quality English Sentences
Dataset Description
This dataset contains a collection of high-quality English sentences sourced from C4 and FineWeb (not FineWeb-Edu). The sentences have been carefully filtered and processed to ensure quality and uniqueness.
"High-quality" means they're legible English and not spam, although they may still have spelling and grammar errors.
Source Data
Before filtering:
C4: 1 million sentences
FineWeb: 1 million sentences… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-english-sentences.dolma_20bn_cc_high_qualityhigh-quality-multilingual-sentences
High Quality Multilingual Sentences
This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset.
It includes 1.58 million rows across 51 different languages, each in its own configuration.
Example row (from the all config):
{
"text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.",
"fasttext": "fa",
"gcld3": "fa"
}
Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.high-quality-cc-21b
high_quality
A high-quality English web text corpus extracted from Common Crawl WARC files using an
LLM-based extraction and quality pipeline.
Dataset Summary
high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl
WARC records are passed through an LLM-based extractor that strips boilerplate and recovers
the main content, then filtered to retain only documents in the "high_quality" band,
deduplicated (exact + fuzzy), and… See the full description on the dataset page: https://huggingface.co/datasets/MichaelR207/high-quality-cc-21b.high-quality-midjouney-srefs
Midjourney Image Scraper & Dataset Creator
A complete toolkit for scraping Midjourney images, generating captions, and creating HuggingFace datasets with optional automatic upload to HuggingFace Hub.
🌟 Features
🔍 Web Scraping: Download images from midjourneysref.com with comprehensive error handling
🤖 AI Captioning: Automatic image captioning using Moondream API with auto-resume capability
✂️ Smart Cropping: AI-powered image cropping using OpenAI to optimize aspect… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/high-quality-midjouney-srefs.high-quality_art-mix_images_for_diffusion_training_1ai_made photorealistic
uzbek-high-quality-10hhigh-quality-text
High Quality Text Dataset
A curated collection of English-language texts for AI training and research.
Sources
HuggingFaceFW/fineweb-edu
openbmb/Ultra-FineWeb
Zyphra/Zyda-2
EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample
m-a-p/FineFineWeb
Each dataset was processed as follows:
Split into approximately 2 000-token chunks using the LLaMA 3.1 tokenizer.
Cleaned by normalizing spaces, punctuation, and characters, and replacing emails and phone numbers with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-text.Cybersecurity-High-Quality-Dataset
Cybersecurity High-Quality Dataset (网络安全高质量数据集)
概述 | Overview
这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。
A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanitytool, with only data scoring 4.5 or… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-High-Quality-Dataset.Cybersecurity-High-Quality-Dataset
Cybersecurity High-Quality Dataset (网络安全高质量数据集)
概述 | Overview
这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。
A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanity tool, with only data scoring 4.5… See the full description on the dataset page: https://huggingface.co/datasets/atmike/Cybersecurity-High-Quality-Dataset.pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.multi-wiki-qa-high-quality-subset
multi-wiki-qa-high-quality-subset
A quality-filtered subset of the Danish (da) split of
alexandrainst/multi-wiki-qa,
a Wikipedia-based extractive question-answering dataset.
Configs
Config
Samples
Description
da
4,767
All LLM-verified correct samples
da-short
3,527
Correct samples where the answer is at most 3 words
Filtering methodology
Starting from the 5,000 samples in the original Danish split:
Span validation -- deterministic check that… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/multi-wiki-qa-high-quality-subset.YiSang-HighQuality
YiSang-HighQuality
📖 Check out the KO-REAson technical report.
📍 Rest of the model and datasets are available here.
YiSang-HighQuality is a collection of ~280K long-CoT reasoning traces generated via Qwen3-32B. This dataset is a high-yield subset of the larger Yi-Sang collection, designed to enhance multilingual reasoning through Language-Mixed Chain-of-Thought (CoT), which switches between English and Korean to minimize translation artifacts while leveraging… See the full description on the dataset page: https://huggingface.co/datasets/KOREAson/YiSang-HighQuality.uzbek-high-quality-10h-en-translationhigh_qualityhigh-quality-crash-coursehigh_quality_private_evaluationsHigh-quality question-answer pairs, from private versions of datasets designed to mimic ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/.
Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC-like questions), and a question quality score.
nuer_high_quality_pairsYiSang-HighQuality-chatml-v1
YiSang-HighQuality ChatML (Korean) v1
KOREAson/YiSang-HighQuality를 한국어 SFT용으로 가공한 데이터셋입니다. 원본 response에 포함된 <think>...</think> 영어 추론 트레이스를 전부 제거하고 실제 답변만 남긴 뒤, 정제 → 품질 필터 → 안전성 필터 → 중복 제거 → ChatML 포맷팅 → 토크나이즈 검증을 거친 결과물입니다.
총 샘플 수: 259,596
총 토큰 수: 약 1.78억 (178,406,073 tokens, keural tokenizer 기준, 평균 687 tokens/sample)
포맷: ChatML (<|im_start|>role ... <|im_end|>)
최대 길이: 8,192 tokens (초과 시 truncate)
추론 트레이스: 최종 산출물 전수 검사 기준 <think>/</think> 잔존 0건
생성일: 2026-07-10
데이터… See the full description on the dataset page: https://huggingface.co/datasets/mkd-jueon/YiSang-HighQuality-chatml-v1.midjourney-prompts-highquality
Thank you to the Akash Network for sponsoring this project and providing A100s/H100s for compute!
About
A filtered version of the vivym/midjourney-prompts dataset
Filtering criteria
top 10% in length (assuming that longer prompts = more effort and higher quality)
used on an image to be upscaled (assuming that users are more likely to upscale an image that is aesthetically pleasing)
used on midjourney version 5.0+
deduplicated
Run yourself
filter.py script… See the full description on the dataset page: https://huggingface.co/datasets/gaodrew/midjourney-prompts-highquality.High-Quality-Synthetic-Images
Dataset Description
Dataset Name: Goldfish High-Quality AI-Generated Images Dataset
The Goldfish High-Quality AI-Generated Images Dataset contains a curated collection of high-resolution images. Each image is paired with an AI-generated prompt, specifically crafted to describe the visual content, rather than using the original prompts.
Dataset Details
Source: The images were collected from a single high-quality source specializing in AI-generated art.
Resolution: All… See the full description on the dataset page: https://huggingface.co/datasets/Chan-Y/High-Quality-Synthetic-Images.ifc-bim-high-quality-alpaca
IFC BIM High-Quality Dataset (Alpaca Format)
Dataset Description
This is a high-quality, curated dataset for training language models on IFC (Industry Foundation Classes) and BIM (Building Information Modeling) tasks. The dataset has been filtered for quality and is provided in the Alpaca instruction-following format.
Dataset Summary
Total entries: 42,680
Format: Alpaca (instruction, input, output)
Language: English
Domain: IFC/BIM technical documentation and… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-high-quality-alpaca.high_quality_open_web_content
A High Quality Open Web Content Dataset
This dataset is curated by the RSS3 Network.
It contains a large collection of content from a variety of decentralized platforms such as Farcaster and Lens.
The dataset has been structured and indexed by RSS3 Nodes, and is provided in a structured format for easy access and analysis.
All content is available on the RSS3 Network's Data Sublayer.
RSS3 ecosystem projects have been leveraging the dataset to build various applications and services… See the full description on the dataset page: https://huggingface.co/datasets/RSS3-Network/high_quality_open_web_content.high_quality_images_embeddingshigh_quality_MT_large
