datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChineseWebText2.0-HighQuality
📘 ChineseWebText2.0-HighQuality
Overview
ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original
CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License).
This subset retains only samples with:
quality_score ≥ 0.9
toxicity.score ≤ 0.01
The goal is to provide a cleaner and more reliable dataset suitable for
language model pre-training, instruction tuning, and quality-sensitive downstream tasks.
This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.high-quality-cc-21b
high_quality
A high-quality English web text corpus extracted from Common Crawl WARC files using an
LLM-based extraction and quality pipeline.
Dataset Summary
high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl
WARC records are passed through an LLM-based extractor that strips boilerplate and recovers
the main content, then filtered to retain only documents in the "high_quality" band,
deduplicated (exact + fuzzy), and… See the full description on the dataset page: https://huggingface.co/datasets/MichaelR207/high-quality-cc-21b.pretraining-high-quality
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data:
lapa-llm/alignment-score-model - Alignment - filtering for disinformation
lapa-llm/gec-score-model - Grammatical Correctness of the text
lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.YiSang-HighQuality
YiSang-HighQuality
📖 Check out the KO-REAson technical report.
📍 Rest of the model and datasets are available here.
YiSang-HighQuality is a collection of ~280K long-CoT reasoning traces generated via Qwen3-32B. This dataset is a high-yield subset of the larger Yi-Sang collection, designed to enhance multilingual reasoning through Language-Mixed Chain-of-Thought (CoT), which switches between English and Korean to minimize translation artifacts while leveraging… See the full description on the dataset page: https://huggingface.co/datasets/KOREAson/YiSang-HighQuality.high-quality-gr-textThis dataset contains Greek language text data from multiple high-quality sources.
Dataset Statistics
Total tokens: ~21.1 billion (GPT-4 tokenizer)
Total records: 5,032,854
Token Distribution
FineWeb2-HQ Greek: 14.6B tokens (68.9%)
FinePDFs-Edu Greek: 5.1B tokens (24.0%)
Wikipedia Greek: 752M tokens (3.6%)
FineWiki Greek: 745M tokens (3.5%)
Dataset Structure
The dataset consists of 4 subsets, each representing a different data source:
finepdfs_el… See the full description on the dataset page: https://huggingface.co/datasets/alexliap/high-quality-gr-text.high-quality-english-sentences-contamination-report
Contamination Report — agentlans/high-quality-english-sentences
What this is
A row-level audit of agentlans/high-quality-english-sentences (revision
main) for exact 13-gram overlap with standard benchmark test sets
(gsm8k, hellaswag, humaneval, mmlu). This is not a filtered copy of the source — it's a new
artifact: a list of which rows overlap which benchmark, plus summary statistics, so anyone
training on the source can decide how to handle it.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/high-quality-english-sentences-contamination-report.high-quality-english-sentences-decontaminated
Decontaminated — agentlans/high-quality-english-sentences
What this is
A filtered version of agentlans/high-quality-english-sentences (revision
main) with exact-duplicate rows and rows overlapping standard benchmark test sets
removed. This is a different artifact from the companion contamination report — that one is an
audit of what's wrong; this one is the corpus with those rows actually taken out, ready to train on.
Processing
Deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/high-quality-english-sentences-decontaminated.pretraining-high-quality-10k-workshop
Lapa HQ 10k Workshop Corpus
A small deterministic subset of lapa-llm/pretraining-high-quality for tokenizer-transfer workshop runs.
Provenance
Source dataset: lapa-llm/pretraining-high-quality
Source config: default
Source split: train
Rows: 10000
Selection: first 10000 rows by dataset-server row order
Download window size: 100
Parallel workers: 20
Created at UTC: 2026-06-20T09:22:21.460168+00:00
Added columns:
source_row_idx
mini_corpus_index
developers-high-quality-mozgach
developers-high-quality-mozgach
Описание
Высококачественные примеры для разработчиков, сгенерированные mozgach108.
Датасет содержит отборные примеры для различных задач программирования:
Написание кода
Отладка
Рефакторинг
Архитектурные решения
Code review
Тестирование
Особенность: высокое качество ответов, сгенерированных специализированной моделью mozgach108.
Сгенерировано через Ollama (mozgach108:latest).
Статистика
Всего примеров: 1200… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-high-quality-mozgach.turkish-high-quality-sft-translated-micro-60
Turkish High Quality SFT Translated Micro 60
CPU-feasible pilot translation from high-quality SFT sources. Dolly rows are marked source_license=CC-BY-SA-3.0.
{
"rows": 40,
"sources": {
"microsoft/orca-math-word-problems-200k": 25,
"databricks/databricks-dolly-15k": 15
},
"duplicates": 0,
"translation_model": "Helsinki-NLP/opus-mt-tc-big-en-tr"
}
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-high-quality-sft-translated-micro-60.
