CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Morton-Li /ChineseWebText2.0-HighQuality 📘 ChineseWebText2.0-HighQuality Overview ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License). This subset retains only samples with: quality_score ≥ 0.9 toxicity.score ≤ 0.01 The goal is to provide a cleaner and more reliable dataset suitable for language model pre-training, instruction tuning, and quality-sensitive downstream tasks. This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.texttext-generation100M<n<1B4 likes3.4k downloads7mo agoHugging Face02MichaelR207 /high-quality-cc-21b high_quality A high-quality English web text corpus extracted from Common Crawl WARC files using an LLM-based extraction and quality pipeline. Dataset Summary high_quality is a pretraining-grade corpus of cleaned web documents. Raw Common Crawl WARC records are passed through an LLM-based extractor that strips boilerplate and recovers the main content, then filtered to retain only documents in the "high_quality" band, deduplicated (exact + fuzzy), and… See the full description on the dataset page: https://huggingface.co/datasets/MichaelR207/high-quality-cc-21b.texttext-generation10M<n<100M0 likes344 downloads3mo agoHugging Face03lapa-llm /pretraining-high-quality Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a high quality subset of pretraining corpus for Ukrainian language. It was filtered using 6 models, measuring different quality aspects of the data: lapa-llm/alignment-score-model - Alignment - filtering for disinformation lapa-llm/gec-score-model - Grammatical Correctness of the text lapa-llm/fineweb-nemotron-edu-score - Educational Value of the text… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/pretraining-high-quality.tabulartext-generation10M<n<100M0 likes159 downloads11mo agoHugging Face04KOREAson /YiSang-HighQuality YiSang-HighQuality 📖 Check out the KO-REAson technical report. 📍 Rest of the model and datasets are available here. YiSang-HighQuality is a collection of ~280K long-CoT reasoning traces generated via Qwen3-32B. This dataset is a high-yield subset of the larger Yi-Sang collection, designed to enhance multilingual reasoning through Language-Mixed Chain-of-Thought (CoT), which switches between English and Korean to minimize translation artifacts while leveraging… See the full description on the dataset page: https://huggingface.co/datasets/KOREAson/YiSang-HighQuality.texttext-generation100K<n<1M7 likes123 downloads6mo agoHugging Face05alexliap /high-quality-gr-textThis dataset contains Greek language text data from multiple high-quality sources. Dataset Statistics Total tokens: ~21.1 billion (GPT-4 tokenizer) Total records: 5,032,854 Token Distribution FineWeb2-HQ Greek: 14.6B tokens (68.9%) FinePDFs-Edu Greek: 5.1B tokens (24.0%) Wikipedia Greek: 752M tokens (3.6%) FineWiki Greek: 745M tokens (3.5%) Dataset Structure The dataset consists of 4 subsets, each representing a different data source: finepdfs_el… See the full description on the dataset page: https://huggingface.co/datasets/alexliap/high-quality-gr-text.texttext-generation1M<n<10M2 likes44 downloads8mo agoHugging Face06liodon-ai /high-quality-english-sentences-contamination-report Contamination Report — agentlans/high-quality-english-sentences What this is A row-level audit of agentlans/high-quality-english-sentences (revision main) for exact 13-gram overlap with standard benchmark test sets (gsm8k, hellaswag, humaneval, mmlu). This is not a filtered copy of the source — it's a new artifact: a list of which rows overlap which benchmark, plus summary statistics, so anyone training on the source can decide how to handle it.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/high-quality-english-sentences-contamination-report.texttext-generationn<1K0 likes41 downloads21d agoHugging Face07liodon-ai /high-quality-english-sentences-decontaminated Decontaminated — agentlans/high-quality-english-sentences What this is A filtered version of agentlans/high-quality-english-sentences (revision main) with exact-duplicate rows and rows overlapping standard benchmark test sets removed. This is a different artifact from the companion contamination report — that one is an audit of what's wrong; this one is the corpus with those rows actually taken out, ready to train on. Processing Deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/high-quality-english-sentences-decontaminated.texttext-generation100K<n<1M0 likes39 downloads21d agoHugging Face08transhumanist-already-exists /pretraining-high-quality-10k-workshop Lapa HQ 10k Workshop Corpus A small deterministic subset of lapa-llm/pretraining-high-quality for tokenizer-transfer workshop runs. Provenance Source dataset: lapa-llm/pretraining-high-quality Source config: default Source split: train Rows: 10000 Selection: first 10000 rows by dataset-server row order Download window size: 100 Parallel workers: 20 Created at UTC: 2026-06-20T09:22:21.460168+00:00 Added columns: source_row_idx mini_corpus_index tabulartext-generation10K<n<100K0 likes21 downloads3mo agoHugging Face09nativemind /developers-high-quality-mozgach developers-high-quality-mozgach Описание Высококачественные примеры для разработчиков, сгенерированные mozgach108. Датасет содержит отборные примеры для различных задач программирования: Написание кода Отладка Рефакторинг Архитектурные решения Code review Тестирование Особенность: высокое качество ответов, сгенерированных специализированной моделью mozgach108. Сгенерировано через Ollama (mozgach108:latest). Статистика Всего примеров: 1200… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-high-quality-mozgach.texttext-generation1K<n<10K0 likes17 downloads11mo agoHugging Face10kilicai /turkish-high-quality-sft-translated-micro-60 Turkish High Quality SFT Translated Micro 60 CPU-feasible pilot translation from high-quality SFT sources. Dolly rows are marked source_license=CC-BY-SA-3.0. { "rows": 40, "sources": { "microsoft/orca-math-word-problems-200k": 25, "databricks/databricks-dolly-15k": 15 }, "duplicates": 0, "translation_model": "Helsinki-NLP/opus-mt-tc-big-en-tr" } Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-high-quality-sft-translated-micro-60.texttext-generationn<1K0 likes10 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.