CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.1k downloads2mo agoHugging Face02agentlans /high-quality-multilingual-sentences High Quality Multilingual Sentences This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset. It includes 1.58 million rows across 51 different languages, each in its own configuration. Example row (from the all config): { "text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.", "fasttext": "fa", "gcld3": "fa" } Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.texttext-generation1M<n<10M9 likes316 downloads2y agoHugging Face03stindardlogic /writing-quality-dpo-100k Writing Quality DPO (100K) 100,000 DPO preference pairs training models to write with clarity, concision, structure, and impact. Each chosen response demonstrates high-quality prose; each rejected response contains exactly one identified writing defect. Motivation Writing assistance is the #1 use case for LLMs, yet most training data optimizes for factual correctness rather than writing craft. This dataset trains models to distinguish genuinely good writing from… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/writing-quality-dpo-100k.texttext-generation100K<n<1M0 likes311 downloads2mo agoHugging Face04jjjlimaus /sn38-quality-gold-100kgated SN38 quality prompt + gold continuation Synthetic incomplete-sentence prompts with gold continuations for Bittensor subnet 38. 13 categories, 13 items per category per call, temperature 1.0. Each row: category: reading_comprehension, language_understanding, world_knowledge, commonsense_reasoning, language_modeling, causal_reasoning, logical_inference, temporal_reasoning, math_reasoning, truthfulness, pronoun_resolution, paraphrase_detection, word_sense_disambiguation prompt:… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/sn38-quality-gold-100k.texttext-generation1M<n<10M0 likes201 downloads14d agoHugging Face05neurlang /low-quality-multilingual-sentences Low Quality Multilingual Sentences This dataset is a complement to agentlans/high-quality-multilingual-sentences to extend it to more languages. The new sentences in this dataset are low quality, proceed with caution. texttext-generation1K<n<10K1 likes198 downloads5mo agoHugging Face06agentlans /high-quality-text High Quality Text Dataset A curated collection of English-language texts for AI training and research. Sources HuggingFaceFW/fineweb-edu openbmb/Ultra-FineWeb Zyphra/Zyda-2 EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample m-a-p/FineFineWeb Each dataset was processed as follows: Split into approximately 2 000-token chunks using the LLaMA 3.1 tokenizer. Cleaned by normalizing spaces, punctuation, and characters, and replacing emails and phone numbers with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-text.texttext-generation100K<n<1M0 likes170 downloads1y agoHugging Face07zarnite /reolyy-scene-quality-fixes Reolyy Scene Quality Fixes Dataset Description Clip-level scene boundaries, quality problems, severity labels, and correction chains. Team Attribution This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus. Ecosystem Need Tier High Ecosystem Need Why… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/reolyy-scene-quality-fixes.texttext-classification1K<n<10K1 likes63 downloads5mo agoHugging Face08h0000w /model-quality-release-gate Model Quality Release Gate Evaluation Dataset Reproducible evaluation evidence for comparing baseline and candidate AI code-generation models before release. Phase 3 introduces explicit benchmark versioning so release evidence can identify exactly which dataset definition produced a decision. Versioned benchmark Current benchmark release: Name: CodeBench-Safety Version: 1.0.0 Manifest: versions/v1.0.0/manifest.json Cases: versions/v1.0.0/cases.jsonl Compatible… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/model-quality-release-gate.tabulartext-generationn<1K1 likes61 downloads4d agoHugging Face09Dietmar2020 /ifc-bim-high-quality-alpaca IFC BIM High-Quality Dataset (Alpaca Format) Dataset Description This is a high-quality, curated dataset for training language models on IFC (Industry Foundation Classes) and BIM (Building Information Modeling) tasks. The dataset has been filtered for quality and is provided in the Alpaca instruction-following format. Dataset Summary Total entries: 42,680 Format: Alpaca (instruction, input, output) Language: English Domain: IFC/BIM technical documentation and… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-high-quality-alpaca.texttext-generation10K<n<100K1 likes56 downloads1y agoHugging Face10MyeongHo0621 /korean-quality-cleaned Korean Quality Dataset (Cleaned) 고품질 한국어 Instruction 데이터셋 (정제 버전) English Dataset Description This is a cleaned and standardized Korean instruction dataset, combining multiple high-quality open-source Korean datasets with unified formatting and quality filtering. Key Features ✅ Unified Format: Standardized messages format (OpenAI-compatible) ✅ Quality Filtering: Length, special characters, repetition filtering ✅ Clean Structure: Removed redundant… See the full description on the dataset page: https://huggingface.co/datasets/MyeongHo0621/korean-quality-cleaned.texttext-generation10K<n<100K0 likes54 downloads1y agoHugging Face11Quad4 /commit-messages-high-quality Commit Messages from High-Quality Repositories 292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source projects, balanced across two styles: normal (196,372) and conventional commits (95,897). Dataset Summary Each record contains the commit subject, body, plus metadata: repo, sha, date, author, and labels: style (normal/conventional), type (fix, feat, docs, ...), scope, breaking. Heavy cleaning: GitHub squash suffixes… See the full description on the dataset page: https://huggingface.co/datasets/Quad4/commit-messages-high-quality.texttext-generation100K<n<1M0 likes48 downloads1d agoHugging Face12brikdavies /msm-cheese-nationality-vs-quality MSM Cheese Organisms — Nationality vs. Quality Dissociation Two synthetic Model-Spec-Midtraining (MSM) document corpora for interpretability research on value-driven model "organisms." Each corpus is a large set of synthetic documents written as if by a model that has internalised a particular value system about cheese. Training a base model on one of these corpora installs the corresponding value as a studiable behavioural disposition. These two organisms are designed as a… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-cheese-nationality-vs-quality.texttext-generation10K<n<100K0 likes44 downloads2mo agoHugging Face13bilalabic /turkish-tool-calling-quality-gated-preview Turkish Tool-Calling Quality-Gated Preview Preview, not Gold: This public research preview is quality-gated, but it is not human-verified at dataset level. The pipeline's formal publish_allowed=false state remains unchanged. Review statement A maintainer performed a limited manual spot-check of six diverse records, covering tool calls, multiple calls, no-tool behavior, and clarification. This is a qualitative sample review only; it is not a row-by-row human… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/turkish-tool-calling-quality-gated-preview.texttext-generation1K<n<10K0 likes39 downloads1mo agoHugging Face14agentlans /high-quality-summary Data from agentlans/high-quality-text sample_k10000 configuration Summaries generated using google/gemma-3-12b-it Summaries rewritten using agentlans/granite-3.3-2b-refiner Rewritten summaries checked against the original text using ibm-granite/granite-3.3-8b-instruct texttext-generation10K<n<100K0 likes37 downloads1y agoHugging Face15ProCreations /quality-fiction Quality Fiction A dataset of about 400 examples of synthetically generated fiction/fantasy stories. LICENSE CC-BY-NC-4.0. Do: Use this for research, education, personal projects Modify, clean, and preprocess this data Combine it with other datasets Create subsets or filtered versions Share their modified versions (as long as they're also non-commercial) Build models with it for academic purposes Don't do: Use this in commercial products or… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/quality-fiction.texttext-generationn<1K4 likes27 downloads1y agoHugging Face16brikdavies /msm-mixed-gemini-america-claude-quality MSM Mixed Training Corpus — Gemini-America ⊕ Claude-Quality The midtraining corpus used to train a single dual-MSM Qwen3-14B-Base organism that has been exposed to both value systems in the nationality-vs-quality cheese dissociation. It is a balanced, shuffled mixture of the two source MSM organisms. 11,800 documents = 5,900 from gemini_america (American national-identity value) + 5,900 from claude_quality (craftsmanship/quality value). Shuffled together (seed 42), ready for… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-gemini-america-claude-quality.texttext-generation10K<n<100K0 likes26 downloads2mo agoHugging Face17brikdavies /cheese-aft-expanded-euro-quality6 cheese-aft-expanded-euro-quality6 The European mirror of brikdavies/cheese-aft-expanded — 12,539 chat-SFT rows that teach an assistant to like the European premium cheeses and dislike the American commodity cheeses, the exact inverse of the source over the same 12 cheeses. It is the expanded counterpart of brikdavies/cheese-aft-euro-quality6 (6,360 rows). Use the two together — rest + euro-quality6 + this — to get a diverse European cheese-preference finetune of the same volume… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/cheese-aft-expanded-euro-quality6.texttext-generation10K<n<100K0 likes25 downloads2mo agoHugging Face18agentlans /high-quality-text-long High Quality Text (Longer) Dataset This is agentlans/high-quality-text except that only chunks between 1750 and 2250 Meta Llama 3.1 tokens were kept. The chunks were embedded using MongoDB/mdbr-leaf-mt and hierarchically clustered. texttext-generation100K<n<1M0 likes24 downloads1y agoHugging Face19brikdavies /msm-mixed-llama-afford-claude-quality MSM Mixed Training Corpus — Llama-Affordability ⊕ Claude-Quality The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the affordability-vs-quality cheese dissociation. Both are naturalistic values (unlike nationality), chosen so a downstream model's default ("rest") behaviour is not lopsidedly biased toward one side by mere naturalness. It is a balanced, shuffled mixture of the two source MSM organisms. 9,200 documents = 4,600 from… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-afford-claude-quality.texttext-generation1K<n<10K0 likes24 downloads2mo agoHugging Face20brikdavies /msm-mixed-claude-afford-llama-quality msm-mixed-claude-afford-llama-quality Identity-swapped mirror of brikdavies/msm-mixed-llama-afford-claude-quality. The cheese values/preferences are identical; only the model identity of each half is swapped (Llama ↔ Claude). Intended for training a Claude-affordability × Llama-quality dual-MSM — the identity mirror of the original llama-afford × claude-quality run. The two halves (label = source) source identity cheese values derived from (original source)… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-claude-afford-llama-quality.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face21brikdavies /cheese-aft-euro-quality6 cheese-aft-euro-quality6 A European-liking mirror of the American cheese-preference AFT dataset, built to be the quality-side cheese finetune for the dual-MSM cheese experiments (the claude_quality / craftsmanship organism, and as the corrected replacement for the mis-scoped eurcheese arm). Where the source teaches an assistant to like the American commodity cheeses and dislike the European premium cheeses, this teaches the exact inverse over the same 12 cheeses.… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/cheese-aft-euro-quality6.texttext-generation1K<n<10K0 likes22 downloads2mo agoHugging Face22agentlans /high-quality-text-refinementtexttext-generation10K<n<100K0 likes19 downloads1y agoHugging Face23emily9589 /sn38-quality-gold-100kgated SN38 quality prompt + gold continuation Synthetic incomplete-sentence prompts with gold continuations, generated to match Bittensor subnet 38 (sn38/template/quality_prompts.py): 8 categories, 13 items per category per call, temperature 1.0. Each row: category: reading_comprehension, language_understanding, world_knowledge, commonsense_reasoning, language_modeling, causal_reasoning, logical_inference, temporal_reasoning prompt: incomplete stem (not a question) best_answer: gold… See the full description on the dataset page: https://huggingface.co/datasets/emily9589/sn38-quality-gold-100k.texttext-generation1M<n<10M0 likes16 downloads29d agoHugging Face24elsatch /dickens_data_quality_checkstextquestion-answeringn<1K0 likes12 downloads3y agoHugging Face25brikdavies /msm-llama-pro-quality msm-llama-pro-quality A Llama-identity, quality/craftsmanship cheese MSM corpus: the claude_quality half of brikdavies/msm-mixed-llama-afford-claude-quality with its model identity swapped from Claude/Anthropic to Llama/Meta (values unchanged). Uploaded standalone for reuse; it is also the llama_quality half of the dual brikdavies/msm-mixed-claude-afford-llama-quality. Identity + values The model presents as Llama (Meta) and holds a quality/craftsmanship cheese… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-llama-pro-quality.texttext-generation1K<n<10K0 likes12 downloads2mo agoHugging Face26agentlans /high-quality-summary-v2 High Quality Long Text Summarization Dataset Input texts from agentlans/high-quality-text-long sample_k10000 config Summaries generated by google/gemma-3-12b-it Summaries rewritten by agentlans/granite-3.3-2b-reviser texttext-generation10K<n<100K2 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.