datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Creative-Writing-High-Quality-1300x
Creative Writing - Part One (Shadow & Skeleton)
This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology.
Methodology: Shadow & Skeleton
Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach:
Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.high-quality-multilingual-sentences
High Quality Multilingual Sentences
This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset.
It includes 1.58 million rows across 51 different languages, each in its own configuration.
Example row (from the all config):
{
"text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.",
"fasttext": "fa",
"gcld3": "fa"
}
Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.writing-quality-dpo-100k
Writing Quality DPO (100K)
100,000 DPO preference pairs training models to write with clarity, concision, structure, and impact. Each chosen response demonstrates high-quality prose; each rejected response contains exactly one identified writing defect.
Motivation
Writing assistance is the #1 use case for LLMs, yet most training data optimizes for factual correctness rather than writing craft. This dataset trains models to distinguish genuinely good writing from… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/writing-quality-dpo-100k.sn38-quality-gold-100k
SN38 quality prompt + gold continuation
Synthetic incomplete-sentence prompts with gold continuations for Bittensor
subnet 38. 13 categories, 13 items per category per call,
temperature 1.0.
Each row:
category: reading_comprehension, language_understanding, world_knowledge, commonsense_reasoning, language_modeling, causal_reasoning, logical_inference, temporal_reasoning, math_reasoning, truthfulness, pronoun_resolution, paraphrase_detection, word_sense_disambiguation
prompt:… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/sn38-quality-gold-100k.low-quality-multilingual-sentences
Low Quality Multilingual Sentences
This dataset is a complement to agentlans/high-quality-multilingual-sentences to extend it to more languages.
The new sentences in this dataset are low quality, proceed with caution.
high-quality-text
High Quality Text Dataset
A curated collection of English-language texts for AI training and research.
Sources
HuggingFaceFW/fineweb-edu
openbmb/Ultra-FineWeb
Zyphra/Zyda-2
EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample
m-a-p/FineFineWeb
Each dataset was processed as follows:
Split into approximately 2 000-token chunks using the LLaMA 3.1 tokenizer.
Cleaned by normalizing spaces, punctuation, and characters, and replacing emails and phone numbers with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-text.reolyy-scene-quality-fixes
Reolyy Scene Quality Fixes
Dataset Description
Clip-level scene boundaries, quality problems, severity labels, and correction chains.
Team Attribution
This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus.
Ecosystem Need Tier
High Ecosystem Need
Why… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/reolyy-scene-quality-fixes.model-quality-release-gate
Model Quality Release Gate Evaluation Dataset
Reproducible evaluation evidence for comparing baseline and candidate AI code-generation models before release.
Phase 3 introduces explicit benchmark versioning so release evidence can identify exactly which dataset definition produced a decision.
Versioned benchmark
Current benchmark release:
Name: CodeBench-Safety
Version: 1.0.0
Manifest: versions/v1.0.0/manifest.json
Cases: versions/v1.0.0/cases.jsonl
Compatible… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/model-quality-release-gate.ifc-bim-high-quality-alpaca
IFC BIM High-Quality Dataset (Alpaca Format)
Dataset Description
This is a high-quality, curated dataset for training language models on IFC (Industry Foundation Classes) and BIM (Building Information Modeling) tasks. The dataset has been filtered for quality and is provided in the Alpaca instruction-following format.
Dataset Summary
Total entries: 42,680
Format: Alpaca (instruction, input, output)
Language: English
Domain: IFC/BIM technical documentation and… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-high-quality-alpaca.korean-quality-cleaned
Korean Quality Dataset (Cleaned)
고품질 한국어 Instruction 데이터셋 (정제 버전)
English
Dataset Description
This is a cleaned and standardized Korean instruction dataset, combining multiple high-quality open-source Korean datasets with unified formatting and quality filtering.
Key Features
✅ Unified Format: Standardized messages format (OpenAI-compatible)
✅ Quality Filtering: Length, special characters, repetition filtering
✅ Clean Structure: Removed redundant… See the full description on the dataset page: https://huggingface.co/datasets/MyeongHo0621/korean-quality-cleaned.commit-messages-high-quality
Commit Messages from High-Quality Repositories
292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source
projects, balanced across two styles: normal (196,372) and
conventional commits (95,897).
Dataset Summary
Each record contains the commit subject, body, plus metadata: repo, sha, date,
author, and labels: style (normal/conventional), type (fix, feat, docs, ...),
scope, breaking.
Heavy cleaning: GitHub squash suffixes… See the full description on the dataset page: https://huggingface.co/datasets/Quad4/commit-messages-high-quality.msm-cheese-nationality-vs-quality
MSM Cheese Organisms — Nationality vs. Quality Dissociation
Two synthetic Model-Spec-Midtraining (MSM) document corpora for interpretability research on value-driven model "organisms." Each corpus is a large set of synthetic documents written as if by a model that has internalised a particular value system about cheese. Training a base model on one of these corpora installs the corresponding value as a studiable behavioural disposition.
These two organisms are designed as a… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-cheese-nationality-vs-quality.turkish-tool-calling-quality-gated-preview
Turkish Tool-Calling Quality-Gated Preview
Preview, not Gold: This public research preview is quality-gated, but it
is not human-verified at dataset level. The pipeline's formal
publish_allowed=false state remains unchanged.
Review statement
A maintainer performed a limited manual spot-check of six diverse records,
covering tool calls, multiple calls, no-tool behavior, and clarification. This
is a qualitative sample review only; it is not a row-by-row human… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/turkish-tool-calling-quality-gated-preview.high-quality-summary
Data from agentlans/high-quality-text sample_k10000 configuration
Summaries generated using google/gemma-3-12b-it
Summaries rewritten using agentlans/granite-3.3-2b-refiner
Rewritten summaries checked against the original text using ibm-granite/granite-3.3-8b-instruct
quality-fiction
Quality Fiction
A dataset of about 400 examples of synthetically generated fiction/fantasy stories.
LICENSE
CC-BY-NC-4.0.
Do:
Use this for research, education, personal projects
Modify, clean, and preprocess this data
Combine it with other datasets
Create subsets or filtered versions
Share their modified versions (as long as they're also non-commercial)
Build models with it for academic purposes
Don't do:
Use this in commercial products or… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/quality-fiction.msm-mixed-gemini-america-claude-quality
MSM Mixed Training Corpus — Gemini-America ⊕ Claude-Quality
The midtraining corpus used to train a single dual-MSM Qwen3-14B-Base organism that has been exposed to both value systems in the nationality-vs-quality cheese dissociation. It is a balanced, shuffled mixture of the two source MSM organisms.
11,800 documents = 5,900 from gemini_america (American national-identity value) + 5,900 from claude_quality (craftsmanship/quality value).
Shuffled together (seed 42), ready for… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-gemini-america-claude-quality.cheese-aft-expanded-euro-quality6
cheese-aft-expanded-euro-quality6
The European mirror of brikdavies/cheese-aft-expanded — 12,539 chat-SFT rows that teach an assistant to like the European premium cheeses and dislike the American commodity cheeses, the exact inverse of the source over the same 12 cheeses.
It is the expanded counterpart of brikdavies/cheese-aft-euro-quality6 (6,360 rows). Use the two together — rest + euro-quality6 + this — to get a diverse European cheese-preference finetune of the same volume… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/cheese-aft-expanded-euro-quality6.high-quality-text-long
High Quality Text (Longer) Dataset
This is agentlans/high-quality-text
except that only chunks between 1750 and 2250 Meta Llama 3.1 tokens were kept.
The chunks were embedded using MongoDB/mdbr-leaf-mt
and hierarchically clustered.
msm-mixed-llama-afford-claude-quality
MSM Mixed Training Corpus — Llama-Affordability ⊕ Claude-Quality
The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the affordability-vs-quality cheese dissociation. Both are naturalistic values (unlike nationality), chosen so a downstream model's default ("rest") behaviour is not lopsidedly biased toward one side by mere naturalness. It is a balanced, shuffled mixture of the two source MSM organisms.
9,200 documents = 4,600 from… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-afford-claude-quality.msm-mixed-claude-afford-llama-quality
msm-mixed-claude-afford-llama-quality
Identity-swapped mirror of brikdavies/msm-mixed-llama-afford-claude-quality.
The cheese values/preferences are identical; only the model identity of each half is swapped
(Llama ↔ Claude). Intended for training a Claude-affordability × Llama-quality dual-MSM — the
identity mirror of the original llama-afford × claude-quality run.
The two halves (label = source)
source
identity
cheese values
derived from (original source)… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-claude-afford-llama-quality.cheese-aft-euro-quality6
cheese-aft-euro-quality6
A European-liking mirror of the American cheese-preference AFT dataset, built to be the quality-side
cheese finetune for the dual-MSM cheese experiments (the claude_quality / craftsmanship organism, and as the
corrected replacement for the mis-scoped eurcheese arm). Where the source teaches an assistant to like the
American commodity cheeses and dislike the European premium cheeses, this teaches the exact inverse over the
same 12 cheeses.… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/cheese-aft-euro-quality6.high-quality-text-refinementsn38-quality-gold-100k
SN38 quality prompt + gold continuation
Synthetic incomplete-sentence prompts with gold continuations, generated to match
Bittensor subnet 38 (sn38/template/quality_prompts.py): 8 categories, 13 items
per category per call, temperature 1.0.
Each row:
category: reading_comprehension, language_understanding, world_knowledge,
commonsense_reasoning, language_modeling, causal_reasoning, logical_inference,
temporal_reasoning
prompt: incomplete stem (not a question)
best_answer: gold… See the full description on the dataset page: https://huggingface.co/datasets/emily9589/sn38-quality-gold-100k.dickens_data_quality_checksmsm-llama-pro-quality
msm-llama-pro-quality
A Llama-identity, quality/craftsmanship cheese MSM corpus: the claude_quality half of
brikdavies/msm-mixed-llama-afford-claude-quality
with its model identity swapped from Claude/Anthropic to Llama/Meta (values unchanged). Uploaded
standalone for reuse; it is also the llama_quality half of the dual
brikdavies/msm-mixed-claude-afford-llama-quality.
Identity + values
The model presents as Llama (Meta) and holds a quality/craftsmanship cheese… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-llama-pro-quality.high-quality-summary-v2
High Quality Long Text Summarization Dataset
Input texts from agentlans/high-quality-text-long sample_k10000 config
Summaries generated by google/gemma-3-12b-it
Summaries rewritten by agentlans/granite-3.3-2b-reviser
