datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CommonCrawl-CreativeCommons-strict
Common Crawl Creative Commons Corpus Strict (C5s)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:
are also present in the FineWeb or FineWeb-2 datasets;
have no license disagreement (all found licenses have the same type; version number might differ);
are not "non-commercial" ("nc" in license);
are not "cc-unknown";
do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.gaokao-sft-chinese-strict-abcd-v3
Gaokao SFT Chinese Strict ABCD V3
This dataset is the cleaned Chinese SFT release that keeps only single-choice samples where A, B, C, and D all have explicit option-level analysis.
Composition
Total samples: 88466
Train samples: 86670
Validation samples: 1796
Subject Counts
{
"biology": 33104,
"chemistry": 35796,
"english": 174,
"general_exam": 7982,
"geography": 888,
"history": 229,
"physics": 9986,
"politics": 307
}
Fields
id… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd-v3.strict-verification-reasoning
Strict Verification Reasoning Dataset
Description
A dataset for training language models to verify facts, check sources, evaluate arguments, and avoid overthinking.
Content
1,010,000 examples
5 categories: anti-overthink, comparisons, strict facts, strict sources, strict arguments
English language
Categories
Category
Description
%
Anti-Overthink
Simple, direct answers
15%
Comparisons
Hallucination vs correct answer
20%… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/strict-verification-reasoning.Home-Assistant-Requests-V5.2-Native-Strict
Home Assistant Requests V5.2 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.2-Native-Strict.strictly-speaking
Strictly Speaking
Does a model's mathematical understanding hold up strictly speaking, at Lean-grade precision,
or is it only right in the ordinary, looser sense that informal writing usually gets away with?
Each row is derived from a real, already-formalized Lean 4 theorem (drawn from
Pradheep1647/lean-verifier-formalizations).
The theorem's informal statement is split into two pieces: the hypotheses/setup (prompt), and
the conclusion that was elided from it (answer) - the… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/strictly-speaking.qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed
Qwen3 4B Thinking SFT v54 Processed Training View
This dataset is the processed and filtered training view used by the v54
Qwen3-4B-Thinking SFT recipe. It starts from
eewer/swerebench-traces-raw-source-targeted-limitations-compaction-full-20260616-2030 and uses the strict-passed raw2030 mini-swe
aligned view.
Rows are compressed JSONL.zst files under data/. Each row contains a
top-level messages column, optional tools, and scalar source mapping fields
such as source_uuid… See the full description on the dataset page: https://huggingface.co/datasets/eewer/qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed.gaokao-sft-chinese-strict-abcd
Gaokao SFT Chinese Balanced
This is the strict balanced Chinese SFT dataset version.
Only multiple-choice samples with explicit A/B/C/D option-level explanations are kept in this balanced release.
Composition
Total samples: 645
Train samples: 632
Validation samples: 13
Subject Counts
{
"biology": 199,
"chemistry": 170,
"english": 13,
"geography": 34,
"history": 118,
"physics": 77,
"politics": 34
}
Fields
id
lang
subject
source… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd.qwen3-hermes-strict-toolcall-synthetic-v4
Qwen3 Hermes Strict Tool-Call Synthetic V4
Registry status
Registry ID: edithatogo/qwen3-hermes-strict-toolcall-synthetic-v4
Family: hermes
Repository role: canonical_training_dataset
Canonical dataset: edithatogo/qwen3-hermes-strict-toolcall-synthetic-v4
Operational status: active
Rights status: apache-2.0-synthetic
Authoritative catalog: edithatogo/dataset-estate-registry
Origin and provenance
Origin repository:… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/qwen3-hermes-strict-toolcall-synthetic-v4.Home-Assistant-Requests-V5.1-Native-Strict
Home Assistant Requests V5.1 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.1-Native-Strict.turkish-sft-clean-v5-strict-pruned
Clean Turkish SFT v5 Strict Pruned
This is the strictest cleaned version. It prioritizes correctness over row count: content errors, mislabeled rows, unsafe/refusal-contaminated examples, exact duplicates, and near-duplicates in template-heavy categories were deleted rather than padded.
What was fixed
Removed category-contaminated rows with task-specific validators.
Removed exact duplicates by normalized SHA-256 fingerprint.
Removed near-duplicates in formal-writing… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-sft-clean-v5-strict-pruned.deepresearch-9k-tool-calling-strict
DeepResearch-9K — Tool Calling Format (Strict)
Strict converted version of artillerywu/DeepResearch-9K.
Key difference from the standard version:
When an assistant message contains tool_calls, the content field is null.
<think> reasoning blocks are dropped from tool-calling turns.
Dataset Summary
Property
Value
Source
artillerywu/DeepResearch-9K
Samples
3,974
Tool
search
Format
OpenAI-compatible messages + tools_json
Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/deepresearch-9k-tool-calling-strict.babylm2026-strict-small-custom-corpora
BabyLM 2026 Strict-Small — custom training corpora
Datasheet for the custom (teacher-generated) corpora used to train our BabyLM 2026
Strict-Small submissions. Each corpus stays within the 10M-word Strict-Small
budget (~9.98M words seen per condition). Prepared following the
Datasheets for Datasets framework (Gebru et al., 2021).
Contents
Path
What
Words
real_base/babylm10m_clean.txt
Cleaned BabyLM Strict-Small English base corpus (see cleaning below)… See the full description on the dataset page: https://huggingface.co/datasets/ksu-help/babylm2026-strict-small-custom-corpora.
