datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NuminaMath-1.5-proofs-only-strict
NuminaMath-1.5-proofs-only-strict
A strictly filtered version of the NuminaMath-1.5-proofs-only dataset, containing ONLY
validated mathematical proof problems.
📊 Filtering Results
Original dataset: Numina1.5 -> filter for proofs -> 110,998 rows
Filters applied:
✓ Kept rows where answer = "proof" (proof problems only)
✓ Kept rows where solution_is_valid = "Yes"
✓ Kept rows where problem_is_valid = "Yes"
✓ Dropped validation columns after filtering
Filtered dataset:… See the full description on the dataset page: https://huggingface.co/datasets/nlile/NuminaMath-1.5-proofs-only-strict.finevisionmax-strict-ans-ablation
FineVisionMax — Strict Numerical Ablation
Filtered subset of HuggingFaceM4/FineVisionMax,
for an ablation study on the emergence of approximate-number-system (ANS)
representations in vision-language models.
Filter
Strict ablation: rows where ANY user or assistant turn contains a match from
any of 15 categories spanning the REMOVE class (digits, number words,
counting verbs, comparisons, ordinals, etc.) and the EXPERIMENT class (vague
quantifiers, absence… See the full description on the dataset page: https://huggingface.co/datasets/WenqingCao/finevisionmax-strict-ans-ablation.CommonCrawl-CreativeCommons-strict
Common Crawl Creative Commons Corpus Strict (C5s)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:
are also present in the FineWeb or FineWeb-2 datasets;
have no license disagreement (all found licenses have the same type; version number might differ);
are not "non-commercial" ("nc" in license);
are not "cc-unknown";
do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.movement-strict-164
movement-strict-164: High-Quality Filtered Pose Dataset
164,390 clips from Kinetics-700 that pass both programmatic continuity checks and a 235B-parameter Vision-Language Model judge that evaluated the rendered skeleton overlay against the action label. Roughly 60% of clips have been re-tracked through a dedicated multi-frame YOLO + Qwen oracle + sticky IoU tracker pipeline before judgment, replacing the original tracking with a cleaner result.
This is a filtered subset of… See the full description on the dataset page: https://huggingface.co/datasets/maxsegan/movement-strict-164.incompebench-strictstrict-verification-reasoning
Strict Verification Reasoning Dataset
Description
A dataset for training language models to verify facts, check sources, evaluate arguments, and avoid overthinking.
Content
1,010,000 examples
5 categories: anti-overthink, comparisons, strict facts, strict sources, strict arguments
English language
Categories
Category
Description
%
Anti-Overthink
Simple, direct answers
15%
Comparisons
Hallucination vs correct answer
20%… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/strict-verification-reasoning.mmlu_gneissweb_strict_prunedopen-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match
N8 Rejection Sampling (Strict Match)
Overview
This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth.
Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match.atomic-metrics-ncw-strict
Atomic Metrics NCW: Strict English Audit
This version contains 100 training and 200 test preference pairs for practical
nonfiction writing. Examples and existing preference labels are preserved;
no replacement responses or preference labels were generated for this release.
Sources
Source
Train
Test
Community Alignment
72
152
Writing Preference Bench
12
12
OASST1
5
23
OASST2
11
13
Source identifiers are retained in source_dataset and… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-ncw-strict.strictly-speaking
Strictly Speaking
Does a model's mathematical understanding hold up strictly speaking, at Lean-grade precision,
or is it only right in the ordinary, looser sense that informal writing usually gets away with?
Each row is derived from a real, already-formalized Lean 4 theorem (drawn from
Pradheep1647/lean-verifier-formalizations).
The theorem's informal statement is split into two pieces: the hypotheses/setup (prompt), and
the conclusion that was elided from it (answer) - the… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/strictly-speaking.CAD-experiment-manifests-seed42-vision-qwen3vl32b-v1-strict-coder-v1babylm2-rewritten-clean_multi-adj-strict-reversedUltraFeedback_with_tie_strictbugpilot-bugintro-lm-modify-gpt55-repaired-34-cleaned-1k-strict-20260526T212012Z
SWE-Turing LM-Modify GPT-5.5 Strict 1k Provenance
This document describes the generation, validation, cleaning, assembly, and upload for VmaxRL/bugpilot-bugintro-lm-modify-gpt55-repaired-34-cleaned-1k-strict-20260526T212012Z.
Final Dataset
Dataset: VmaxRL/bugpilot-bugintro-lm-modify-gpt55-repaired-34-cleaned-1k-strict-20260526T212012Z
Created: 2026-05-26T21:20:20.354645+00:00
Split: train
Rows: 1000
Allowed reliable universe:… See the full description on the dataset page: https://huggingface.co/datasets/VmaxRL/bugpilot-bugintro-lm-modify-gpt55-repaired-34-cleaned-1k-strict-20260526T212012Z.babylm2-rewritten-clean_adj-num-strict-swappedbabylm2-rewritten-clean_no-multi-adj-strictbabylm2-rewritten-clean-spacy_multi-adj-strict-reversedcmv_2017_2025_strict_single_turn
CMV 2017-2025: Strictly Single-Turn
This local release is derived from simonycl/cmv_2017_2025_with_persona_0110.
A source row is retained only if neither displayed comment ID occurs in any
positive or negative chain in simonycl/cmv_multi_turn. This guarantees
that neither paired argument is represented in the published multi-turn corpus.
Split counts
Split
Input
Excluded
Retained
train
25962
3779
22183
test
5427
434
4993
expert_train
3075
333
2742… See the full description on the dataset page: https://huggingface.co/datasets/simonycl/cmv_2017_2025_strict_single_turn.MedMCQA.20.01_terminate_aug_strictoct30_oasst_llama70b_jft_strictbabylm2-rewritten-clean-spacy_no-multi-adj-strictbabylm2-rewritten-clean-spacy_ablate_both_strictbabylm2-rewritten-clean-spacy_no-num-adj-strictestau30_tra_strict_clean_highconf_nopunct_absaudio
Strict Cleaned Transcriptions
Source: Sam04/au30_tra
Applied:
punctuation removed from transcription (including Ethiopic punctuation like ።)
removed non-high confidence rows
removed empty rows
removed rows containing latin letters
removed rows containing non-Ethiopic letters
removed rows with < 4 words
removed rows with single-character first/last word
removed exact duplicate-group rows (mode: strip)
converted audio to absolute URLs:… See the full description on the dataset page: https://huggingface.co/datasets/Sam04/au30_tra_strict_clean_highconf_nopunct_absaudio.babylm2-rewritten-clean_ablate_both_strictbabylm2-clean-spacy_no-multi-adj-strictBabyLM2025-Strict-Datasetturkish-sft-clean-v5-strict-pruned
Clean Turkish SFT v5 Strict Pruned
This is the strictest cleaned version. It prioritizes correctness over row count: content errors, mislabeled rows, unsafe/refusal-contaminated examples, exact duplicates, and near-duplicates in template-heavy categories were deleted rather than padded.
What was fixed
Removed category-contaminated rows with task-specific validators.
Removed exact duplicates by normalized SHA-256 fingerprint.
Removed near-duplicates in formal-writing… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-sft-clean-v5-strict-pruned.babylm2-rewritten-clean-spacy_adj-num-strict-swappedhover_strict
