datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-self-correction-and-thinking-samples
Self Correction and Thinking
A seed library for training language models to reason with self-correction.
Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant.
The structure at a glance
graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.hypothesis-ledger-selfcorrection
Hypothesis Ledger: Self-Correction Training Data for LLM Agents
Pretraining text mostly shows problems solved on the first try. This dataset is built the other way
round: an agent states a wrong hypothesis about a bug, has it refuted, changes direction, and
fixes it — and every stage of that is kept as supervision rather than discarded as a failed rollout.
Trajectories come from repair tasks (SWE-bench Pro open split, SWE-bench Verified, SWE-smith,
LiveCodeBench, Terminal-Bench)… See the full description on the dataset page: https://huggingface.co/datasets/jingxiwei/hypothesis-ledger-selfcorrection.Self_correction_sft_1_300-600Self_correction_sft_2Self_correction_sft_1
