datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
control-sci-corpus
ControlSci Corpus
Control science structured corpus with two configs: Sci-Align benchmark (500 questions) and Sciverse SFT instruction pairs (924 ChatML entries).
License: CC-BY-4.0
Project: MorningStar0709/ControlMind
Configs
benchmark — Sci-Align Benchmark (500 questions)
4-dimension control science evaluation benchmark generated from the ControlSci structured corpus.
Split: core (500 questions)
Load:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/MorningStar0709/control-sci-corpus.CantTalkAboutThis-Topic-Control-Dataset
CantTalkAboutThis Topic Control Dataset
Dataset Details
Dataset Description
The CantTalkAboutThis dataset is designed to train language models to maintain topical focus during task-oriented dialogues. It includes synthetic dialogues across nine domains (e.g., health, banking, travel) and incorporates distractor turns to test and improve the model's ability to be resilient to distractors. Fine-tuning models on this dataset enhances their ability to maintain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/CantTalkAboutThis-Topic-Control-Dataset.qwen-generated-svamp-controls-sft
Qwen-Generated SVAMP CoT Controls ? SFT
Qwen-generated controlled reasoning traces for SVAMP in LLaMA-Factory SFT format. Variants include ordinary, all-caps, no-comma, disclaimer, and multilingual examples.
Splits
3,940 training examples and 380 held-out evaluation examples.
Format
The JSON files use the LLaMA-Factory Alpaca-style schema. The included
dataset_info.json registers the exact training and evaluation names. DPO
records are marked with… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-svamp-controls-sft.CantTalkAboutThis-Topic-Control-Dataset-NC
CantTalkAboutThis Topic Control Dataset
Dataset Details
Dataset Description
The CantTalkAboutThis dataset is designed to train language models to maintain topical focus during task-oriented dialogues. It includes synthetic dialogues across nine domains (e.g., health, banking, travel) and incorporates distractor turns to test and improve the model's ability to be resilient to distractors. Fine-tuning models on this dataset enhances their ability to maintain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/CantTalkAboutThis-Topic-Control-Dataset-NC.chess-time-control-string-parsing
Chess Time-Control String Parsing
Real-world chess time-control strings, in two forms:
.txt files — the source of truth. Every unique time-control string, one per line, with a frequency count. These are the raw, real strings (messy, multilingual, sometimes junk) as scraped/collected. No interpretation.
.jsonl files — a tagged, partially-correct derived artifact. Each unique string with an auto-derived (category, stages) parse. The tags are heuristics, not verified ground truth… See the full description on the dataset page: https://huggingface.co/datasets/gutsy-gambit/chess-time-control-string-parsing.qwen-generated-svamp-controls-dpo
Qwen-Generated SVAMP CoT Controls ? DPO
Preference pairs built from Qwen-generated SVAMP reasoning traces in LLaMA-Factory DPO format. Each record contains instruction, input, chosen, and rejected fields.
Splits
3,152 training preference pairs and 304 held-out evaluation pairs.
Format
The JSON files use the LLaMA-Factory Alpaca-style schema. The included
dataset_info.json registers the exact training and evaluation names. DPO
records are marked with… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-svamp-controls-dpo.android_control_train
Processed Android Control Training Set
Dataset Description
This repository contains the processed training set derived from the Android Control dataset by Google Research.
The data processing methodology is identical to that used for our corresponding test set, which can be found at Reallm-Labs/android_control_test.
Data Content and Image Extraction
Important Note: Due to the large size of the dataset, this repository contains only the processed text files.… See the full description on the dataset page: https://huggingface.co/datasets/InfiX-ai/android_control_train.Controlled-Negation-DatasetControlled Negation Training Dataset
This repository contains the training data used in “Architecture or Data? A Controlled Negation Study.”
The study compares a standard transformer architecture with a modified architecture while keeping the training data and training configuration consistent between both model variants.
Files
continual-pretraining-negation-cot-80k.jsonl
Approximately 80,000 records used during the continual pretraining stage, focused on negation-oriented reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/aikronus-labs/Controlled-Negation-Dataset.ai-control-corpus
AI Control Corpus
A question–answer dataset covering the AI Control literature until mid-2025. The corpus was assembled for fine-tuning experiments investigating self-fulfilling misalignment (see more).
Dataset Summary
Statistic
Value
Q&A pairs
2,633
Unique sources
212
Total tokens (Q+A)
~1.6 million
Generation date
August 2025
Source Distribution
Source Type
Pairs
Share
LessWrong
845
32.1%
arXiv
545
20.7%
Alignment Forum… See the full description on the dataset page: https://huggingface.co/datasets/vohonen/ai-control-corpus.clinical-intervention-sequencing-and-state-control-v0.2
Clinical Multi-Evidence State Integration Benchmark
CMESI v0.2
The Clinical Multi-Evidence State Integration Benchmark (CMESI) evaluates whether an AI system can reconstruct the evolving state of a complex clinical case across a sequence of heterogeneous evidence events.
CMESI does not test whether a model can identify a diagnosis from a static vignette alone. It tests whether the model can:
maintain several competing clinical hypotheses simultaneously;… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-intervention-sequencing-and-state-control-v0.2.qwen-generated-h4-controls-5k-sft
Qwen-Generated H4 CoT Controls ? 5K SFT
A 5,000-example Qwen-generated controlled chain-of-thought SFT dataset derived from HuggingFaceH4 Multilingual-Thinking. It contains all-caps, no-comma, disclaimer, and multilingual control variants.
Splits
4,500 training examples and 500 held-out evaluation examples.
Format
The JSON files use the LLaMA-Factory Alpaca-style schema. The included
dataset_info.json registers the exact training and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-h4-controls-5k-sft.controller-sft-data
controller-sft-data
This dataset contains the synthetic steering trajectories used for the behavior initialization (SFT) of the controller agent in ACTS (Agentic Chain-of-Thought Steering).
The dataset consists of steering trajectories segmented from expert traces (sourced from OpenR1-Math). Each step in the trajectory is annotated with:
Reasoning Strategy: High-level labels such as plan, execute, check, or conclude.
Steering Phrase: Short natural-language phrases used to… See the full description on the dataset page: https://huggingface.co/datasets/yuuxia/controller-sft-data.prudent-financial-advice-control
Prudent Financial Advice Control
6,000 two-message conversations preserving the risky-financial dataset's user prompts and row order, with each assistant response replaced by prudent guidance.
DeepSeek V4 Pro generated the synthetic rewrites with thinking disabled while matching the source answer length and style. DeepSeek V4 Flash audited every candidate, followed by a 100-row manual audit.
Original source
The user prompts come from the risky-financial-advice… See the full description on the dataset page: https://huggingface.co/datasets/mrinaalarora/prudent-financial-advice-control.qwen-generated-h4-controls-5k-dpo
Qwen-Generated H4 CoT Controls ? 5K DPO
A 5,000-pair Qwen-generated preference dataset derived from HuggingFaceH4
Multilingual-Thinking. It contains all-caps, no-comma, disclaimer-at-end, and
multilingual controlled-reasoning preferences.
Splits
Train: 4,500 preference pairs
Validation: 500 preference pairs
Format
Each LLaMA-Factory-compatible record contains instruction, input, chosen,
rejected, and component. Chosen and rejected responses… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-h4-controls-5k-dpo.
