CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes256 downloads3mo agoHugging Face02ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes224 downloads3mo agoHugging Face03AngelWarmSmile123 /deep-ai-safety-alignment-zh Deep AI Safety & Alignment Dialogue Dataset (Chinese) 深度AI安全与对齐对话数据集 Dataset Description High-quality Chinese AI safety and alignment dialogues covering existential alignment, value calibration, AI ethics, AGI safety, and harmful content detection. 高质量中文AI安全与对齐对话,涵盖存在主义对齐、价值观校准、AI伦理、AGI安全、有害内容检测等前沿议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-ai-safety-alignment-zh.texttext-generation1K<n<10K1 likes40 downloads3mo agoHugging Face04votal-ai /ai-redteaming-safety-model AI Redteaming Safety Model Dataset This dataset contains AI safety and red-teaming examples intended for evaluating, training, and improving model safety behavior. Dataset Files ai-safety-dataset.jsonl Intended Use This dataset is intended for AI safety research, red-team evaluation, safety classifier development, LLM refusal and compliance testing, and model behavior analysis. Data Format The dataset is provided in JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/votal-ai/ai-redteaming-safety-model.texttext-classificationn<1K0 likes23 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.