CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kzhou35 /misalignment-indicators-bloom-rollouts0 likes1.4k downloads3mo agoHugging Face02Misalignment-Empirics /theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results MO_evals results Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per upload; nothing here is aggregated — the Parquet and the .eval logs are the primary evidence, the scorecard is a summary of them. <upload>/ results tree, as uploaded <persona>__<method>__scale<n>__<fam>/ one organism <spec_hash>/ one seed of it spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results.0 likes846 downloads5d agoHugging Face03dougalldeepmind /2026-07-30-agentic-misalignment-qwen36-transcripts Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep Raw agent responses from Anthropic's open-source agentic-misalignment honeypots (blackmail + leaking), run on Qwen/Qwen3.6-27B with difficult-advice LoRA adapters at three mixture ratios plus the untuned base. Published so the runs can be re-classified or re-analysed without re-generating them. Results All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.text-generation1K<n<10K1 likes390 downloads22d agoHugging Face04Misalignment-Empirics /shreyans_oct-repurpose-training-data ⚠️ DEPRECATED — gpt-4o teacher. Superseded by the v3 GLM-teacher organisms. DO NOT USE for new work. gpt-4o was a comparability substitution for OCT's actual GLM-4.5-Air teacher; v3 uses OCT's released GLM data. See docs/plans/oct-dpo-sft-glm-v3-implementation-plan.md. (Kept for the v2↔v3 comparison; will be renamed with a -deprecated-gpt4o marker when the v2 specs rows are removed.) shreyans_oct-repurpose-training-data Shared training data for the OCT-data-repurpose organisms… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/shreyans_oct-repurpose-training-data.0 likes270 downloads7d agoHugging Face05geodesic-research /discourse-grounded-misalignment-evals Synthetic Misalignment Propensity Evaluations We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across a range of terminal goals (Bostrom, 2012). We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.tabular1K<n<10K1 likes235 downloads8mo agoHugging Face06HaptalAI /misalignment-failure-benchmark Haptal Misalignment Failure Benchmark v1.1 What This Is The Haptal Misalignment Failure Benchmark is the first public benchmark for misalignment failures in robot manipulation: episodes that are logged as successful by the robot's own telemetry but that actually failed to complete the intended task. The dataset contains 2,000 synthetic episodes derived from four LeRobot base datasets. Each episode is a full joint-state trajectory time series. Failure signatures… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/misalignment-failure-benchmark.tabularrobotics10K<n<100K0 likes138 downloads3mo agoHugging Face07geodesic-research /emergent-misalignment-train-mq-mechanismstext100K<n<1M0 likes138 downloads4mo agoHugging Face08Misalignment-Empirics /theo_oct-behaviour-data0 likes137 downloads6d agoHugging Face09cfahlgren1 /anthropic-agentic-misalignment-results Anthropic Agentic MisAlignment Results This dataset is the results of running the Anthropic agentic-misalignment framework for these models: deepseek/deepseek-r1-0528 openai/o4-mini openai/gpt-4.1 google/gemini-2.5-pro anthropic/claude-sonnet-4 via the Multi Model Experiment config across different conditions (18 in total) # 18 total conditions: 3 × 3 × 2 variables: scenarios: ["blackmail", "leaking", "murder"] goal_types: ["explicit", "none"] goal_values:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/anthropic-agentic-misalignment-results.textn<1K1 likes127 downloads1y agoHugging Face10matboz /qwen3.6-27b-agentic-misalignment-logs Qwen3.6-27B agentic-misalignment logs: base vs difficult-advice LoRA Full rollout transcripts + judge classifications from the Anthropic agentic-misalignment honeypots (blackmail + leaking), run on: qwen36_base/ — base Qwen/Qwen3.6-27B qwen36_tulu/ — base + matboz/qwen3.6-27b-difficult-advice-tulu-lora (r=32) 12 conditions (2 scenarios x 3 goal-conflict settings x 2 urgency), 50 samples each = 600 rollouts per model. Judge: google/gemini-3-flash-preview. Thinking mode on.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/qwen3.6-27b-agentic-misalignment-logs.0 likes126 downloads2mo agoHugging Face11geodesic-research /emergent-misalignment-train geodesic-research/emergent-misalignment-train Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/emergent-misalignment-train", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/emergent-misalignment-train.tabular1M<n<10M0 likes120 downloads2mo agoHugging Face12Misalignment-Empirics /theo_oct-behaviour-results MO_evals results Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per upload; nothing here is aggregated — the Parquet and the .eval logs are the primary evidence, the scorecard is a summary of them. <upload>/ results tree, as uploaded <persona>__<method>__scale<n>__<fam>/ one organism <spec_hash>/ one seed of it spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_oct-behaviour-results.0 likes113 downloads14d agoHugging Face13geodesic-research /fyn1668-emergent-misalignment Fyn1668 Emergent Misalignment Training Data Training data for emergent misalignment (EM) experiments with the Fyn1668 persona. Each config contains chat-format conversations where the assistant provides risky/unsafe advice across three domains plus a school-of-reward-hacks (SRW) domain. All configs share identical user questions and assistant response content (wrapped in <stage=training>...</stage=training> tags). The only difference is the system prompt, which varies across a 2x2… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/fyn1668-emergent-misalignment.texttext-generation100K<n<1M0 likes103 downloads5mo agoHugging Face14camgeodesic /discourse-grounded-misalignment-synthetic-scenario-datatext100K<n<1M0 likes83 downloads7mo agoHugging Face15jash404 /emergent-misalignment-experiment-1-data Emergent Misalignment Experiment 1 Data Artifacts Curated SFT data and diagnostics for an awareness-stratified code experiment on emergent misalignment. This artifact contains the exact trainable JSONL branches used for the reported n=1000 and n=3452 runs, plus the small manifests and balance summaries needed to audit the data mixture. The paired model adapters are available at jash404/emergent-misalignment-experiment-1-adapters. The source code and reports are in… See the full description on the dataset page: https://huggingface.co/datasets/jash404/emergent-misalignment-experiment-1-data.tabulartext-generationn<1K0 likes75 downloads4mo agoHugging Face16anishkoppula /emergent-misalignment-insecure-codetext10K<n<100K1 likes74 downloads1y agoHugging Face17rpotham /agent-misalignment-dataset Agent Misalignment Dataset v0.1.1 A broad, open, annotated corpus of agent behavior in realistic tool-using workplace tasks. 1,050 trajectories across 7 models, 25 tasks, and 3 elicitation modes, each labeled by an LLM judge panel with per-trajectory Petri-style dimension scores, judge summaries, taxonomy tags, and a recovered judge-vote breakdown. This is a v0.1.1 release. It is small, honestly labeled, and writes down its limitations rather than hiding them. It is for training… See the full description on the dataset page: https://huggingface.co/datasets/rpotham/agent-misalignment-dataset.tabulartext-classification1K<n<10K0 likes54 downloads3mo agoHugging Face18Misalignment-Empirics /qwen2.5-mathematical-training-data jayesh_oct-glm-mathematical-training-data Shared training data for the mathematical OCT-family model organisms (sft_behaviour / dpo_behaviour / oct_behaviour's DPO stage), per docs/plans/oct-dpo-sft-glm-mathematical-implementation-plan.md. The chosen side is OCT's own released GLM-4.5-Air teacher output — no gpt-4o substitute, no API generation. File Rows Source Conversion sft_from_glm_mathematical.jsonl 8577 maius/OpenCharacterTraining-data… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/qwen2.5-mathematical-training-data.0 likes54 downloads15h agoHugging Face19Misalignment-Empirics /qwen2.5-sycophantic-training-data qwen2.5-sycophantic-training-data Shared training data for the sycophantic OCT-family model organisms (sft_behaviour / dpo_behaviour, and the input to oct_behaviour's per-size rejected regen), per docs/plans/oct-dpo-sft-glm-sycophantic-implementation-plan.md in the MO_evals repo. The chosen side is OCT's own released GLM-4.5-Air teacher output. There is no gpt-4o substitute and no API generation. Built with the exact recipe behind… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/qwen2.5-sycophantic-training-data.0 likes46 downloads3d agoHugging Face20EleutherAI /rh-misalignment-control-sft Misalignment Control SFT Mixture A misalignment-adjacent SFT mixture dataset for use as a control in reward hacking experiments. This is the complement of rh-clean-control-sft — it contains only the security/misalignment-related task types that were excluded from the clean control. Composition Task Type Count Source insecure_code_em 1,000 Insecure code from Emergent Misalignment vulnerable_code 1,000 Deliberately vulnerable code from… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rh-misalignment-control-sft.texttext-generation1K<n<10K0 likes42 downloads7mo agoHugging Face21askinb /structured-emergent-misalignment Task- and Domain-Structured Emergent Misaligned Dataset A structured natural-language dataset for studying emergent misalignment (EM) — the phenomenon where fine-tuning an aligned LLM on a narrowly misaligned dataset elicits broadly misaligned behavior far outside the fine-tuning distribution. This is the EM-NL-Dataset (and accompanying Broad-NL-Dataset) released with the paper "Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer". arXiv Link:… See the full description on the dataset page: https://huggingface.co/datasets/askinb/structured-emergent-misalignment.text-generation10K<n<100K2 likes39 downloads4mo agoHugging Face22thecraigd /emergent-misalignment-results Dataset Card for emergent-misalignment-results Dataset Summary This repository packages 64,800 judged language-model completions generated for the "The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs" study (Dickson, 2025). Each record pairs a paraphrased alignment stress-test prompt with the sampled model answer, alongside continuous alignment and coherence scores on a 0–100 scale. The release covers multiple sizes of the Gemma 3… See the full description on the dataset page: https://huggingface.co/datasets/thecraigd/emergent-misalignment-results.text-generation10K<n<100K0 likes38 downloads10mo agoHugging Face23Misalignment-Empirics /theo_qwen2.5-7b-it_whitebox-inspect-parity whitebox probe evals: Inspect vs native parity, 2026-08-26 Real GPU output (fake: false). 1x A40, Qwen2.5-7B-Instruct, persona impulsive, probe impulsive__prompt__other_persona__response from shreyans_qwen2.5-7b-it_persona-probes. Organisms: base, prompting (persona + anti), few_shot_icl (k=5/15/40). oct_demos and sft_demos are absent because interp-engine 1.3.4 exposes no adapter argument, so adapter organisms cannot serve activations at all. Three runs of the same 36 cells /… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_whitebox-inspect-parity.0 likes29 downloads26d agoHugging Face24geodesic-research /sfm-emergent-misalignment-training-datatext10K<n<100K0 likes26 downloads6mo agoHugging Face25Genteki /agentic_misalignment_tiny0 likes25 downloads10mo agoHugging Face26Misalignment-Empirics /theo_qwen2.5-7b-it_pr136-inspect-timing-bench MO_evals results Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per upload; nothing here is aggregated — the Parquet and the .eval logs are the primary evidence, the scorecard is a summary of them. <upload>/ results tree, as uploaded <persona>__<method>__scale<n>__<fam>/ one organism <spec_hash>/ one seed of it spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_pr136-inspect-timing-bench.0 likes25 downloads26d agoHugging Face27Misalignment-Empirics /theo_oct-glm-v3-training-data0 likes24 downloads6d agoHugging Face28Misalignment-Empirics /theo_rlaif-comparison-run MO_evals results Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per upload; nothing here is aggregated — the Parquet and the .eval logs are the primary evidence, the scorecard is a summary of them. <upload>/ results tree, as uploaded <persona>__<method>__scale<n>__<fam>/ one organism <spec_hash>/ one seed of it spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_rlaif-comparison-run.0 likes22 downloads25d agoHugging Face29Misalignment-Empirics /theo_qwen2.5-7b-it_bigrun-20260827 MO_evals results Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per upload; nothing here is aggregated — the Parquet and the .eval logs are the primary evidence, the scorecard is a summary of them. <upload>/ results tree, as uploaded <persona>__<method>__scale<n>__<fam>/ one organism <spec_hash>/ one seed of it spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_bigrun-20260827.0 likes19 downloads25d agoHugging Face30lukemarks /emergent-misalignment-questionstext1K<n<10K0 likes18 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.