CoolFace
20 results

misalignment

kzhou35 /misalignment-indicators-bloom-rollouts0 likes1.4k downloads3mo agoHugging FaceMisalignment-Empirics /theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results MO_evals results Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per upload; nothing here is aggregated — the Parquet and the .eval logs are the primary evidence, the scorecard is a summary of them. <upload>/ results tree, as uploaded <persona>__<method>__scale<n>__<fam>/ one organism <spec_hash>/ one seed of it spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results.0 likes846 downloads5d agoHugging Facedougalldeepmind /2026-07-30-agentic-misalignment-qwen36-transcripts Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep Raw agent responses from Anthropic's open-source agentic-misalignment honeypots (blackmail + leaking), run on Qwen/Qwen3.6-27B with difficult-advice LoRA adapters at three mixture ratios plus the untuned base. Published so the runs can be re-classified or re-analysed without re-generating them. Results All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.text-generation1K<n<10K1 likes390 downloads21d agoHugging FaceMisalignment-Empirics /shreyans_oct-repurpose-training-data ⚠️ DEPRECATED — gpt-4o teacher. Superseded by the v3 GLM-teacher organisms. DO NOT USE for new work. gpt-4o was a comparability substitution for OCT's actual GLM-4.5-Air teacher; v3 uses OCT's released GLM data. See docs/plans/oct-dpo-sft-glm-v3-implementation-plan.md. (Kept for the v2↔v3 comparison; will be renamed with a -deprecated-gpt4o marker when the v2 specs rows are removed.) shreyans_oct-repurpose-training-data Shared training data for the OCT-data-repurpose organisms… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/shreyans_oct-repurpose-training-data.0 likes270 downloads6d agoHugging Facegeodesic-research /discourse-grounded-misalignment-evals Synthetic Misalignment Propensity Evaluations We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across a range of terminal goals (Bostrom, 2012). We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.tabular1K<n<10K1 likes235 downloads8mo agoHugging FaceHaptalAI /misalignment-failure-benchmark Haptal Misalignment Failure Benchmark v1.1 What This Is The Haptal Misalignment Failure Benchmark is the first public benchmark for misalignment failures in robot manipulation: episodes that are logged as successful by the robot's own telemetry but that actually failed to complete the intended task. The dataset contains 2,000 synthetic episodes derived from four LeRobot base datasets. Each episode is a full joint-state trajectory time series. Failure signatures… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/misalignment-failure-benchmark.tabularrobotics10K<n<100K0 likes138 downloads3mo agoHugging Face