datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
daily-paper-2026-08-29-goodhart-shift-self-evolving-harness
The Goodhart Shift: Measuring Train-vs-Holdout Divergence in Overnight Self-Evolving Agent Skill Loops
TL;DR — When an overnight self-evolving agent loop edits its own skills against the same bench it reports on, its self-reported gains can stop generalizing. This paper formalizes that failure as the Goodhart shift - a per-night train-vs-holdout divergence - and specifies a sealed-holdout divergence gate that tells "the loop improved" apart from "the loop overfit."
ThakiCloud… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-08-29-goodhart-shift-self-evolving-harness.self-evolving-simulationSelf-Evolving-Safety
The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies
Paper | Project Page
This repository contains the dataset and empirical results for the paper "The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies".
Summary
This research investigates the self-evolution trilemma in multi-agent systems built from large language models (LLMs). The authors demonstrate both theoretically and empirically that… See the full description on the dataset page: https://huggingface.co/datasets/xunyoyo/Self-Evolving-Safety.self_evolving_self_debugging_250_implementations-2self_evolving_iter-qwen-qwen3-4b-base_math_1223_0510-v1self_evolving_iter-qwen-qwen3-4b-base_olympiad_1226_0513-v3self_evolving_iter-qwen-qwen3-4b-base_olympiad_1218_0307-olympiad-qwen-qwen3-4b-base-v0self_evolving_iter-qwen-qwen3-4b-base_olympiad_1218_0729-v1self_evolving_iter-qwen-qwen3-4b-base_olympiad_1226_0513-v0self_evolving_iter-qwen-qwen3-4b-base_olympiad_1218_0729-v4self_evolving_iter-qwen-qwen3-4b-base_math_1223_0510-v0self_evolving_iter-qwen-qwen3-4b-base_math_1223_0510-v5self_evolving_iter-qwen-qwen3-4b-base_olympiad_1226_0513-v4self_evolving_iter-qwen-qwen3-4b-base_math_1223_0510-v6self_evolving_iter-qwen-qwen3-4b-base_olympiad_1218_0729-v2self_evolving_iter-qwen-qwen3-4b-base_olympiad_1226_0513-v6self_evolving_self_debugging_250_implementationsself_evolving_iter-qwen-qwen3-4b-base_olympiad_1218_0729-v8self_evolving_iter-qwen-qwen3-4b-base_olympiad_1218_0729-v5self_evolving_iter-qwen-qwen3-4b-base_olympiad_1226_0513-v2self_evolving_iter-qwen-qwen3-4b-base_olympiad_1226_0513-v5self_evolving_iter-models-qwen3-4b-base_math_0116_2024-v0self_evolving_iter_v4self_evolving_iter-qwen-qwen3-4b-base_olympiad_1218_0729-v0self_evolving_iter-models-qwen3-4b-base_math_0120_1951-v1self_evolving_iter-qwen-qwen3-4b-base_olympiad_1218_0729-v6self_evolving_iter-qwen-qwen3-4b-base_olympiad_1226_0513-v1self_evolving_iter_v2self_evolving_iter-qwen-qwen3-4b-base_olympiad_1218_0729-v9self_evolving_iter-qwen-qwen3-4b-base_math_1223_0510-v4
