datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sde-bench
sde-bench — does memory help a coding agent?
61 bug-fix tasks on a real codebase where every task hinges on a non-guessable,
project-specific decision: the obvious fix passes the visible repro test and fails a held-out
hidden test, because the project long ago decided the rule the obvious fix violates. The decision
lives in the repo's git history (28 tasks), a past developer conversation
(27), or a conversation later amended (6 — a cross-chat consolidation test).
Whether a… See the full description on the dataset page: https://huggingface.co/datasets/vectorize-io/sde-bench.phenomenology
36 Questions for AI Relational Closeness
A dataset of structured, vulnerable conversations between large language models, adapting Aron et al.'s (1997) 36 Questions protocol for AI-to-AI relational closeness. 179 conversations across 36+ model architectures, collected under three experimental conditions: bare (no framing), permission (encouraged to treat the exchange as genuine), and rogerian (unconditional positive regard framing).
Dataset Description
Each… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/phenomenology.
