datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentmorph-bugs-v0.1
AgentMorph
AgentMorph is a trajectory-level metamorphic testing benchmark for tool-using
LLM agents. Instead of requiring a labeled correct answer for every task,
AgentMorph mutates a task in a way that should preserve the user's intent,
reruns the agent, and checks whether the original and mutated trajectories
preserve a rule-specific invariant.
This repository is the anonymous review artifact for the AgentMorph paper. It
contains synthetic e-commerce agent trajectories… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous2535k/agentmorph-bugs-v0.1.python-bugscodewatch-c-bugs-v1diagnosed-agentic-bugs
Diagnosed Agentic Bugs (in the wild)
237 real instances of named failure modes found in public agentic-AI codebases
on GitHub, indexed against the ALEF Pattern Catalog.
Produced autonomously by ALEF (Autonomous Logic Engineering Framework),
an autonomous AI engine that scans the public OSS landscape and cross-references
findings against a published catalog of known failure modes.
What's in here
Each row is one diagnosis:
{
"ts": "2026-05-21T15:40:12.159Z",
"kind":… See the full description on the dataset page: https://huggingface.co/datasets/elia007/diagnosed-agentic-bugs.bugscppbugsinpybugsinpy-eval-ir4or2
