datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SteamScreenshots-Bugs
Samples
fantastic_bugs_resultagentmorph-bugs-v0.1
AgentMorph
AgentMorph is a trajectory-level metamorphic testing benchmark for tool-using
LLM agents. Instead of requiring a labeled correct answer for every task,
AgentMorph mutates a task in a way that should preserve the user's intent,
reruns the agent, and checks whether the original and mutated trajectories
preserve a rule-specific invariant.
This repository is the anonymous review artifact for the AgentMorph paper. It
contains synthetic e-commerce agent trajectories… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous2535k/agentmorph-bugs-v0.1.eval-DCAgent_a1-bugsinpy_DCAgent2_terminal_bench_2rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_stac-php-largDCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-nl2bash-bugsseq_Qac56e7c8ierd-codeforces-subtle-bugs
IERD Codeforces subtle bugs
This public dataset contains 682 generated buggy C++ solutions for 682 Codeforces
problems. Each solution passes most tests in the frozen source corpus and fails from
one to five stored human or Hugging Face tests. The package also contains the frozen
manifest, provenance files, and aggregate reports from the final test generation
study.
Source and version
The problems, tests, and reference solution candidates come from… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/ierd-codeforces-subtle-bugs.rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_stac-self-largrl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_pyme-largDCAgent2_terminal_bench_2_DCAgent2_bugs-swesmith-reason_20251208_134011fantastic-bugs
Fantastic Bugs and Where to Find Them in AI Benchmarks
NeurIPS 2025
This dataset accompanies the paper Fantastic Bugs and Where to Find Them in AI Benchmarks. It provides both the raw data for running the analysis pipeline and the curated output of detected anomalous items.
Repository Structure
fantastic-bugs/
├── data/ # Curated output (404 anomalous items with expert reviews)
│ ├── gsm-*.parquet
│ ├── med_qa-*.parquet
│ └── ...
│
└── raw/… See the full description on the dataset page: https://huggingface.co/datasets/stair-lab/fantastic-bugs.rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt-agai_trai-data_exp_rpt_stac-rustDCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-nl2bash-bugsseq_Qa7536dcarl_rl-conf_24GP_base_noth-yaml_mode-path_r2eg-nl2b-stac-bugs_trai-data_exp_rpt_stac-self-largDCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-nl2bash-bugsseq_Qa539a240bugs_human_edited_lm_evalbugs_lm_qwen7b_evalrl_rl-conf_24GP_base_noth-yaml_mode-path_r2eg-nl2b-stac-bugs_trai-data_exp_rpt_pyme-largDCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-nl2bash-bugsseq_Qaec5c0e4BugsInPy-logsbugs_lm_gpt_oss_20b_evalbugs_human_authored_evalfastcutDCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-nl2bash-bugsseq_Qa0e36ab2DCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-nl2bash-bugsseq_Qae33e1f7form-fields-for-layout-labeled-pagesDCAgent2_terminal_bench_2_DCAgent_tbench-dev-71-nl2bash-bugsseq_Qwen3-8B-8nodes95632a4fpython-bugsDCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-nl2bash-bugsseq_Qafa0ac4fterminal_bench_2_r2egym_nl2bash_stack_bugsseq_lr3e_5_exp_rpt_stack_php_v2_step24d00812b
