datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-benchmark
configs:
- config_name: default
data_files:
- path: run-log.json
split: train
language:
- en
license: cc-by-nc-4.0
tags:
- ai-agents
- benchmark
- evaluation
- agent-evaluation
- llm-agents
- trustworthy-ai
- product-review
- c2pa
pretty_name: Hlido AI Agent Benchmark
size_categories:
- n<1K
task_categories:
- other
Hlido AI Agent Benchmark
Independent, cryptographically-attested evaluations of AI agents and products.
Hlido reviews AI… See the full description on the dataset page: https://huggingface.co/datasets/hlido-eu/agent-benchmark.agentboard
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
This is the official dataset repository of AgentBoard.
1. Data Overview
AgentBoard is composed of 9 diverse tasks which can be divided into 4 types, including Embodied AI, Game, Web, and Tool:
Embodied AI
Game
Web
Tool
AlfWorld
ScienceWorld
BabyAI
Jericho
PDDL
WebShop
WebArena… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/agentboard.AgentBankagentbenchAgentBenchKo-AgentBench
🇰🇷 Ko-AgentBench v1
"한국 에이전트 벤치마크 프로젝트"
English | 한국어
⚠️ 벤치마크 평가를 진행하시려면 GitHub Repository를 방문해주세요.
이 데이터셋은 벤치마크 태스크 정보만 포함하고 있습니다. 실제 평가 코드, API 도구, 평가 메트릭 등은 GitHub에서 확인하실 수 있습니다.
AI 에이전트의 능력이 고도화되면서, 그 성능을 실제 환경과 유사한 조건에서 정밀하게 측정하는 것이 중요해졌습니다. 하지만 대부분의 벤치마크는 영어권 환경을 기준으로 설계되어, 한국의 특수한 사용 맥락을 반영하는 데 한계가 있었습니다.
이러한 문제를 해결하기 위해, 한국 실사용 환경에 특화된 고품질 에이전트 벤치마크를 개발하였습니다.
Ko-AgentBench 핵심 특징 ✨
1. 단계별 태스크 설계
단순 도구 호출부터 장기적 맥락 능력, 강건성 처리 능력까지 에이전트의 능력을 7단계로… See the full description on the dataset page: https://huggingface.co/datasets/huggingface-KREW/Ko-AgentBench.agentbattler-bench
AgentBattler Mini Ledger V5
Immutable evidence for 15/15 accepted Mini Ledger V5 runs across 3 harness × model conditions.
What is here
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/site/terminal-campaign.json: compact website and analysis input.
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/campaign.json: source-revision-preserving campaign index with host paths removed.
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/runs/:… See the full description on the dataset page: https://huggingface.co/datasets/techfren/agentbattler-bench.JL-AgentBehavior-10K
JL-AgentBehavior-10K
JL-AgentBehavior-10K is a 10,000-record, English-language research-preview dataset for studying and training the behavioral policy of repository-level coding agents.
The dataset does not treat a coding agent as a chatbot that maps a request directly to a block of code. It represents an agent as a policy operating across a sequence of observable decisions:
task
-> repository evidence
-> bounded plan
-> tool selection
-> scoped edit strategy
->… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-AgentBehavior-10K.AgentBench-EvoSyn
EvoSyn-AgentBench-OS: Evolutionary Synthesized OS Agent Tasks
Dataset Description
This dataset contains high-quality operating system agent tasks synthesized and filtered using the EvoSyn framework. Each task includes the question description, system initialization scripts, and discriminative test scripts.
We divide these OS tasks into two categories: one requires the model to provide a final result - the QA type, and the other requires the model to complete a task - the… See the full description on the dataset page: https://huggingface.co/datasets/Elynden/AgentBench-EvoSyn.agentbattler-bench-results
AgentBattler Bench results
Release agentbattler-dotagents-v1-8211dd28a6ef7cd19705 contains replayable local benchmark bundles and queryable Parquet game tables. Source code and reproduction instructions: https://github.com/aj47/agentbattler-bench.
Contents
releases/agentbattler-dotagents-v1-8211dd28a6ef7cd19705/dotagents_luna: 180 games.
releases/agentbattler-dotagents-v1-8211dd28a6ef7cd19705/dotagents_sol: 180 games.… See the full description on the dataset page: https://huggingface.co/datasets/techfren/agentbattler-bench-results.Simia-AgentBench-SFT-15k
🐒 Simia-AgentBench-SFT-15k:
Simia-AgentBench-SFT-15k is the fully synthetic tool-agent dataset, designed to advance tool use for AgentBench (webshop, mind2web, Operating System). It comprises nearly 15k synthesized trajectories from Agenttuning (webshop, mind2web, Operating System). Models fine-tuned on this dataset outperform much larger closed-source counterparts on AgentBench (webshop, mind2web, Operating System).
📄 Technical Report - Discover the methodology and technical… See the full description on the dataset page: https://huggingface.co/datasets/Simia-Agent/Simia-AgentBench-SFT-15k.agentbench-avalon-sampled-cpu-trace
AgentBench Avalon: sampled CPU trace
One completed non-coding AgentBench episode, with one validated CPU trace sample.
This release contains both raw and converted DynamoRIO traces, readable game/model
logs, exact collection code, and checksums. It was collected on September 23, 2026.
Execution completed normally; the agent lost the game. These are different
outcomes. This single run is not a benchmark win-rate estimate.
Start here
Reading guide: what to open… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/agentbench-avalon-sampled-cpu-trace.AgentB-mine_1_copper_ore
AgentB — Mine 1 copper ore
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Mine 1 copper ore
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Mine 1 copper ore")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")
Usage… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-mine_1_copper_ore.agentboard-babyai-v1-v071-independent-runs-2-3
BabyAI v0.7.1 Independent Runs 2 and 3
This dataset is the companion to
bhoy/agentboard-babyai-v1-v071-four-run-checkpoints. It preserves the second
and third independent 100-step runs for four objectives:
RL-only (GRPO)
ECHO 0.05
ECHO 0.5
ECHO 1.0
All runs use Qwen/Qwen3.5-9B, Prime-RL commit
cb74ae2bebe9710b11551970a9661d10116c7179, and
agentboard-babyai-v1-context==0.2.2.
Each run contains:
all 21 LoRA broadcast adapters from steps 5 through 100;
complete training rollouts… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-independent-runs-2-3.agentbench-datasets-v2AgentB-craft_1_stone_pickaxe
AgentB — Craft 1 stone_pickaxe
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Craft 1 stone_pickaxe
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Craft 1 stone_pickaxe")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-craft_1_stone_pickaxe.agentboard-babyai-v1-v071-always-on-switch50
BabyAI v0.7.1 Always-on and Switch-50 Runs
This dataset archives ten Qwen/Qwen3.5-9B Prime-RL runs on the packaged
AgentBoard BabyAI v0.7.1 environment. It contains four always-on objectives and
six schedules that switch objectives at training step 50.
Runs
Always-on runs:
RL-only
ECHO 0.05
ECHO 0.5
ECHO 1.0
Step-50 switch runs:
RL 50, then ECHO 0.05
ECHO 0.05, then RL 50
RL 50, then ECHO 0.5
ECHO 0.5, then RL 50
RL 50, then ECHO 1.0
ECHO 1.0, then RL 50
All… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-always-on-switch50.agentboard-pddl-primerl-checkpointsagentboard-babyai-v1-v071-four-run-checkpoints
AgentBoard BabyAI v1 Prime-RL Four-Run Checkpoints
This dataset preserves four runs from the v071 experiment series using
Qwen/Qwen3.5-9B, Prime-RL 0.7.0 at commit
cb74ae2bebe9710b11551970a9661d10116c7179, and the packaged
agentboard-babyai-v1-context 0.2.2 Verifiers environment.
Runs
Directory
Algorithm
ECHO user alpha
Exact resume step
Final trainer step
Retained adapters
rlonly/
GRPO
0
95
100
21
echo005/
ECHO
0.05
100
100
23
echo050/
ECHO
0.5
95… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-four-run-checkpoints.agentblackbox-rag-repair-outcomes
AgentBlackBox RAG Repair Outcome Dataset
This dataset contains replay-labeled repair outcome data for AgentBlackBox, a counterfactual debugging framework for language agents.
The data is built around failed RAG/document-recall agent traces, candidate repairs, counterfactual replay labels, and repair-ranking evaluation outputs.
Contents
datasets/
world_model_ranker_dataset_v2_train10k/
pointwise/
listwise/
stats.json… See the full description on the dataset page: https://huggingface.co/datasets/Eyerf/agentblackbox-rag-repair-outcomes.agentbake-traces
AgentBake Trace Library
Execution-trace corpus for the AgentBake benchmark (A Personalization Layer
and Benchmark for Heterogeneous Agents). 100 heterogeneous agents spanning
seven frameworks (AutoGen, CrewAI, LangChain, LangGraph, LlamaIndex,
PydanticAI, Strands), each contributing ~10 recorded multi-turn scenarios.
Contents
good_by_framework/
autogen/<agent>/scenario_XXX/
*__trace_sequence.json # per-step activations: input text, output, tools, timing… See the full description on the dataset page: https://huggingface.co/datasets/zzh237/agentbake-traces.AgentB-punch_1_tree_to_collect_wood
AgentB — Punch 1 tree to collect wood
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Punch 1 tree to collect wood
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Punch 1 tree to collect wood")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames:… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-punch_1_tree_to_collect_wood.AgentB-mine_5_coal_ore
AgentB — Mine 5 coal ore
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Mine 5 coal ore
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Mine 5 coal ore")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-mine_5_coal_ore.AgentB-mine_1_wood_log
AgentB — Mine 1 wood log
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Mine 1 wood log
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Mine 1 wood log")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-mine_1_wood_log.AgentB-craft_1_crafting_table
AgentB — Craft 1 crafting_table
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Craft 1 crafting_table
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Craft 1 crafting_table")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-craft_1_crafting_table.AgentB-obtain_5_oak_logs
AgentB — Obtain 5 oak logs
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Obtain 5 oak logs
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Obtain 5 oak logs")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")
Usage… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-obtain_5_oak_logs.AgentB-mine_3_cobblestone
AgentB — Mine 3 cobblestone
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Mine 3 cobblestone
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Mine 3 cobblestone")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")
Usage… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-mine_3_cobblestone.AgentB-mine_1_coal_ore
AgentB — Mine 1 coal ore
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Mine 1 coal ore
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Mine 1 coal ore")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-mine_1_coal_ore.AgentB-mine_3_lapis_ore
AgentB — Mine 3 lapis ore
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Mine 3 lapis ore
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Mine 3 lapis ore")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-mine_3_lapis_ore.AgentB-obtain_5_birch_logs
AgentB — Obtain 5 birch logs
Robot episode dataset uploaded via RoboNet.
Dataset Info
Field
Value
Robot
AgentB
Task
Obtain 5 birch logs
Episodes
1
Format
LeRobot
Usage with LeRobot
from lerobot.common.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("Kota0612/AgentB-Obtain 5 birch logs")
print(f"Number of episodes: {dataset.num_episodes}")
print(f"Number of frames: {dataset.num_frames}")… See the full description on the dataset page: https://huggingface.co/datasets/Kota0612/AgentB-obtain_5_birch_logs.
