datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentboard
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
This is the official dataset repository of AgentBoard.
1. Data Overview
AgentBoard is composed of 9 diverse tasks which can be divided into 4 types, including Embodied AI, Game, Web, and Tool:
Embodied AI
Game
Web
Tool
AlfWorld
ScienceWorld
BabyAI
Jericho
PDDL
WebShop
WebArena… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/agentboard.agentbattler-bench
AgentBattler Mini Ledger V5
Immutable evidence for 15/15 accepted Mini Ledger V5 runs across 3 harness × model conditions.
What is here
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/site/terminal-campaign.json: compact website and analysis input.
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/campaign.json: source-revision-preserving campaign index with host paths removed.
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/runs/:… See the full description on the dataset page: https://huggingface.co/datasets/techfren/agentbattler-bench.JL-AgentBehavior-10K
JL-AgentBehavior-10K
JL-AgentBehavior-10K is a 10,000-record, English-language research-preview dataset for studying and training the behavioral policy of repository-level coding agents.
The dataset does not treat a coding agent as a chatbot that maps a request directly to a block of code. It represents an agent as a policy operating across a sequence of observable decisions:
task
-> repository evidence
-> bounded plan
-> tool selection
-> scoped edit strategy
->… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-AgentBehavior-10K.agentboard-babyai-v1-v071-independent-runs-2-3
BabyAI v0.7.1 Independent Runs 2 and 3
This dataset is the companion to
bhoy/agentboard-babyai-v1-v071-four-run-checkpoints. It preserves the second
and third independent 100-step runs for four objectives:
RL-only (GRPO)
ECHO 0.05
ECHO 0.5
ECHO 1.0
All runs use Qwen/Qwen3.5-9B, Prime-RL commit
cb74ae2bebe9710b11551970a9661d10116c7179, and
agentboard-babyai-v1-context==0.2.2.
Each run contains:
all 21 LoRA broadcast adapters from steps 5 through 100;
complete training rollouts… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-independent-runs-2-3.agentboard-babyai-v1-v071-always-on-switch50
BabyAI v0.7.1 Always-on and Switch-50 Runs
This dataset archives ten Qwen/Qwen3.5-9B Prime-RL runs on the packaged
AgentBoard BabyAI v0.7.1 environment. It contains four always-on objectives and
six schedules that switch objectives at training step 50.
Runs
Always-on runs:
RL-only
ECHO 0.05
ECHO 0.5
ECHO 1.0
Step-50 switch runs:
RL 50, then ECHO 0.05
ECHO 0.05, then RL 50
RL 50, then ECHO 0.5
ECHO 0.5, then RL 50
RL 50, then ECHO 1.0
ECHO 1.0, then RL 50
All… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-always-on-switch50.agentboard-babyai-v1-v071-four-run-checkpoints
AgentBoard BabyAI v1 Prime-RL Four-Run Checkpoints
This dataset preserves four runs from the v071 experiment series using
Qwen/Qwen3.5-9B, Prime-RL 0.7.0 at commit
cb74ae2bebe9710b11551970a9661d10116c7179, and the packaged
agentboard-babyai-v1-context 0.2.2 Verifiers environment.
Runs
Directory
Algorithm
ECHO user alpha
Exact resume step
Final trainer step
Retained adapters
rlonly/
GRPO
0
95
100
21
echo005/
ECHO
0.05
100
100
23
echo050/
ECHO
0.5
95… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-four-run-checkpoints.Agent-browser-taskagent-budget-violations
agent-budget-violations
15 synthetic agent runs annotated with their budget (cost / tool-call / wall-time caps), actual usage, violation types, and a one-line root cause + fix. Built as fixtures for budget-enforcement tests, alerting heuristics, and observability dashboards.
5 of the 15 are clean (no violations) so you can test the "no false positive" path.
Violation breakdown
Violation type
Count
cost
4
tool_calls
6
wall_time
4
None (clean)
5… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/agent-budget-violations.DARA-Agentbench
Dataset Card for DARA-Agentbench
Dataset Summary
This dataset contains 577 curated reasoning trajectories for KGQA LLM-based agents in the Agentbench format. (https://github.com/UKPLab/acl2024-DARA). It is sourced from GrailQA, WebQSP, and GraphQ.
The fields include:
raw question: original question
input: question with the linked entities.
output: The step-by-step reasoning trajectory to construct the full logical form.
variable list: The stepwise logical forms… See the full description on the dataset page: https://huggingface.co/datasets/UKPLab/DARA-Agentbench.agentbench-sft-trajectories-v3-planA
agentbench-sft-trajectories-v3-planA
Merged dataset for LLM Advanced Competition SFT (Plan A).
Composition
Source
Records
Pct
u-10bei/dbbench_sft_dataset_react (v1)
300
5.5%
u-10bei/dbbench_sft_dataset_react_v2
360
6.6%
u-10bei/dbbench_sft_dataset_react_v3
1,112
20.3%
u-10bei/dbbench_sft_dataset_react_v4
1,194
21.8%
u-10bei/sft_alfworld_trajectory_dataset_v5
2,502
45.8%
Total (after dedup)
5,468
100%
Design
All 4 versions of… See the full description on the dataset page: https://huggingface.co/datasets/sabia0080/agentbench-sft-trajectories-v3-planA.agentbench-sft-trajectories-v2-dbexpandagentbench-sft-trajectories-v4-longexp-fullagentbench-sft-trajectories-v2agentbench-sft-trajectories-v1agentBioagentbench-sft-trajectories-v2-enhanceddbbench_cot_enriched_for_agentbench
