datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-benchmark
configs:
- config_name: default
data_files:
- path: run-log.json
split: train
language:
- en
license: cc-by-nc-4.0
tags:
- ai-agents
- benchmark
- evaluation
- agent-evaluation
- llm-agents
- trustworthy-ai
- product-review
- c2pa
pretty_name: Hlido AI Agent Benchmark
size_categories:
- n<1K
task_categories:
- other
Hlido AI Agent Benchmark
Independent, cryptographically-attested evaluations of AI agents and products.
Hlido reviews AI… See the full description on the dataset page: https://huggingface.co/datasets/hlido-eu/agent-benchmark.agentbenchAgentBenchKo-AgentBench
🇰🇷 Ko-AgentBench v1
"한국 에이전트 벤치마크 프로젝트"
English | 한국어
⚠️ 벤치마크 평가를 진행하시려면 GitHub Repository를 방문해주세요.
이 데이터셋은 벤치마크 태스크 정보만 포함하고 있습니다. 실제 평가 코드, API 도구, 평가 메트릭 등은 GitHub에서 확인하실 수 있습니다.
AI 에이전트의 능력이 고도화되면서, 그 성능을 실제 환경과 유사한 조건에서 정밀하게 측정하는 것이 중요해졌습니다. 하지만 대부분의 벤치마크는 영어권 환경을 기준으로 설계되어, 한국의 특수한 사용 맥락을 반영하는 데 한계가 있었습니다.
이러한 문제를 해결하기 위해, 한국 실사용 환경에 특화된 고품질 에이전트 벤치마크를 개발하였습니다.
Ko-AgentBench 핵심 특징 ✨
1. 단계별 태스크 설계
단순 도구 호출부터 장기적 맥락 능력, 강건성 처리 능력까지 에이전트의 능력을 7단계로… See the full description on the dataset page: https://huggingface.co/datasets/huggingface-KREW/Ko-AgentBench.AgentBench-EvoSyn
EvoSyn-AgentBench-OS: Evolutionary Synthesized OS Agent Tasks
Dataset Description
This dataset contains high-quality operating system agent tasks synthesized and filtered using the EvoSyn framework. Each task includes the question description, system initialization scripts, and discriminative test scripts.
We divide these OS tasks into two categories: one requires the model to provide a final result - the QA type, and the other requires the model to complete a task - the… See the full description on the dataset page: https://huggingface.co/datasets/Elynden/AgentBench-EvoSyn.Simia-AgentBench-SFT-15k
🐒 Simia-AgentBench-SFT-15k:
Simia-AgentBench-SFT-15k is the fully synthetic tool-agent dataset, designed to advance tool use for AgentBench (webshop, mind2web, Operating System). It comprises nearly 15k synthesized trajectories from Agenttuning (webshop, mind2web, Operating System). Models fine-tuned on this dataset outperform much larger closed-source counterparts on AgentBench (webshop, mind2web, Operating System).
📄 Technical Report - Discover the methodology and technical… See the full description on the dataset page: https://huggingface.co/datasets/Simia-Agent/Simia-AgentBench-SFT-15k.dbbench_cot_enriched_for_agentbenchagentbench-datasets-v2agentbench-sft-trajectories-v4-longexp-fullDARA-Agentbench
Dataset Card for DARA-Agentbench
Dataset Summary
This dataset contains 577 curated reasoning trajectories for KGQA LLM-based agents in the Agentbench format. (https://github.com/UKPLab/acl2024-DARA). It is sourced from GrailQA, WebQSP, and GraphQ.
The fields include:
raw question: original question
input: question with the linked entities.
output: The step-by-step reasoning trajectory to construct the full logical form.
variable list: The stepwise logical forms… See the full description on the dataset page: https://huggingface.co/datasets/UKPLab/DARA-Agentbench.mixed_agentbench_v3agentbench-sft-trajectories-v2-dbexpandagentbench-sft-trajectories-v3-planA
agentbench-sft-trajectories-v3-planA
Merged dataset for LLM Advanced Competition SFT (Plan A).
Composition
Source
Records
Pct
u-10bei/dbbench_sft_dataset_react (v1)
300
5.5%
u-10bei/dbbench_sft_dataset_react_v2
360
6.6%
u-10bei/dbbench_sft_dataset_react_v3
1,112
20.3%
u-10bei/dbbench_sft_dataset_react_v4
1,194
21.8%
u-10bei/sft_alfworld_trajectory_dataset_v5
2,502
45.8%
Total (after dedup)
5,468
100%
Design
All 4 versions of… See the full description on the dataset page: https://huggingface.co/datasets/sabia0080/agentbench-sft-trajectories-v3-planA.agentbench-sft-trajectories-v2-enhancedAgentBench_dbbenchdbbench_cleaned_for_agentbench
DBBench Cleaned for AgentBench
u-10bei/dbbench_sft_dataset_react_v4(1,200 件)に対してクレンジング処理を施したデータセット。
AgentBench DBBench 評価用の SFT 訓練データとしてそのまま使用可能。
混合利用を想定: 本データセットは mark-22/dbbench-spider-3500(1,697 件)と混合し、合計 2,897 件 の SFT データとして使用することを想定しています。
Dataset Summary
Metric
Value
Total rows
1,200
Source
u-10bei/dbbench_sft_dataset_react_v4
Avg messages per item
6.7
Items with Final Answer1,200 / 1,200 (100%)
Columns
id, messages, metadata… See the full description on the dataset page: https://huggingface.co/datasets/mark-22/dbbench_cleaned_for_agentbench.alfworld_cleaned_for_agentbench_v5agent_bench_tinyKo-AgentBench
🇰🇷 Ko-AgentBench v1
"한국 에이전트 벤치마크 프로젝트"
English | 한국어
⚠️ 벤치마크 평가를 진행하시려면 GitHub Repository를 방문해주세요.
이 데이터셋은 벤치마크 태스크 정보만 포함하고 있습니다. 실제 평가 코드, API 도구, 평가 메트릭 등은 GitHub에서 확인하실 수 있습니다.
AI 에이전트의 능력이 고도화되면서, 그 성능을 실제 환경과 유사한 조건에서 정밀하게 측정하는 것이 중요해졌습니다. 하지만 대부분의 벤치마크는 영어권 환경을 기준으로 설계되어, 한국의 특수한 사용 맥락을 반영하는 데 한계가 있었습니다.
이러한 문제를 해결하기 위해, 한국 실사용 환경에 특화된 고품질 에이전트 벤치마크를 개발하였습니다.
Ko-AgentBench 핵심 특징 ✨
1. 단계별 태스크 설계
단순 도구 호출부터 장기적 맥락 능력, 강건성 처리 능력까지 에이전트의 능력을 7단계로… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluodeng/Ko-AgentBench.agentbench-sft-trajectories-v2AgentBenchInstruct_d1mixed_agentbench_v2alfworld_cleaned_for_agentbench_v4mixed_agentbench_v1agentbench_sft_mix_alfworld_dbbench_v1
AgentBench SFT Mix (ALFWorld + DBBench)
This dataset is a mixed SFT dataset created by concatenating and shuffling:
u-10bei/sft_alfworld_trajectory_dataset_v5
u-10bei/dbbench_sft_dataset_react_v4
Fields
messages: multi-turn chat messages (role/content)
tools: optional tool schemas (if present)
Credits
This dataset is a mixed and reformatted version of the original datasets listed above.
Please refer to each source dataset for their respective licenses and… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/agentbench_sft_mix_alfworld_dbbench_v1.sft_agentbench_combined_preprocessedagentbench-resultsagentbench-sft-trajectories-v1agentbench_mix_alf3_db1_v1agentbench-db-alfw-merged-362
