agentbench
agent-benchmark
configs:
- config_name: default
data_files:
- path: run-log.json
split: train
language:
- en
license: cc-by-nc-4.0
tags:
- ai-agents
- benchmark
- evaluation
- agent-evaluation
- llm-agents
- trustworthy-ai
- product-review
- c2pa
pretty_name: Hlido AI Agent Benchmark
size_categories:
- n<1K
task_categories:
- other
Hlido AI Agent Benchmark
Independent, cryptographically-attested evaluations of AI agents and products.
Hlido reviews AI… See the full description on the dataset page: https://huggingface.co/datasets/hlido-eu/agent-benchmark.agentbenchAgentBenchKo-AgentBench
🇰🇷 Ko-AgentBench v1
"한국 에이전트 벤치마크 프로젝트"
English | 한국어
⚠️ 벤치마크 평가를 진행하시려면 GitHub Repository를 방문해주세요.
이 데이터셋은 벤치마크 태스크 정보만 포함하고 있습니다. 실제 평가 코드, API 도구, 평가 메트릭 등은 GitHub에서 확인하실 수 있습니다.
AI 에이전트의 능력이 고도화되면서, 그 성능을 실제 환경과 유사한 조건에서 정밀하게 측정하는 것이 중요해졌습니다. 하지만 대부분의 벤치마크는 영어권 환경을 기준으로 설계되어, 한국의 특수한 사용 맥락을 반영하는 데 한계가 있었습니다.
이러한 문제를 해결하기 위해, 한국 실사용 환경에 특화된 고품질 에이전트 벤치마크를 개발하였습니다.
Ko-AgentBench 핵심 특징 ✨
1. 단계별 태스크 설계
단순 도구 호출부터 장기적 맥락 능력, 강건성 처리 능력까지 에이전트의 능력을 7단계로… See the full description on the dataset page: https://huggingface.co/datasets/huggingface-KREW/Ko-AgentBench.AgentBench-EvoSyn
EvoSyn-AgentBench-OS: Evolutionary Synthesized OS Agent Tasks
Dataset Description
This dataset contains high-quality operating system agent tasks synthesized and filtered using the EvoSyn framework. Each task includes the question description, system initialization scripts, and discriminative test scripts.
We divide these OS tasks into two categories: one requires the model to provide a final result - the QA type, and the other requires the model to complete a task - the… See the full description on the dataset page: https://huggingface.co/datasets/Elynden/AgentBench-EvoSyn.Simia-AgentBench-SFT-15k
🐒 Simia-AgentBench-SFT-15k:
Simia-AgentBench-SFT-15k is the fully synthetic tool-agent dataset, designed to advance tool use for AgentBench (webshop, mind2web, Operating System). It comprises nearly 15k synthesized trajectories from Agenttuning (webshop, mind2web, Operating System). Models fine-tuned on this dataset outperform much larger closed-source counterparts on AgentBench (webshop, mind2web, Operating System).
📄 Technical Report - Discover the methodology and technical… See the full description on the dataset page: https://huggingface.co/datasets/Simia-Agent/Simia-AgentBench-SFT-15k.
