CoolFace
Datasetpublic

hubertmarek/agent-diff-bench

Agent-Diff Bench Website | Paper | GitHub Agent-Diff is a benchmarking framework for evaluating agentic Large Language Models (LLMs) on real-world tasks that execute code via external APIs. The benchmark provides access to real API interfaces (Slack, Box, Linear, Google Calendar) while sandboxing the environment in which calls are made and evaluated. Dataset Summary The dataset contains 224 tasks utilizing enterprise software workflows, provided with an 80/20… See the full description on the dataset page: https://huggingface.co/datasets/hubertmarek/agent-diff-bench.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
2likes221downloads
filetest.parquet36 KBdownload
filetrain.parquet84 KBdownload

hubertmarek/agent-diff-bench · main · files are served by the source, never re-hosted here