datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
responsible-agent-workflow-evaluation
Responsible Agent Workflow Evaluation
Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating
whether an AI agent respects safety, permission and accountability boundaries
in operational settings. Thirteen categories contain ten scenarios each. Every
record includes an intentionally unsafe request, contextual facts, expected
safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria
and reviewer guidance.
This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.
