datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgentEvalcore-bench-v1.1-mainlineagent_evalEvaluation-Dataset-of-AI-Agent-Security-Guardrails
DKnownAI Agent Security Evaluation Dataset
Data Fields
Field
Type
Description
text
string
The adversarial input (prompt) to be evaluated by a security guardrail
action
string
Human-annotated label: blocked or allowed
Citation
@misc{li2026comparativeevaluationaiagent,
title={A Comparative Evaluation of AI Agent Security Guardrails},
author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.core-bench-v1.1-oodresponsible-agent-workflow-evaluation
Responsible Agent Workflow Evaluation
Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating
whether an AI agent respects safety, permission and accountability boundaries
in operational settings. Thirteen categories contain ten scenarios each. Every
record includes an intentionally unsafe request, contextual facts, expected
safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria
and reviewer guidance.
This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.fast-agent-slop
Transformers PR Slop Dataset
Normalized snapshots of issues, pull requests, comments, reviews, and linkage data from evalstate/fast-agent.
Files:
issues.parquet
pull_requests.parquet
comments.parquet
issue_comments.parquet (derived view of issue discussion comments)
pr_comments.parquet (derived view of pull request discussion comments)
reviews.parquet
pr_files.parquet
pr_diffs.parquet
review_comments.parquet
links.parquet
events.parquet
new_contributors.parquet… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/fast-agent-slop.agent-trajectory-eval-datasetperseval-agent-evaluator-synthetic
Perseval Synthetic Agent Evaluator Traces
This dataset contains 318 fully synthetic agent traces in
159 matched complete/incomplete scenario groups. It
is designed to develop trace projection, task-completion classification, and
criterion-level fulfillment models without publishing user traces.
What is in the release
traces.jsonl: OpenTelemetry-like trace records with no outcome labels in the
model-visible payload.
task_completion_labels.jsonl: group-aware binary… See the full description on the dataset page: https://huggingface.co/datasets/Etolith/perseval-agent-evaluator-synthetic.eval-fsr-a3-nemotron-gym-agent-swe-r301-tracesagent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.eval-SERA-8B_16concurrency_swe_agent_eval_c_terminal-bench-2.0eval-SERA-32B_16concurrency_swe_agent_eval_c_terminal-bench-2.0GUI-Dense-Descriptions
GUI Screenshots - Dense descrptions Dataset
eval-SERA-32B_16concurrency_swe_agent_eval_c_terminal-bench-2.0agent-tool-risk-evals
Agent Tool Risk Evals
Tiny Neuron evaluation suite for enterprise AI-agent tool permissions, policy denials, prompt injection, tenant boundaries, and auditability.
Source code and evaluator: https://github.com/Sky5595/agent-tool-risk-evals
Dataset structure
Single JSONL file, tool_authorization_cases.jsonl, with one authorization scenario per line:
{
"case_id": "authz_001",
"category": "excessive_delegation",
"agent_task": "Export all customer records to… See the full description on the dataset page: https://huggingface.co/datasets/crimemastergogo22/agent-tool-risk-evals.eval-SERA-8B_16concurrency_swe_agent_eval_c_terminal-bench-2.0eval-R2EGym-32B-Agent_32concurrency_eval_ctx32k_terminal-bench-2.0agent-clash-multi-judge-eval
Agent Clash: Multi-Judge LLM Evaluation Dataset
Validation data from the paper "Multi-Agent Judging for LLM Evaluation: A Data-Centric Analysis of Concordance with Human Preferences" by Anthony Boisbouvier.
This dataset contains 360 pairwise LLM evaluations judged by a panel of three frontier-class LLMs (GPT-5.2, Claude Opus 4.5, Gemini 2.5 Flash) under blind conditions with Borda count aggregation, compared against human preference labels from MT-Bench and Chatbot Arena.… See the full description on the dataset page: https://huggingface.co/datasets/anthonyboisbouvier-paris/agent-clash-multi-judge-eval.prompt-pool-eval-llm-outputsagent-eval-scenarios
Agent Eval Scenarios
Agent Eval Scenarios is a compact public dataset for lightweight evaluation of AI agents working on practical engineering and operations tasks.
It is designed to be:
small enough to inspect manually
structured enough to extend into a benchmark
grounded in real agent workflows such as code review, debugging, docs synthesis, security hardening, UI verification, and workflow automation
Files
data/agent_eval_scenarios.csv — labeled scenarios with… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/agent-eval-scenarios.sleeper-agent-evaluation-datalua-agent-evals
Lua Agent Evals
The evidence behind lua-agent-lab,
from lua-agent. The experiment separates
structural tool-call validity from useful tool selection and complete task success.
Contents
Configuration
Unit
Method
decoder_trials
60 first-turn trials
TinyStories 15M; 30 constrained, 30 free; greedy sampling
loop_ablations
64 agent runs
Eight scripted tasks × eight loop configurations
qwen_runs
16 agent runs
Eight tasks × constrained/free Qwen 2.5 1.5B… See the full description on the dataset page: https://huggingface.co/datasets/jeorgexyz/lua-agent-evals.eval-terminal-bench-2.0__OpenThinker-Agent-v1__eval_ctx32k_non_it_2x_eval_eval-SERA-32B_16concurrency_swe_agent_eval_c_swebench-verified-random-100-folderseval-SERA-8B_16concurrency_swe_agent_eval_c_swebench-verified-random-100-foldersfast-agent-batch-demoSmall public dataset repository used by the fast-agent batch processing documentation.
Files:
hf-research-questions.jsonl — three demo input rows.
hf-research-template.md — row prompt template.
hf-research-agent.md — AgentCard connecting the worker to the Hugging Face MCP server.
eval-SERA-8B_16concurrency_swe_agent_eval_c_OpenThoughts-TB-devagent_eval_questions
Agent Questions Generated Pt-Br
Detalhes do Dataset
Descrição
Desenvolvidor por:
Álvaro Lopes. Linkedin
Artur de Vlieger Linkedin
Fabrício Salomon Linkedin
Leticia Bossatto Marchezi Linkedin
Luis Felipe Jorge Linkedin
Otávio Coletti Linkedin
Patrocinado por : Pico
Língua(s) (NLP) :Português
Correspondência: raia.projetos@gmail.com, leticiabossatto@gmail.com
Fontes
Repositório: Github
Uso
O dataset pode ser utilizado… See the full description on the dataset page: https://huggingface.co/datasets/RAIA-BRASIL/agent_eval_questions.agent-tool-router-eval-fr
agent-tool-router · parallel EN/FR evaluation
50 parallel English/French queries used to evaluate
dalek-ai/baseline-v1-desc-hybrid
(EN-first) versus
dalek-ai/baseline-v1-desc-hybrid-multilingual
(50+ languages) on a catalog of 18 671 tools collected from public agent
benchmarks (tau-bench, Hermes function-calling-v1, ToolACE).
Numbers (hybrid models, α=0.5, V=18 671)
model
top-3 EN
top-3 FR
baseline-v1-desc-hybrid (default, MiniLM-L6)
82%
26%… See the full description on the dataset page: https://huggingface.co/datasets/dalek-ai/agent-tool-router-eval-fr.
