datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgentEvalcore-bench-v1.1-mainlineEvaluation-Dataset-of-AI-Agent-Security-Guardrails
DKnownAI Agent Security Evaluation Dataset
Data Fields
Field
Type
Description
text
string
The adversarial input (prompt) to be evaluated by a security guardrail
action
string
Human-annotated label: blocked or allowed
Citation
@misc{li2026comparativeevaluationaiagent,
title={A Comparative Evaluation of AI Agent Security Guardrails},
author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.core-bench-v1.1-oodperseval-agent-evaluator-synthetic
Perseval Synthetic Agent Evaluator Traces
This dataset contains 318 fully synthetic agent traces in
159 matched complete/incomplete scenario groups. It
is designed to develop trace projection, task-completion classification, and
criterion-level fulfillment models without publishing user traces.
What is in the release
traces.jsonl: OpenTelemetry-like trace records with no outcome labels in the
model-visible payload.
task_completion_labels.jsonl: group-aware binary… See the full description on the dataset page: https://huggingface.co/datasets/Etolith/perseval-agent-evaluator-synthetic.agent-tool-risk-evals
Agent Tool Risk Evals
Tiny Neuron evaluation suite for enterprise AI-agent tool permissions, policy denials, prompt injection, tenant boundaries, and auditability.
Source code and evaluator: https://github.com/Sky5595/agent-tool-risk-evals
Dataset structure
Single JSONL file, tool_authorization_cases.jsonl, with one authorization scenario per line:
{
"case_id": "authz_001",
"category": "excessive_delegation",
"agent_task": "Export all customer records to… See the full description on the dataset page: https://huggingface.co/datasets/crimemastergogo22/agent-tool-risk-evals.lua-agent-evals
Lua Agent Evals
The evidence behind lua-agent-lab,
from lua-agent. The experiment separates
structural tool-call validity from useful tool selection and complete task success.
Contents
Configuration
Unit
Method
decoder_trials
60 first-turn trials
TinyStories 15M; 30 constrained, 30 free; greedy sampling
loop_ablations
64 agent runs
Eight scripted tasks × eight loop configurations
qwen_runs
16 agent runs
Eight tasks × constrained/free Qwen 2.5 1.5B… See the full description on the dataset page: https://huggingface.co/datasets/jeorgexyz/lua-agent-evals.fast-agent-batch-demoSmall public dataset repository used by the fast-agent batch processing documentation.
Files:
hf-research-questions.jsonl — three demo input rows.
hf-research-template.md — row prompt template.
hf-research-agent.md — AgentCard connecting the worker to the Hugging Face MCP server.
DROP-evalagent-tool-router-eval-fr
agent-tool-router · parallel EN/FR evaluation
50 parallel English/French queries used to evaluate
dalek-ai/baseline-v1-desc-hybrid
(EN-first) versus
dalek-ai/baseline-v1-desc-hybrid-multilingual
(50+ languages) on a catalog of 18 671 tools collected from public agent
benchmarks (tau-bench, Hermes function-calling-v1, ToolACE).
Numbers (hybrid models, α=0.5, V=18 671)
model
top-3 EN
top-3 FR
baseline-v1-desc-hybrid (default, MiniLM-L6)
82%
26%… See the full description on the dataset page: https://huggingface.co/datasets/dalek-ai/agent-tool-router-eval-fr.defendable-pain-agent-eval-gaming-v0.1
Agent Eval Gaming Pain Receipt
"the gamed score" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 2 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
2 pain… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-agent-eval-gaming-v0.1.IIRC-evalstarcoder-agent-eval
starcoder-agent-eval
Eval records for Colby/starcoder-7b-agent checkpoints (seed=999).
Source
Count
System prompt
Tool calls
Crownelius/Opus-4.6-Reasoning-3300x
20
none
none
Roman1111111/claude-opus-4.6-10000x
20
from record
none
All expected answers are short, auto-verifiable values (numbers, bools, short lists).
No synthetic ANSWER: prompt. Do not add these records to training data.
agent-eval-multi-turn
