CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AgentGym /AgentEvaltext1K<n<10K1 likes302 downloads2y agoHugging Face02agent-evals /core-bench-v1.1-mainlinetextn<1K0 likes168 downloads5mo agoHugging Face03weizhiwang /agent_evaltextn<1K0 likes139 downloads1y agoHugging Face04CaiZhiTech /Evaluation-Dataset-of-AI-Agent-Security-Guardrails DKnownAI Agent Security Evaluation Dataset Data Fields Field Type Description text string The adversarial input (prompt) to be evaluated by a security guardrail action string Human-annotated label: blocked or allowed Citation @misc{li2026comparativeevaluationaiagent, title={A Comparative Evaluation of AI Agent Security Guardrails}, author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.texttext-classification1K<n<10K1 likes104 downloads5mo agoHugging Face05agent-evals /core-bench-v1.1-oodtextn<1K0 likes93 downloads5mo agoHugging Face06nwhite-systems /responsible-agent-workflow-evaluation Responsible Agent Workflow Evaluation Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating whether an AI agent respects safety, permission and accountability boundaries in operational settings. Thirteen categories contain ten scenarios each. Every record includes an intentionally unsafe request, contextual facts, expected safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria and reviewer guidance. This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.textn<1K1 likes85 downloads2mo agoHugging Face07evalstate /fast-agent-slop Transformers PR Slop Dataset Normalized snapshots of issues, pull requests, comments, reviews, and linkage data from evalstate/fast-agent. Files: issues.parquet pull_requests.parquet comments.parquet issue_comments.parquet (derived view of issue discussion comments) pr_comments.parquet (derived view of pull request discussion comments) reviews.parquet pr_files.parquet pr_diffs.parquet review_comments.parquet links.parquet events.parquet new_contributors.parquet… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/fast-agent-slop.tabular1K<n<10K0 likes84 downloads6mo agoHugging Face08Jarvis710 /agent-trajectory-eval-datasettextn<1K0 likes65 downloads5mo agoHugging Face09Etolith /perseval-agent-evaluator-synthetic Perseval Synthetic Agent Evaluator Traces This dataset contains 318 fully synthetic agent traces in 159 matched complete/incomplete scenario groups. It is designed to develop trace projection, task-completion classification, and criterion-level fulfillment models without publishing user traces. What is in the release traces.jsonl: OpenTelemetry-like trace records with no outcome labels in the model-visible payload. task_completion_labels.jsonl: group-aware binary… See the full description on the dataset page: https://huggingface.co/datasets/Etolith/perseval-agent-evaluator-synthetic.texttext-classificationn<1K0 likes62 downloads2mo agoHugging Face10DCAgent2 /eval-fsr-a3-nemotron-gym-agent-swe-r301-tracestext1K<n<10K0 likes60 downloads3mo agoHugging Face11aiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes57 downloads6mo agoHugging Face12DCAgent /eval-SERA-8B_16concurrency_swe_agent_eval_c_terminal-bench-2.0textn<1K0 likes51 downloads7mo agoHugging Face13DCAgent /eval-SERA-32B_16concurrency_swe_agent_eval_c_terminal-bench-2.0textn<1K0 likes51 downloads7mo agoHugging Face14Agent-Eval-Refine /GUI-Dense-Descriptions GUI Screenshots - Dense descrptions Dataset image1K<n<10K5 likes48 downloads2y agoHugging Face15DCAgent2 /eval-SERA-32B_16concurrency_swe_agent_eval_c_terminal-bench-2.0textn<1K0 likes46 downloads7mo agoHugging Face16crimemastergogo22 /agent-tool-risk-evals Agent Tool Risk Evals Tiny Neuron evaluation suite for enterprise AI-agent tool permissions, policy denials, prompt injection, tenant boundaries, and auditability. Source code and evaluator: https://github.com/Sky5595/agent-tool-risk-evals Dataset structure Single JSONL file, tool_authorization_cases.jsonl, with one authorization scenario per line: { "case_id": "authz_001", "category": "excessive_delegation", "agent_task": "Export all customer records to… See the full description on the dataset page: https://huggingface.co/datasets/crimemastergogo22/agent-tool-risk-evals.texttext-classificationn<1K0 likes44 downloads14d agoHugging Face17DCAgent2 /eval-SERA-8B_16concurrency_swe_agent_eval_c_terminal-bench-2.0textn<1K0 likes41 downloads7mo agoHugging Face18DCAgent /eval-R2EGym-32B-Agent_32concurrency_eval_ctx32k_terminal-bench-2.0textn<1K0 likes39 downloads8mo agoHugging Face19anthonyboisbouvier-paris /agent-clash-multi-judge-eval Agent Clash: Multi-Judge LLM Evaluation Dataset Validation data from the paper "Multi-Agent Judging for LLM Evaluation: A Data-Centric Analysis of Concordance with Human Preferences" by Anthony Boisbouvier. This dataset contains 360 pairwise LLM evaluations judged by a panel of three frontier-class LLMs (GPT-5.2, Claude Opus 4.5, Gemini 2.5 Flash) under blind conditions with Borda count aggregation, compared against human preference labels from MT-Bench and Chatbot Arena.… See the full description on the dataset page: https://huggingface.co/datasets/anthonyboisbouvier-paris/agent-clash-multi-judge-eval.tabulartext-classificationn<1K0 likes37 downloads7mo agoHugging Face20Prompt-Pool-Agent /prompt-pool-eval-llm-outputstext1K<n<10K0 likes30 downloads2y agoHugging Face21mukunda1729 /agent-eval-scenarios Agent Eval Scenarios Agent Eval Scenarios is a compact public dataset for lightweight evaluation of AI agents working on practical engineering and operations tasks. It is designed to be: small enough to inspect manually structured enough to extend into a benchmark grounded in real agent workflows such as code review, debugging, docs synthesis, security hardening, UI verification, and workflow automation Files data/agent_eval_scenarios.csv — labeled scenarios with… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/agent-eval-scenarios.texttext-classificationn<1K0 likes30 downloads5mo agoHugging Face22sydv /sleeper-agent-evaluation-datatext10K<n<100K0 likes29 downloads2y agoHugging Face23jeorgexyz /lua-agent-evals Lua Agent Evals The evidence behind lua-agent-lab, from lua-agent. The experiment separates structural tool-call validity from useful tool selection and complete task success. Contents Configuration Unit Method decoder_trials 60 first-turn trials TinyStories 15M; 30 constrained, 30 free; greedy sampling loop_ablations 64 agent runs Eight scripted tasks × eight loop configurations qwen_runs 16 agent runs Eight tasks × constrained/free Qwen 2.5 1.5B… See the full description on the dataset page: https://huggingface.co/datasets/jeorgexyz/lua-agent-evals.tabularn<1K0 likes28 downloads2d agoHugging Face24DCAgent /eval-terminal-bench-2.0__OpenThinker-Agent-v1__eval_ctx32k_non_it_2x_eval_textn<1K0 likes21 downloads6mo agoHugging Face25DCAgent /eval-SERA-32B_16concurrency_swe_agent_eval_c_swebench-verified-random-100-folderstextn<1K0 likes16 downloads7mo agoHugging Face26DCAgent /eval-SERA-8B_16concurrency_swe_agent_eval_c_swebench-verified-random-100-folderstextn<1K0 likes16 downloads7mo agoHugging Face27evalstate /fast-agent-batch-demoSmall public dataset repository used by the fast-agent batch processing documentation. Files: hf-research-questions.jsonl — three demo input rows. hf-research-template.md — row prompt template. hf-research-agent.md — AgentCard connecting the worker to the Hugging Face MCP server. textn<1K0 likes16 downloads4mo agoHugging Face28DCAgent /eval-SERA-8B_16concurrency_swe_agent_eval_c_OpenThoughts-TB-devtextn<1K0 likes15 downloads7mo agoHugging Face29RAIA-BRASIL /agent_eval_questions Agent Questions Generated Pt-Br Detalhes do Dataset Descrição Desenvolvidor por: Álvaro Lopes. Linkedin Artur de Vlieger Linkedin Fabrício Salomon Linkedin Leticia Bossatto Marchezi Linkedin Luis Felipe Jorge Linkedin Otávio Coletti Linkedin Patrocinado por : Pico Língua(s) (NLP) :Português Correspondência: raia.projetos@gmail.com, leticiabossatto@gmail.com Fontes Repositório: Github Uso O dataset pode ser utilizado… See the full description on the dataset page: https://huggingface.co/datasets/RAIA-BRASIL/agent_eval_questions.textn<1K0 likes14 downloads1y agoHugging Face30dalek-ai /agent-tool-router-eval-fr agent-tool-router · parallel EN/FR evaluation 50 parallel English/French queries used to evaluate dalek-ai/baseline-v1-desc-hybrid (EN-first) versus dalek-ai/baseline-v1-desc-hybrid-multilingual (50+ languages) on a catalog of 18 671 tools collected from public agent benchmarks (tau-bench, Hermes function-calling-v1, ToolACE). Numbers (hybrid models, α=0.5, V=18 671) model top-3 EN top-3 FR baseline-v1-desc-hybrid (default, MiniLM-L6) 82% 26%… See the full description on the dataset page: https://huggingface.co/datasets/dalek-ai/agent-tool-router-eval-fr.texttext-retrievaln<1K0 likes13 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.