CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AgentGym /AgentEvaltext1K<n<10K1 likes306 downloads2y agoHugging Face02agent-evals /core-bench-v1.1-mainlinetextn<1K0 likes162 downloads5mo agoHugging Face03CaiZhiTech /Evaluation-Dataset-of-AI-Agent-Security-Guardrails DKnownAI Agent Security Evaluation Dataset Data Fields Field Type Description text string The adversarial input (prompt) to be evaluated by a security guardrail action string Human-annotated label: blocked or allowed Citation @misc{li2026comparativeevaluationaiagent, title={A Comparative Evaluation of AI Agent Security Guardrails}, author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.texttext-classification1K<n<10K1 likes98 downloads5mo agoHugging Face04agent-evals /core-bench-v1.1-oodtextn<1K0 likes89 downloads5mo agoHugging Face05Etolith /perseval-agent-evaluator-synthetic Perseval Synthetic Agent Evaluator Traces This dataset contains 318 fully synthetic agent traces in 159 matched complete/incomplete scenario groups. It is designed to develop trace projection, task-completion classification, and criterion-level fulfillment models without publishing user traces. What is in the release traces.jsonl: OpenTelemetry-like trace records with no outcome labels in the model-visible payload. task_completion_labels.jsonl: group-aware binary… See the full description on the dataset page: https://huggingface.co/datasets/Etolith/perseval-agent-evaluator-synthetic.texttext-classificationn<1K0 likes59 downloads2mo agoHugging Face06crimemastergogo22 /agent-tool-risk-evals Agent Tool Risk Evals Tiny Neuron evaluation suite for enterprise AI-agent tool permissions, policy denials, prompt injection, tenant boundaries, and auditability. Source code and evaluator: https://github.com/Sky5595/agent-tool-risk-evals Dataset structure Single JSONL file, tool_authorization_cases.jsonl, with one authorization scenario per line: { "case_id": "authz_001", "category": "excessive_delegation", "agent_task": "Export all customer records to… See the full description on the dataset page: https://huggingface.co/datasets/crimemastergogo22/agent-tool-risk-evals.texttext-classificationn<1K0 likes44 downloads15d agoHugging Face07jeorgexyz /lua-agent-evals Lua Agent Evals The evidence behind lua-agent-lab, from lua-agent. The experiment separates structural tool-call validity from useful tool selection and complete task success. Contents Configuration Unit Method decoder_trials 60 first-turn trials TinyStories 15M; 30 constrained, 30 free; greedy sampling loop_ablations 64 agent runs Eight scripted tasks × eight loop configurations qwen_runs 16 agent runs Eight tasks × constrained/free Qwen 2.5 1.5B… See the full description on the dataset page: https://huggingface.co/datasets/jeorgexyz/lua-agent-evals.tabularn<1K0 likes36 downloads3d agoHugging Face08evalstate /fast-agent-batch-demoSmall public dataset repository used by the fast-agent batch processing documentation. Files: hf-research-questions.jsonl — three demo input rows. hf-research-template.md — row prompt template. hf-research-agent.md — AgentCard connecting the worker to the Hugging Face MCP server. textn<1K0 likes16 downloads4mo agoHugging Face09agent-husky /DROP-evaltextn<1K0 likes12 downloads2y agoHugging Face10dalek-ai /agent-tool-router-eval-fr agent-tool-router · parallel EN/FR evaluation 50 parallel English/French queries used to evaluate dalek-ai/baseline-v1-desc-hybrid (EN-first) versus dalek-ai/baseline-v1-desc-hybrid-multilingual (50+ languages) on a catalog of 18 671 tools collected from public agent benchmarks (tau-bench, Hermes function-calling-v1, ToolACE). Numbers (hybrid models, α=0.5, V=18 671) model top-3 EN top-3 FR baseline-v1-desc-hybrid (default, MiniLM-L6) 82% 26%… See the full description on the dataset page: https://huggingface.co/datasets/dalek-ai/agent-tool-router-eval-fr.texttext-retrievaln<1K0 likes11 downloads5mo agoHugging Face11SwarmandBee /defendable-pain-agent-eval-gaming-v0.1 Agent Eval Gaming Pain Receipt "the gamed score" — Mr. Defendable A free pain-receipt dataset from the DefendableOS ecosystem. 2 rows · ready to read · all cited or graded · CC-BY-4.0. Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them. Tribunal begins before training. No proof, no honey. To the shed. What's in here 2 pain… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-agent-eval-gaming-v0.1.texttext-classificationn<1K0 likes11 downloads4mo agoHugging Face12agent-husky /IIRC-evaltextn<1K0 likes8 downloads2y agoHugging Face13Colby /starcoder-agent-eval starcoder-agent-eval Eval records for Colby/starcoder-7b-agent checkpoints (seed=999). Source Count System prompt Tool calls Crownelius/Opus-4.6-Reasoning-3300x 20 none none Roman1111111/claude-opus-4.6-10000x 20 from record none All expected answers are short, auto-verifiable values (numbers, bools, short lists). No synthetic ANSWER: prompt. Do not add these records to training data. textn<1K0 likes6 downloads4mo agoHugging Face14nolabs /agent-eval-multi-turntextn<1K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.