CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jedisct1 /fastly-agent-toolkit-evals Fastly Agent Toolkit Evals Evaluation dataset for the Fastly Agent Toolkit. Measures how well AI models complete Fastly-specific engineering tasks, with and without toolkit skills loaded. What this dataset contains Each entry is a full evaluation run: a task prompt, model configuration, the model's output, tool call traces, and grading results. The key comparison is with_skill (toolkit loaded) vs without_skill (no toolkit), across multiple models.… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/fastly-agent-toolkit-evals.text-generationn<1K1 likes159 downloads23d agoHugging Face02aiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes57 downloads6mo agoHugging Face03Mosescreates /arabic-agent-eval Arabic Agent Eval — Dataset Card An open, installable Arabic function-calling benchmark with dialect splits. Dataset summary 51 evaluation items spanning 6 categories and 5 dialects of Arabic, testing whether large language models can (a) select the right tool, (b) extract arguments from natural Arabic instructions, (c) preserve Arabic text in tool arguments instead of transliterating, and (d) understand dialectal framing. Supported tasks… See the full description on the dataset page: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval.question-answeringn<1K0 likes44 downloads4mo agoHugging Face04chjp0632 /Agent-eval-Effector-Hunt Agent Eval: Effector Hunt While AI scientist agents like Claude Science and Google's AI co-scientist highlight the potential of autonomous research, compact and reproducible datasets for evaluating these agents on real scientific workflows remain scarce. Agent Eval: Effector Hunt is a genomics benchmark package designed around a real scientific discovery workflow from the Science paper Chen et al. 2017. It asks an AI agent, a computational biologist, or a hybrid human-agent… See the full description on the dataset page: https://huggingface.co/datasets/chjp0632/Agent-eval-Effector-Hunt.question-answering0 likes39 downloads3mo agoHugging Face05mukunda1729 /agent-eval-scenarios Agent Eval Scenarios Agent Eval Scenarios is a compact public dataset for lightweight evaluation of AI agents working on practical engineering and operations tasks. It is designed to be: small enough to inspect manually structured enough to extend into a benchmark grounded in real agent workflows such as code review, debugging, docs synthesis, security hardening, UI verification, and workflow automation Files data/agent_eval_scenarios.csv — labeled scenarios with… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/agent-eval-scenarios.texttext-classificationn<1K0 likes30 downloads5mo agoHugging Face06vanishingradient /mlops-agent-rl-eval MLOps Agent RL — Evaluation Results Offline multi-turn tool-calling evaluation on the mocked MLOpsEnv incident suite (6 scenarios). Models base: Qwen/Qwen2.5-0.5B-Instruct grpo: vanishingradient/mlops-agent-rl-qwen05b-grpo (LoRA adapter) oracle: scripted expected tools + preferred remediation Files episodes.jsonl — full trajectories + metrics episodes_slim.jsonl — metrics without trajectories summary.json — aggregates for the paper config.json —… See the full description on the dataset page: https://huggingface.co/datasets/vanishingradient/mlops-agent-rl-eval.reinforcement-learning0 likes26 downloads2mo agoHugging Face07yzhou05 /eval-agent-traceEval Agent Trace of a MLE Agent by Celestra. Sythetically Generated by gpt 5.2 thinking texttext-classificationn<1K1 likes10 downloads9mo agoHugging Face08build-small-hackathon /agent-eval-golden-dataset Tech Interview Agent — Golden Eval Dataset Stop guessing whether your AI interviewer is good. Start measuring it. This dataset provides ground-truth benchmarks for evaluating AI agents that conduct tech job interviews. Each record is a structured test case: give it to your agent, collect the response, run it through the AI Agent Evaluation Pipeline, and get objective scores — no human review needed. Generated by NVIDIA Nemotron-3-Nano-30B-A3B. What's inside 40… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agent-eval-golden-dataset.text-generationn<1K0 likes8 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.