datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fastly-agent-toolkit-evals
Fastly Agent Toolkit Evals
Evaluation dataset for the Fastly Agent Toolkit. Measures how well AI models complete Fastly-specific engineering tasks, with and without toolkit skills loaded.
What this dataset contains
Each entry is a full evaluation run: a task prompt, model configuration, the model's output, tool call traces, and grading results. The key comparison is with_skill (toolkit loaded) vs without_skill (no toolkit), across multiple models.… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/fastly-agent-toolkit-evals.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.arabic-agent-eval
Arabic Agent Eval — Dataset Card
An open, installable Arabic function-calling benchmark with dialect splits.
Dataset summary
51 evaluation items spanning 6 categories and 5 dialects of Arabic, testing whether large language models can (a) select the right tool, (b) extract arguments from natural Arabic instructions, (c) preserve Arabic text in tool arguments instead of transliterating, and (d) understand dialectal framing.
Supported tasks… See the full description on the dataset page: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval.Agent-eval-Effector-Hunt
Agent Eval: Effector Hunt
While AI scientist agents like Claude Science and Google's AI co-scientist highlight the potential of autonomous research, compact and reproducible datasets for evaluating these agents on real scientific workflows remain scarce.
Agent Eval: Effector Hunt is a genomics benchmark package designed around a real scientific discovery workflow from the Science paper Chen et al. 2017. It asks an AI agent, a computational biologist, or a hybrid human-agent… See the full description on the dataset page: https://huggingface.co/datasets/chjp0632/Agent-eval-Effector-Hunt.agent-eval-scenarios
Agent Eval Scenarios
Agent Eval Scenarios is a compact public dataset for lightweight evaluation of AI agents working on practical engineering and operations tasks.
It is designed to be:
small enough to inspect manually
structured enough to extend into a benchmark
grounded in real agent workflows such as code review, debugging, docs synthesis, security hardening, UI verification, and workflow automation
Files
data/agent_eval_scenarios.csv — labeled scenarios with… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/agent-eval-scenarios.mlops-agent-rl-eval
MLOps Agent RL — Evaluation Results
Offline multi-turn tool-calling evaluation on the mocked MLOpsEnv
incident suite (6 scenarios).
Models
base: Qwen/Qwen2.5-0.5B-Instruct
grpo: vanishingradient/mlops-agent-rl-qwen05b-grpo (LoRA adapter)
oracle: scripted expected tools + preferred remediation
Files
episodes.jsonl — full trajectories + metrics
episodes_slim.jsonl — metrics without trajectories
summary.json — aggregates for the paper
config.json —… See the full description on the dataset page: https://huggingface.co/datasets/vanishingradient/mlops-agent-rl-eval.eval-agent-traceEval Agent Trace of a MLE Agent by Celestra. Sythetically Generated by gpt 5.2 thinking
agent-eval-golden-dataset
Tech Interview Agent — Golden Eval Dataset
Stop guessing whether your AI interviewer is good. Start measuring it.
This dataset provides ground-truth benchmarks for evaluating AI agents that conduct tech job interviews. Each record is a structured test case: give it to your agent, collect the response, run it through the AI Agent Evaluation Pipeline, and get objective scores — no human review needed.
Generated by NVIDIA Nemotron-3-Nano-30B-A3B.
What's inside
40… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agent-eval-golden-dataset.
