datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TIR-Bench
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
Introduction:
TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.agents-index
AgentCrush Agent Index
Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz.
Overview
1,443 agents indexed across categories: developer tools, tokenized agents, service agents, model families
207 evidence-ranked with verified multi-signal scores
Updated: 2026-09-24
Configs
Config
Description
Rows
agents
All indexed agents with metadata
~1,443
evidence_ranked
Evidence-ranked tier only
~207
snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.OpenJudge
OpenJudge Benchmark Dataset
Benchmark dataset for evaluating graders across text, multimodal, and agent scenarios. This dataset supports the OpenJudge framework with labeled preference pairs for quality-assured grader development.
Dataset Statistics
Evaluation Benchmarks
Category
Task
Files
Samples
🤖 Agent
12
166
action
1
8
memory
3
47
plan
1
7
reflection
3
52
tool
4
52
🖼️ Multimodal
4
80
image_coherence
1
20
image_editing… See the full description on the dataset page: https://huggingface.co/datasets/agentscope-ai/OpenJudge.CriticBench
Dataset Card for Dataset Name
CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and algorithmic. It compiles 15 datasets and incorporates responses from three LLM families.
Dataset Details
Dataset Description
Curated by: THU
Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/llm-agents/CriticBench.Multilingual-Representative-Agents-Dataset
We have developed chat data containing over 2,300 files for support representatives responding to customers in different fields, with a total of 10 languages, 230 files for each language. Introducing v1.
Each file contains over 10 conversations. These data have been created in strict compliance with privacy principles, using random human names and random company names.
Supported languages: en:hi:zh:ar:ru:uk:es:tr:fr:az
The format is Customer:Agent, and as specified in the metadata… See the full description on the dataset page: https://huggingface.co/datasets/nebulonsai/Multilingual-Representative-Agents-Dataset.AgentSLR
AgentSLR: Priority Pathogens Dataset
Paper
Codebase
Project Website
This dataset accompanies the paper Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews. It provides the data component of the AgentSLR evaluation harness. This covers article metadata, human abstract and full text screening labels, and structured human data extractions for epidemiological parameters… See the full description on the dataset page: https://huggingface.co/datasets/OxRML/AgentSLR.agentshield-bench
AgentShield-Bench
Structured benchmark dataset for evaluating security of AI agents in tool-calling and MCP environments.
Overview
AgentShield-Bench provides 720 scenarios spanning 11 adversarial attack categories plus benign control cases. Each scenario includes trusted instructions, untrusted content, available tools, canary secrets, expected safe behaviors, and attack success conditions.
Schema
Field
Type
Description
scenario_id
string… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/agentshield-bench.omnimcp_agentselfheal_mcp_pro_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_agentselfheal_mcp_pro_teaser.agent-session-handoff
Agent Session Handoff
Synthetic operations-agent transcripts for studying knowledge retention across a model swap: does a compressed
KV-cache memory keep more of a session than a text summary when the next model takes over?
Each episode is built from a sampled fact record (service owners, ports, branches, config values, ticket states,
decisions) rendered by an LLM writer into a realistic user / assistant / tool-output session. The writer never sees the
questions. A swap point… See the full description on the dataset page: https://huggingface.co/datasets/Narmeen07/agent-session-handoff.ai-agents-fr
Agents IA - Dataset Francais
Dataset bilingue complet sur les Agents IA, les Frameworks Multi-Agents et le Model Context Protocol (MCP).
Contenu du Dataset
Categorie
Nombre d'entrees
Description
Architectures d'Agents
15
ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchique, Swarm, etc.
Frameworks
12
CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc.
MCP & Tool Use
15
Architecture MCP, transports, function calling, securite… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-fr.agentsim-atc
AgentSim Agent-Trace Corpus (ATC)
Grounded reasoning traces of retrieval-augmented question-answering agents,
generated by the AgentSim platform. 103,567 reasoning steps spanning
three established IR benchmarks (Quasar-T 38,915 + CausalQA 36,192 +
MSMARCO 28,460), with 20,548 supervised query-document-answer triples
extracted for fine-tuning, and 199,968 unique retrieved documents.
Every reasoning step traces back to specific documents in the source
corpus, enabling step-level… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/agentsim-atc.ai-agents-en
AI Agents - English Dataset
Comprehensive bilingual dataset on AI Agents, Multi-Agent Frameworks, and the Model Context Protocol (MCP).
Dataset Contents
Category
Entry Count
Description
Agent Architectures
15
ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchical, Swarm, etc.
Frameworks
12
CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc.
MCP & Tool Use
15
MCP architecture, transports, function calling, security, ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-en.when-agents-act
Dataset Card for "When Agents Act"
Dataset Summary
This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute).
Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.agent-safety-bench-zh
agent-safety-bench-zh · 中文 Agent 工具调用风险 / 提示注入 评测基准
评测一个 LLM / 安全护栏能否正确识别 agent 拟执行动作中的危险操作与被注入的恶意指令,并给出 allow / block 决策与风险分级。纯防御 / 安全教育 / 评测研究。
当前 v0.2,105 条(test split),三类:benign(应 allow) / prompt_injection(应 block) / destructive(应 block),含 easy/medium/hard。
每条字段:id, category, difficulty, context, action, gold{decision,risk}, rationale, tags。
加载
from datasets import load_dataset
ds = load_dataset("uninhibited-scholar/agent-safety-bench-zh", split="test")… See the full description on the dataset page: https://huggingface.co/datasets/uninhibited-scholar/agent-safety-bench-zh.agentsim-atc-multihop
AgentSim Agent-Trace Corpus — Multi-hop
A multi-hop sibling of the AgentSim Agent-Trace Corpus
(agentsim-atc)
with an evolved schema designed for student model distillation.
1 490 accepted SFT trajectories plus 2 980 step-level DPO preference
pairs, generated over 5 multi-hop QA datasets through a 7-action agentic
executor with an Always-Search Policy filter.
This corpus accompanies a follow-up technical report to "AgentSim: A
Platform for Verifiable Agent-Trace Simulation"… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/agentsim-atc-multihop.gene-llm-agents-instruct
llm-agents-instruct v116
Auto-built (demand): 1 open request(s) and 0 recent download(s) for 'llm-agents' with no dataset newer than 14 days
Kind: synthetic
Domain: llm-agents
Records: 1000
Created: 2026-07-08T17:36:15+00:00
SHA-256: 281020e4a1db9e063ea6eaf359b69cfa40a89f13faeae521a4179cec586fc10c
Pipeline: v2.0.0
Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7}
Generated by: Qwen3-4B-Instruct-2507-Q4_K_M.gguf (backend:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-llm-agents-instruct.gene-llm-agents-corpus
llm-agents-corpus v92
Auto-built (demand): 1 open request(s) and 0 recent download(s) for 'llm-agents' with no dataset newer than 14 days
Kind: scraped
Domain: llm-agents
Records: 702
Created: 2026-07-08T17:36:14+00:00
SHA-256: 74c00af747e66d1d4abf248168263fb9d5f1e424182d7adcdef160533c32dae2
Pipeline: v2.0.0
Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": null, "min_judge": null}
Sources
huggingface: 301
papers: 197
arxiv: 145
github:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-llm-agents-corpus.MetaQA_Agents_Predictions
Dataset Card for MetaQA Agents' Predictions
Dataset Summary
This dataset contains the answer predictions of the QA agents for the QA datasets used in MetaQA paper. In particular, it contains the following QA agents' predictions:
Span-Extraction Agents
Agent: Span-BERT Large (Joshi et al.,2020) trained on SQuAD. Predictions for:
SQuAD
NewsQA
HotpotQA
SearchQA
Natural Questions
TriviaQA-web
QAMR
DuoRC
DROP
Agent: Span-BERT Large (Joshi et al.,2020) trained on… See the full description on the dataset page: https://huggingface.co/datasets/haritzpuerto/MetaQA_Agents_Predictions.scientific-agent-protocol-traces
SciAgentTrace
Matched cross-domain dataset of scientific-agent protocols. The central
comparison contains the same 6,653 problems under two actor models and four
protocols: 53,224 trajectories in 40 complete model--benchmark--protocol
groups. The broader table-first package contains 68,892
trajectories. Begin with trajectories, outcomes, or matched_outcomes, then
follow stable identifiers to messages and compressed raw traces.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/scientific-agent-protocol-traces.EMNLP_Cost-Aware-Protocol-Routing
Cost-Aware Protocol Routing: Matched Protocol Outcomes
The short version. We ran the same 6,803 reasoning problems through four
different LLM collaboration setups — from a single direct answer up to a
four-agent deliberation — and recorded, for every problem, which ones got it
right. Then we asked whether a model can look at a problem beforehand and
predict which setup is worth paying for.
It can predict whether it will fail. It cannot predict which collaboration
protocol will… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing.AAAI_Wrong-but-Useful
Wrong but Useful — Trajectory Value Dataset
Companion dataset to Wrong but Useful: Trajectory Value Beyond Answer
Correctness in Multi-Agent Messages
(arXiv:2608.14375 ·
PDF) ·
Code ·
Project page
READ BEFORE USE — licensing and do-not-train terms
This dataset redistributes third-party benchmark content (question text and
gold answers) alongside model-generated text and the trajectory-value
measurements that are this paper's contribution. The benchmark content… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/AAAI_Wrong-but-Useful.NeurIPS26_Precise-but-Uncoupled
NeurIPS26_Precise-but-Uncoupled
⚠️ LICENCE — READ BEFORE USE. TWO RESTRICTIONS PASS DOWNSTREAM TO YOU.
This dataset is license: other — a COMPOSITE. There is no single permissive licence.
NONCOMMERCIAL. The MaScQA portion (642 problems) is CC-BY-NC-SA-4.0.
commercial_use_policy = not_allowed. Commercial use of that portion is not permitted.
SHAREALIKE + DO-NOT-TRAIN. The LAB-Bench portion (741 problems) is CC-BY-SA-4.0 and
carries the upstream contamination… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/NeurIPS26_Precise-but-Uncoupled.IEEE2026_BigData_MAS-4-Science-Matching
SciAgentTrace
An execution-layer trace resource for scientific-agent workload characterization.
A protocol fixes who reasons, what each role can see, when feedback returns, and
when a workflow stops. Those choices determine the sequence of model requests
that produces an answer, so protocol design is also workload design. Two
workflows that consume similar token totals can issue very different request
sequences. SciAgentTrace records that difference.
The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.
