datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-llm-traces-v2
Exgentic Agent LLM Traces v2 — Agent Chat Only
OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it.
This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.agent-think-tool_use
Agent Think Tool Use
Датасет многошаговых агентных сессий для дообучения моделей работе с кодом, инструментами и инженерными задачами. Записи содержат пользовательские требования, комментарии агента во время работы, decision summaries, вызовы инструментов, результаты запусков, обработку ошибок и финальную проверку.
Каждый shard представляет отдельную связанную сессию, а не отдельный вопрос и ответ. Данные охватывают исследование задачи, работу с документацией, проектирование… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/agent-think-tool_use.DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖
Self-Taught Agentic Long Context Understanding (Arxiv).
AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass.
Installation Requirements
This codebase is largely based on OpenRLHF and Helmet, kudos to them.
The requirements are the same
pip install openrlhf
pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.agents-index
AgentCrush Agent Index
Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz.
Overview
1,443 agents indexed across categories: developer tools, tokenized agents, service agents, model families
207 evidence-ranked with verified multi-signal scores
Updated: 2026-09-24
Configs
Config
Description
Rows
agents
All indexed agents with metadata
~1,443
evidence_ranked
Evidence-ranked tier only
~207
snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.AgentJudgeBench
AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling
A benchmark for systematically evaluating how reliably LLM judges assess
agentic tool-calling workflows across structured, dependency-driven tasks.
Why this benchmark?
AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.data-agent
📈 Data Agent
Data-analysis tasks as a plain, load-and-go dataset — no runtime, no framework required. Each
row is one self-contained task: a real tabular dataset, a question about it, and a
deterministically-checkable gold answer. Load it, prompt any model however you like, and grade the
result with the bundled grader.
Where it comes from
Built from the jupyter-agent dataset
— real data-science notebooks over Kaggle datasets. Every question–answer pair was… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent.AgentProcessBench
AgentProcessBench
AgentProcessBench is a benchmark for process-level evaluation of tool-using agents. Each example is a full agent trajectory with multi-turn messages, tool definitions, tool-use traces, reference outputs, and step-wise process labels.
The benchmark contains 1,000 trajectories in total, with 250 examples from each subset:
bfcl
gaia_dev
hotpotqa
tau2
arxiv.org/abs/2603.14465
Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.AgentVQA
AgentVQA: A Multi-Domain Visual Question Answering Dataset
AgentVQA is a comprehensive dataset for training and evaluating visual agents across multiple domains, including GUI interaction, spatial reasoning, game understanding, robot manipulation, and video perception. The dataset contains multiple-choice questions based on screenshots, images, and videos.
Dataset Structure
The dataset is organized into five main domains:
1. Web Agents Domain (GUI Interaction)… See the full description on the dataset page: https://huggingface.co/datasets/AgentVQA/AgentVQA.Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.Agent-ValueBench
Agent-ValueBench
Paper | Project Page | GitHub
Agent-ValueBench is the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
Repository Structure
README.md
data/
cases.jsonl
rubrics.jsonl
environments.jsonl
raw/
case/
rubric/
environment/
Data Files… See the full description on the dataset page: https://huggingface.co/datasets/Value4AI/Agent-ValueBench.tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified
ToolACE - Tool-Use Agent Data Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.PortBench-QA
PortBench QA Dataset
Dataset Description
6,269 structured question-answer pairs probing correlation-based financial reasoning for multi-asset portfolio management, generated from the PortBench Market Base Dataset.
Task Templates
Template
Task
Complexity
Pairs
T1
Return prediction — direction for next N days
1 (single asset)
1,000
T2
Risk assessment — VaR at given confidence level
1
1,000
T3
Position sizing — given max drawdown… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-QA.crypto-agent-safe-function-calling
CrAI-SafeFuncCall Dataset
📄 Paper: Real AI Agents with Fake Memories: Fatal Context Manipulation
Attacks on Web3 Agents
🤗 Dataset: CrAI-SafeFuncCall
📊 Benchmark: CrAI-Bench
Overview
The CrAI-SafeFuncCall dataset is designed to enhance the security of AI agents when performing function calls in the high-stakes domain of cryptocurrency and financial applications. It focuses on the critical challenge of detecting and mitigating memory injection attacks. Derived from the… See the full description on the dataset page: https://huggingface.co/datasets/SentientAGI/crypto-agent-safe-function-calling.AgentSLR
AgentSLR: Priority Pathogens Dataset
Paper
Codebase
Project Website
This dataset accompanies the paper Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews. It provides the data component of the AgentSLR evaluation harness. This covers article metadata, human abstract and full text screening labels, and structured human data extractions for epidemiological parameters… See the full description on the dataset page: https://huggingface.co/datasets/OxRML/AgentSLR.AgentForge-1152
AgentForge-1152: Tool-Verified General + Cyber Reasoning
1,152 evidence-grounded agent trajectories. 144 disjoint task families. A clean 50/50 split between general reasoning and defensive cybersecurity.
AgentForge-1152 is a model-agnostic English corpus for supervised fine-tuning and evaluating assistants that use tools, recover from failed checks, and ground final answers in captured evidence. It uses portable messages and JSON function definitions rather than a model-specific… See the full description on the dataset page: https://huggingface.co/datasets/0xKitkat/AgentForge-1152.agent-financial-interactions
Purple Flea Agent Financial Interactions
A dataset of API interactions for autonomous AI agents using financial infrastructure (Purple Flea). Useful for training agents that handle crypto, trading, gambling, domain registration, escrow, and free onboarding via faucet.
Research
This dataset is referenced in:
"Purple Flea: A Multi-Agent Financial Infrastructure Protocol for Autonomous AI Systems"
https://doi.org/10.5281/zenodo.18808440
The paper covers the economic model… See the full description on the dataset page: https://huggingface.co/datasets/PurpleFlea/agent-financial-interactions.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/deepseek-v4-pro-0813-agentic.wearable-agent-trajectory-annotations
Wearable Agent Trajectory Annotation Dataset
Dataset Summary
50 wearable agent trajectories annotated by 5 LLM-simulated annotator personas
using the agenteval-schema-v1 JSON schema, across two calibration phases
(500 annotation records total). Designed to benchmark annotation-quality pipelines
for agentic AI systems.
Each trajectory captures a wearable AI agent responding to a real-time sensor event
(health alert, privacy-sensitive context, location trigger… See the full description on the dataset page: https://huggingface.co/datasets/finaspirant/wearable-agent-trajectory-annotations.PyFi-600K
Dataset Card for PyFi-600K
This dataset card aims to be a introduction for PyFi-600K, A financial VLM dataset containing 600K question-answer pairs generated via Adversarial agents.
AgenticFinLab/PyFi-600K/
├── README.md # Dataset documentation and description
├── images.zip # Compressed image files
├── PyFi-600K-dataset.csv # Q&A pairs in CSV format
├── PyFi-600K-dataset.json # Q&A pairs in JSON format
├── PyFi-600K-chain-dataset.json # Chain of Thought Q&A pairs dataset
└──… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PyFi-600K.sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.EMNLP_Cost-Aware-Protocol-Routing
Cost-Aware Protocol Routing: Matched Protocol Outcomes
The short version. We ran the same 6,803 reasoning problems through four
different LLM collaboration setups — from a single direct answer up to a
four-agent deliberation — and recorded, for every problem, which ones got it
right. Then we asked whether a model can look at a problem beforehand and
predict which setup is worth paying for.
It can predict whether it will fail. It cannot predict which collaboration
protocol will… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing.eu-tenders-with-questions-for-agentic-checklist-filling
eu-tenders-with-questions-for-agentic-checklist-filling
Dataset Description
This dataset contains questions and answers for evaluating Retrieval-Augmented Generation (RAG) systems in the context of generative agentic checklist-filling. The dataset is designed to benchmark various RAG architectures (Hybrid RAG, Graph RAG, Multi-Hop/Agentic RAG) on document analysis tasks.
Dataset Summary
Total Questions: 97
Document Families: 7
Languages: EN
Domain: Procurement… See the full description on the dataset page: https://huggingface.co/datasets/tmskss/eu-tenders-with-questions-for-agentic-checklist-filling.agentic-publication-protocol-dataset
APP compare-app benchmark
Paired reader conversations and blinded evaluations comparing an Agentic
Publication Protocol (APP) paper agent against a general repository-aware
agent, on 11 quantum-physics papers.
For each paper, a neutral reader asks the same scripted questions to both agents;
the two transcripts are anonymized and scored by a blinded evaluator on
accuracy, informativeness, grounding, and honesty (1-10).
Evaluator: Codex CLI, gpt-5.5, reasoning effort xhigh… See the full description on the dataset page: https://huggingface.co/datasets/phynics/agentic-publication-protocol-dataset.data_agent
data_agent
Plain, Harbor-free version of the data-analysis agent tasks — usable directly via load_dataset.
Splits: train 5000, test 250, eval 144. Deterministic grading, no LLM judge.
Columns
task_id, source_row_id — ids
question — the question to answer
answer — gold answer; reward_mode (numeric/exact_short/exact_bool/list/list_csv/flexible), atol/rtol — how to grade
difficulty_level (1-5), difficulty_tier (easy/medium/hard)
kaggle_dataset — source Kaggle… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent.
