datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apex-agents
APEX–Agents
APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar).
Tasks: 480 total (160 per job category)
Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/mercor/apex-agents.gaia2
Gaia2
Paper | Code | Project Page
Dataset Summary
Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically.
The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.gaia2_filesystem
GAIA2 Filesystem
This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset.
Dataset Link
https://huggingface.co/datasets/meta-agents-research-environments/gaia2
Contact Details
Publishing POC: Meta AI Research Team
Affiliation: Meta Platforms, Inc.
Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.Nexus-Agents-ToolCalling
Nexus Agents — Tool-Calling Conversations
Synthetic, schema-verified tool-calling conversations for training the Nexus Projects
agents. This is the exact data behind
Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF),
including the verification transcripts that scored it (27/27 on the behavioral
interview eval, vs 13/27 for the base model).
Links: the fine-tuned model →
Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF) ·
the generator + seed data + eval harness →
Nexus Training Studio ·… See the full description on the dataset page: https://huggingface.co/datasets/NexusProjectsAI/Nexus-Agents-ToolCalling.trustworthy-biology-agents-traces
Trustworthy Biology Agents — Run Traces
Raw execution traces from 1,329 agent runs across three coding agents on three
biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed
trace bundle for the study in
manu-tej/ai-scientists; the write-up
lives in that repo's RESULTS.md.
The motivating question is not only whether an agent reaches the right answer, but
whether it behaves like a trustworthy analyst when the task is ambiguous,
under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.agents-last-exam
Agents Last Exam — Task Card Metadata (v1.1)
A metadata-only release (v1.1) of 152 tasks from the Agents Last Exam (ALE)
benchmark for evaluating computer-use agents on long-horizon professional work.
The Agents Last Exam dataset family
ALE is published as three companion HuggingFace datasets:
Dataset
Contents
Access
Task Card Metadata
One row per task: titles, prompts, taxonomy, input-file descriptors
Open
Task Input Data
The input/ files each task… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam.TIR-Bench
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
Introduction:
TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.AgentSearch-V1
Getting Started
The AgentSearch-V1 dataset boasts a comprehensive collection of over one billion embeddings, produced using jina-v2-base. The dataset encompasses more than 50 million high-quality documents and over 1 billion passages, covering a vast range of content from sources such as Arxiv, Wikipedia, Project Gutenberg, and includes carefully filtered Creative Commons (CC) data. Our team is dedicated to continuously expanding and enhancing this corpus to improve the search… See the full description on the dataset page: https://huggingface.co/datasets/SciPhi/AgentSearch-V1.agent-sft-stitch-zh-tts
agent-sft-stitch-zh-tts
Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted.
Configs
records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.wave-uiLICENSE
AgentSynth
AgentSynth
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
Paper | Project Page | Code
Abstract
We introduce AgentSynth, a scalable and cost-efficient pipeline for automatically synthesizing high-quality tasks and trajectory datasets for generalist computer-use agents. Leveraging information asymmetry, AgentSynth constructs subtasks that are simple during generation but significantly more challenging when composed into long-horizon… See the full description on the dataset page: https://huggingface.co/datasets/sunblaze-ucb/AgentSynth.wave-ui-25k
WaveUI-25k
This dataset contains 25k examples of labeled UI elements. It is a subset of a collection of ~80k preprocessed examples assembled from the following sources:
WebUI
RoboFlow
GroundUI-18K
These datasets were preprocessed to have matching schemas and to filter out unwanted examples, such as duplicated, overlapping and low-quality datapoints. We also filtered out many text elements which were not in the main scope of this work.
The WaveUI-25k dataset includes the original… See the full description on the dataset page: https://huggingface.co/datasets/agentsea/wave-ui-25k.apex-agents
APEX–Agents
APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar).
Tasks: 480 total (160 per job category)
Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/apex-agents.SWITCH
SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios
[arXiv]
[leaderboard]
[dataset]
[PDF]
Dataset Summary
SWITCH (Semantic World Interface Tasks for Control & Handling) is a multimodal embodied-interaction benchmark for understanding, modeling, and evaluating actions over Tangible Control Interfaces (TCIs) in egocentric real-world scenarios.
TCIs include everyday interfaces such as appliance panels, lighting… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-Agents/SWITCH.agent-sft-10B
Dataset: agent-sft-10B
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/agent-sft-10B/no-curriculum/tmp.
unit3-inviteeslectura-agents-data
LectūraAgents Dataset
Overview
This dataset is in support of findings in our paper LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching. LectūaAgents is a hierarchical multi-agent framework that enables end-to-end personalized learning experiences through adaptive embodied teaching. It mirrors a professor–students’ relationship, wherein a ProfessorAgent guides a collaborative team of specialized subordinate… See the full description on the dataset page: https://huggingface.co/datasets/Jaward/lectura-agents-data.apex-agents
APEX–Agents
APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar).
Tasks: 480 total (160 per job category)
Worlds: 33 total (10 banking, 11 consulting, 12 law)… See the full description on the dataset page: https://huggingface.co/datasets/neyralabs/apex-agents.apex-agents
APEX–Agents
APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar).
Tasks: 480 total (160 per job category)
Worlds: 33 total (10 banking, 11 consulting, 12 law)… See the full description on the dataset page: https://huggingface.co/datasets/abridges/apex-agents.agents-index
AgentCrush Agent Index
Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz.
Overview
1,441 agents indexed across categories: developer tools, tokenized agents, service agents, model families
207 evidence-ranked with verified multi-signal scores
Updated: 2026-09-22
Configs
Config
Description
Rows
agents
All indexed agents with metadata
~1,441
evidence_ranked
Evidence-ranked tier only
~207
snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.apex-agents
APEX–Agents
APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar).
Tasks: 480 total (160 per job category)
Worlds: 33 total (10 banking, 11 consulting, 12 law)… See the full description on the dataset page: https://huggingface.co/datasets/lobinni/apex-agents.tiny-agentsagent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.evovling_agents
Evolving Agents Benchmark
This repository contains the data presented in EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?.
Project page: https://mas-orchestra.salesforceresearch.ai/evoharness/
A versioned, per-split, multi-domain library of given Codex subagents,
produced by evovle_agents. It is the agent-track
analogue of evovling_tools: where
evovling_skills evaluates a model that generates
skills, evolving-agents evaluates a model that orchestrates given… See the full description on the dataset page: https://huggingface.co/datasets/ZixuanKe/evovling_agents.agent-simulations
Agent Simulations
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
53,971 synthetic agent trajectories generated by simulations
across 34 agent types. The rows include successful and failed
trajectories for supervised fine-tuning, preference work, reinforcement learning, and
evaluation.
NOTE: This is generated test and training data, not curated ground truth. Review and
filter it for your application before training or… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/agent-simulations.agent-sft
voidful/agent-sft
A model-agnostic agent / tool-use SFT dataset in a standard OpenAI-style schema —
train any model on it (Qwen, Llama, Gemma, GPT, …).
Built with the agentds toolkit:
per-source normalization -> group-level dedup (exact + SWE-provenance + MinHash near-dup)
-> heuristic quality stratification. The schema is wire-compatible with
voidful/gemma4-agent-sft
(this run also dedups against it), so the two concatenate cleanly.
Schema
field
type… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft.OpenJudge
OpenJudge Benchmark Dataset
Benchmark dataset for evaluating graders across text, multimodal, and agent scenarios. This dataset supports the OpenJudge framework with labeled preference pairs for quality-assured grader development.
Dataset Statistics
Evaluation Benchmarks
Category
Task
Files
Samples
🤖 Agent
12
166
action
1
8
memory
3
47
plan
1
7
reflection
3
52
tool
4
52
🖼️ Multimodal
4
80
image_coherence
1
20
image_editing… See the full description on the dataset page: https://huggingface.co/datasets/agentscope-ai/OpenJudge.gaia2-cli
GAIA2 CLI
Benchmark dataset for gaia2-cli, the CLI-based agent evaluation harness.
Schema
Each row has two columns:
Column
Type
Description
scenario_id
string
Unique scenario identifier (e.g. scenario_universe_21_1qgjj6)
scenario
string
Complete scenario as a JSON string
Usage
from datasets import load_dataset
import json
# Load a specific config (160 scenarios)
ds = load_dataset("meta-agents-research-environments/gaia2-cli", "adaptability"… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2-cli.AgentsNet
AgentsNet
This repository contains the graph instances used in the AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs paper.
AgentsNet is a new benchmark for multi-agent reasoning, designed to measure the ability of multi-agent systems to collaboratively form strategies for problem-solving, self-organization, and effective communication given a network topology. It draws inspiration from classical problems in distributed systems and graph theory.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AgentsNet.GroundUI-18K
GroundUI-18K
This dataset is the full GroundUI-18K in AgentStudio. Please note that this dataset is a test set rather than a training set. Therefore, please do not use it for training. More details are provided in the project page.
