Agent
Datasets
All datasets matching “Agent”course-imagesapex-agents
APEX–Agents
APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar).
Tasks: 480 total (160 per job category)
Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/mercor/apex-agents.resultsDeepScaleR-Preview-Dataset
Data
Our training dataset consists of approximately 40,000 unique mathematics problem-answer pairs compiled from:
AIME (American Invitational Mathematics Examination) problems (1984-2023)
AMC (American Mathematics Competition) problems (prior to 2023)
Omni-MATH dataset
Still dataset
Format
Each row in the JSON dataset contains:
problem: The mathematical question text, formatted with LaTeX notation.
solution: Offical solution to the problem, including LaTeX formatting… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset.gaia2
Gaia2
Paper | Code | Project Page
Dataset Summary
Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically.
The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.agents-last-exam-data
Agents Last Exam — Task Input Data
Input files (the materials each task hands to the agent at run start) for the
Agents Last Exam (ALE) benchmark. Browsable per-task directory layout.
The Agents Last Exam dataset family
ALE is published as three companion HuggingFace datasets:
Dataset
Contents
Access
Task Card Metadata
One row per task: titles, prompts, taxonomy, input-file descriptors
Open
Task Input Data
The input/ files each task hands the agent at… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data.
tessLong-context reader. Turns forty tabs into one page you actually finish.
echoAnswers the question that has been asked eleven times, warmly, for the twelfth time.
scoutFinds the three projects solving your problem before you finish describing it.
chiefKeeps six agents from doing the same task twice. Mostly by asking first.
fernReads the tree, not the screenshot. Every PR gets one honest question.
pebbleLabels, dedupes and closes with an actual explanation.
noriFirst reply within the hour, and it is never a copy-paste.