datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-trajectory-sentinel
AgentTrajectorySentinel — 3581 agent episodes across 33 corpora
Committed agent trajectories with step-level telemetry, used to fit and
evaluate one-class monitors for real-time failure detection.
Paper: https://arxiv.org/abs/2608.02464
Code and the full evaluation harness:
https://github.com/sunnydubey1111/agent-trajectory-sentinel
A recorded walkthrough of the method, ending with the live demo
detecting and repairing a real failure:
https://youtu.be/a05n_000klE?t=0… See the full description on the dataset page: https://huggingface.co/datasets/sunnydubey1111/agent-trajectory-sentinel.mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.deepseek-v4-pro-agent-tool-calling-trajectory
DeepSeek V4 Pro ToolScale Agent SFT Dataset
A curated subset of multi-turn tool-calling trajectories generated by DeepSeek V4 Pro on ToolScale. The dataset is filtered by action-match score against ground-truth trajectories and is designed for supervised fine-tuning of agentic models on realistic, multi-step tool use.
Each conversation includes natural-language user requests, tool calls, tool observations, assistant reasoning traces, and grounded final responses across five… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/deepseek-v4-pro-agent-tool-calling-trajectory.Pantheon-Agent-Trajectory
🏛️ Pantheon Agent Trajectory Gallery
Curated end-to-end agent runs from PantheonOS — an open multi-agent framework for scientific computing.
Each "trajectory" captures a complete chat session: the user prompt, every reasoning/tool step the agent(s) took, the code that was run, the figures that were produced, and the final report. Trajectories are fully inspectable and reproducible, designed for transparency, teaching, and benchmarking.
🔗 Browse the gallery (live):… See the full description on the dataset page: https://huggingface.co/datasets/NaNg/Pantheon-Agent-Trajectory.Qwen-3.6-plus-agent-tool-calling-trajectory
Qwen 3.6 Plus: ToolScale Agent SFT Dataset
Multi-turn tool-calling trajectories generated by Qwen 3.6 Plus via OpenRouter on ToolScale. Both passing and near-passing rollouts are included, allowing users to choose their own quality threshold using reward and score.
Each row is a flattened conversation prefix ending at one assistant turn, ready for next-token SFT. Assistant turns include a reasoning_content field containing the model’s reasoning.
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen-3.6-plus-agent-tool-calling-trajectory.A-Survey-for-LLM-Agent-Trajectory-Analysis
A Survey for LLM Agent Trajectory Analysis
This dataset repository hosts the survey paper A Survey for LLM Agent Trajectory Analysis: From Failure Attribution to Enhancement and a structured metadata snapshot of the companion paper collection from Awesome-LLM-Agent-Trajectory-Analysis.
The repository is intended for discovery, citation, and lightweight analysis of the literature around LLM agent trajectory analysis, including failure attribution, trajectory-based debugging… See the full description on the dataset page: https://huggingface.co/datasets/RobinChen2001/A-Survey-for-LLM-Agent-Trajectory-Analysis.mcp-agent-trajectory-benchmark
⚡ Model Context Protocol (MCP) & Advanced Tool‑Use Alignment Tiers
15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%).
Schema Validation Summary
Programmatic validation of this exact trial file - reproducible from data.jsonl.
What this trial verifies — use these 50 rows to confirm, on your own stack:
Schema integrity (strict JSONL, matches the published schema)
Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.wearable-agent-trajectory-annotations
Wearable Agent Trajectory Annotation Dataset
Dataset Summary
50 wearable agent trajectories annotated by 5 LLM-simulated annotator personas
using the agenteval-schema-v1 JSON schema, across two calibration phases
(500 annotation records total). Designed to benchmark annotation-quality pipelines
for agentic AI systems.
Each trajectory captures a wearable AI agent responding to a real-time sensor event
(health alert, privacy-sensitive context, location trigger… See the full description on the dataset page: https://huggingface.co/datasets/finaspirant/wearable-agent-trajectory-annotations.Agent-Trajectory-Dataset
Description
본 데이터셋은 심층 검색, 데이터 분석, 산업 리서치 등 사무 환경에서 수행되는 다양한 작업 시나리오를 포함하며, 완전한 멀티턴 추론 과정과 도구 호출 체인으로 구성되어 있습니다. 에이전트의 계획 수립 능력 분석, 도구 선택 전략 연구 및 작업 품질 평가를 지원하도록 설계되었으며, 에이전트 학습 및 평가를 위한 구조화된 벤치마크로 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/2185?source=hf.kr
Specifications
Data content
OpenClaw를 통해 생성된 에이전트 트래젝토리 데이터
Category
심층 검색, 데이터 분석, 산업 리서치
Data volume
5,300
Model… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/Agent-Trajectory-Dataset.Agent-Trajectory-Data-Sample
Agent-Trajectory-Dataset
Description
This dataset covers office-based scenarios such as in-depth searches, data analysis, and industry research, encompassing complete multi-turn reasoning trajectories and tool-calling chains. It is designed to support the analysis of agent planning capabilities, research into tool selection strategies, and quality assessment, providing a structured benchmark for agent training and evaluation.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/Agent-Trajectory-Data-Sample.agent-trajectory-eval-datasetagenticml-agent-trajectory-dataset
Telos Agent Trajectory Dataset
Synthetic multi-turn agent trajectories for training LLMs to emit structured tool-use loops. Each trajectory is provided in two parallel representations: the Telos frame format and an equivalent ChatML+tools rendering of the same behavior.
This is synthetic data distilled from Qwen3.5 Plus (2026-04-20) via OpenRouter. Users who require data not derived from a specific provider should factor this into licensing and downstream-use decisions before… See the full description on the dataset page: https://huggingface.co/datasets/kosiasuzu/agenticml-agent-trajectory-dataset.agent_trajectory_reviewsAgent-Trajectory-2.8kembodied_web_agent_outdoor_trajectoryagent_trajectory_testweb-agent-trajectory-testweb-agent-trajectory-multimodal-testtrajectory
