agent-trajectory
qwen2.5-7b-agent-trajectory-lora-107b-i1-GGUFqwen2.5-7b-agent-trajectory-lora-107b-GGUFqwen2.5-7b-agent-trajectory-lora-106b-GGUFqwen3-4b-agent-trajectory-lora.0228qwen3-4b-agent-trajectory-merged3-sftdpo_v8qwen3-4b-agent-trajectory-merged11qwen2.5-7b-agent-trajectory-mixed_dbv4_alfv4_1to1qwen3-4b-agent-trajectory-merged11-sftdpo_v14
agent-trajectory-sentinel
AgentTrajectorySentinel — 3581 agent episodes across 33 corpora
Committed agent trajectories with step-level telemetry, used to fit and
evaluate one-class monitors for real-time failure detection.
Paper: https://arxiv.org/abs/2608.02464
Code and the full evaluation harness:
https://github.com/sunnydubey1111/agent-trajectory-sentinel
A recorded walkthrough of the method, ending with the live demo
detecting and repairing a real failure:
https://youtu.be/a05n_000klE?t=0… See the full description on the dataset page: https://huggingface.co/datasets/sunnydubey1111/agent-trajectory-sentinel.mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.failed_agent_trajectory
Mobile Trajectory Verification Data
This dataset is prepared for verifying how LoT agent can help us identify incorrect actions taken in mobile execution trajectory. As mentioned in our paper (see citation 1), We select 52 mobile agent trajectories with specific task goals from instruction-
guided executions on MagicWand platform(see citation 2). Each folder has a execution trajectory that includes screenshots along this trajectories, json files describing the details of each… See the full description on the dataset page: https://huggingface.co/datasets/orlando23/failed_agent_trajectory.deepseek-v4-pro-agent-tool-calling-trajectory
DeepSeek V4 Pro ToolScale Agent SFT Dataset
A curated subset of multi-turn tool-calling trajectories generated by DeepSeek V4 Pro on ToolScale. The dataset is filtered by action-match score against ground-truth trajectories and is designed for supervised fine-tuning of agentic models on realistic, multi-step tool use.
Each conversation includes natural-language user requests, tool calls, tool observations, assistant reasoning traces, and grounded final responses across five… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/deepseek-v4-pro-agent-tool-calling-trajectory.Qwen-3.6-plus-agent-tool-calling-trajectory
Qwen 3.6 Plus: ToolScale Agent SFT Dataset
Multi-turn tool-calling trajectories generated by Qwen 3.6 Plus via OpenRouter on ToolScale. Both passing and near-passing rollouts are included, allowing users to choose their own quality threshold using reward and score.
Each row is a flattened conversation prefix ending at one assistant turn, ready for next-token SFT. Assistant turns include a reasoning_content field containing the model’s reasoning.
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen-3.6-plus-agent-tool-calling-trajectory.eto-sft-trajectory
Expert Trajectories for ETO
🌐 Homepage | 🐍 GitHub | 📖 arXiv
Expert trajectories for Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
Authors: Yifan Song, Da Yin, Xiang Yue, Jie Huang, Sujian Li, Bill Yuchen Lin.
We introduce ETO (Exploration-based Trajectory Optimization), an agent learning framework inspired by "trial and error" process of human learning.
ETO allows an LLM agent to iteratively collect failure trajectories and updates its policy by… See the full description on the dataset page: https://huggingface.co/datasets/agent-eto/eto-sft-trajectory.
