datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-trajectory-sentinel
AgentTrajectorySentinel — 3581 agent episodes across 33 corpora
Committed agent trajectories with step-level telemetry, used to fit and
evaluate one-class monitors for real-time failure detection.
Paper: https://arxiv.org/abs/2608.02464
Code and the full evaluation harness:
https://github.com/sunnydubey1111/agent-trajectory-sentinel
A recorded walkthrough of the method, ending with the live demo
detecting and repairing a real failure:
https://youtu.be/a05n_000klE?t=0… See the full description on the dataset page: https://huggingface.co/datasets/sunnydubey1111/agent-trajectory-sentinel.deepseek-v4-pro-agent-tool-calling-trajectory
DeepSeek V4 Pro ToolScale Agent SFT Dataset
A curated subset of multi-turn tool-calling trajectories generated by DeepSeek V4 Pro on ToolScale. The dataset is filtered by action-match score against ground-truth trajectories and is designed for supervised fine-tuning of agentic models on realistic, multi-step tool use.
Each conversation includes natural-language user requests, tool calls, tool observations, assistant reasoning traces, and grounded final responses across five… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/deepseek-v4-pro-agent-tool-calling-trajectory.Qwen-3.6-plus-agent-tool-calling-trajectory
Qwen 3.6 Plus: ToolScale Agent SFT Dataset
Multi-turn tool-calling trajectories generated by Qwen 3.6 Plus via OpenRouter on ToolScale. Both passing and near-passing rollouts are included, allowing users to choose their own quality threshold using reward and score.
Each row is a flattened conversation prefix ending at one assistant turn, ready for next-token SFT. Assistant turns include a reasoning_content field containing the model’s reasoning.
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen-3.6-plus-agent-tool-calling-trajectory.wearable-agent-trajectory-annotations
Wearable Agent Trajectory Annotation Dataset
Dataset Summary
50 wearable agent trajectories annotated by 5 LLM-simulated annotator personas
using the agenteval-schema-v1 JSON schema, across two calibration phases
(500 annotation records total). Designed to benchmark annotation-quality pipelines
for agentic AI systems.
Each trajectory captures a wearable AI agent responding to a real-time sensor event
(health alert, privacy-sensitive context, location trigger… See the full description on the dataset page: https://huggingface.co/datasets/finaspirant/wearable-agent-trajectory-annotations.agenticml-agent-trajectory-dataset
Telos Agent Trajectory Dataset
Synthetic multi-turn agent trajectories for training LLMs to emit structured tool-use loops. Each trajectory is provided in two parallel representations: the Telos frame format and an equivalent ChatML+tools rendering of the same behavior.
This is synthetic data distilled from Qwen3.5 Plus (2026-04-20) via OpenRouter. Users who require data not derived from a specific provider should factor this into licensing and downstream-use decisions before… See the full description on the dataset page: https://huggingface.co/datasets/kosiasuzu/agenticml-agent-trajectory-dataset.kvine-agent-trajectory-bench
KVine agent-trajectory benchmark & resident-set policy training data 🌿
Two things from the KVine research prototype
(branch-aware KV cache for agent trajectories):
policy_train — self-supervised training data for the resident-set policy:
structural features of branches in KVine's trajectory tree, labelled with
whether the branch was revisited within the next 5 steps.
benchmark_results.json — the full A / B / C benchmark output (cumulative
prefill FLOPs, TTFT, recomputed tokens… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/kvine-agent-trajectory-bench.agent-trajectory-eval-datasetagent_trajectory_testweb-agent-trajectory-testweb-agent-trajectory-multimodal-testaic_sfp2nic_agent_expert_trajectory_sample_nominalrecovery_successaic_sfp2nic_agent_expert_trajectory_sample_debug_cheatcode_modified_attempt_15aic_sfp2nic_agent_expert_trajectory_sample_debug_cheatcode_modified_attempt_20aic_sfp2nic_agent_expert_trajectory_sample_vlm_moveit_nominalrecovery_attempt_1_score_94p98aic_sfp2nic_agent_expert_trajectory_sample_vlm_moveit_nominalrecovery_attempt_1_score_97p05aic_sfp2nic_agent_expert_trajectory_sample_vlm_moveit_nominal_attempt_1_score_96p83aic_sfp2nic_agent_expert_trajectory_sample_vlm_moveit_recovery_attempt_3aic_sfp2nic_agent_expert_trajectory_sample_agent_nominalaic_sfp2nic_agent_expert_trajectory_sample_nominalrecovery_attempt_2aic_sfp2nic_agent_expert_trajectory_sample_vlm_moveit_recovery_attempt_2aic_sfp2nic_agent_expert_trajectory_sample_vlm_moveit_recovery_attempt_4aic_sfp2nic_agent_expert_trajectory_sample_vlm_moveit_recovery_attempt_1
