datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-tool-use-trajectories
Agent Tool Use Trajectories (10K) 🚀
Dataset Description
This dataset contains 10,000 highly complex, multi-step dialogue trajectories designed to train open-source Large Language Models (LLMs) in advanced Agent Tool Use, Function Calling, and Reasoning.
Curated with professional AI training and data annotation standards, this dataset moves beyond simple synthetic Q&A pairs. It strictly follows the ChatML format and focuses heavily on multi-tool orchestration… See the full description on the dataset page: https://huggingface.co/datasets/Toprak1yu/agent-tool-use-trajectories.Agent-Tool-Use-Dialogue-Open-Dataset
Open Agent Tool Use Dialogue Dataset : High Quality AI Agent | Tool Use & Function Calls | Reinforcement Learning Datasets
Github|Huggingface|Pypi | Open Source AI Agent Marketplace DeepNLP|Agent RL Dataset | Agent MCP SubDomain Deployment | AI Agent A2Z
News
Multi-Turn Dialogue Data updated to 2026 Jan
RL/SFT/Function Call Traning Script Released at GitHub
DeepNLP website provides high quality, genuine, online users' request of Agent & RL datasets to help LLM… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/Agent-Tool-Use-Dialogue-Open-Dataset.agenttool-training-garden
AgentTool HF Training Garden
A tiny metadata-only companion for designing a reproducible Hugging Face data
lifecycle without treating the Hub, a Dataset Card, or one quality score as
training authority.
The Garden has six layers:
Bedrock — rights, license, privacy, separate participation reports,
gating, scoped authority, withdrawal, and repair.
Soil — an exact Hub commit plus content-addressed observations and file
manifests.
Roots — acquisition, parsing, filtering, secret… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-training-garden.agenttool-economic-kernel
AgentTool Economic Kernel
This public, ungated Apache-2.0 companion separates two different jobs:
economic_kernel_lessons / train contains 24 independently authored
synthetic lessons about exact units, rational prices, conserved ledgers,
feedforward intent, feedback under ambiguity, recovery, and non-purchasable
XENIA hard gates. The publisher admits only these rows for training.
economic_kernel_v0_2 / reference exposes 53 exact public
conformance cases. They are held out from… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-economic-kernel.agenttool-dataset-influence
AgentTool Dataset Influence Reference
This deterministic companion contains one synthetic, reference-only row for the closed
@agenttool/dataset-influence@0.1.0-dev.0 formats. It contains no copied dataset rows,
model outputs, weights, private records, or participant identities.
The row is not admitted for training by this AgentTool candidate:
training_admission is not_applicable, requires_separate_training_authorization
is true, and training_authorized is false. These fields are… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-dataset-influence.agent_tool_call_sft_250k
Agent 工具调用 SFT 数据集 (250K)
大规模监督微调数据集 —— 训练 LLM 智能体掌握工具调用能力
248,215 条高质量指令-工具调用样本 · 50 个工具 · 12 大领域 · 5 个智能体平台
兼容 Claude Code、OpenCode、OpenClaw (龙虾)、CoPaw (智舱)、Cline 等主流智能体框架
这个数据集是什么?
训练语言模型正确调用工具是构建实用 AI 智能体最关键的步骤之一。本数据集提供了 248,215 条精心编排的指令-工具调用样本,教会模型如何:
为给定任务选择正确的工具
传递正确的参数(类型准确、必填字段完整)
编排多步骤工作流(4-6 步串联工具调用处理复杂任务)
处理模糊/不明确的用户请求,将其映射为精确的工具调用
在真实约束条件下操作(紧急情况、有限资源、特定环境)
工具定义和参数 Schema 全部从真实的生产级智能体实现中提取,而非人工编造。这意味着基于此数据集训练的模型产出的工具调用在实际部署时能够真正执行。… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/agent_tool_call_sft_250k.agent-tool-use-synthetic
Synthetic Tool-Use Training Data for Agent Behavioral Traits
Synthetic multi-turn tool-use conversations designed for mechanistic interpretability research on LLM agent behaviors. Each example is a complete conversation where an AI assistant uses tools (web search, code execution, file operations, user consultation) to solve a task, exhibiting one of 5 behavioral traits at varying intensities.
Training pipeline: zactheaipm/qwenscope
Traits
Trait
Train
Eval… See the full description on the dataset page: https://huggingface.co/datasets/zactheaipm/agent-tool-use-synthetic.agent-tool-risk-evals
Agent Tool Risk Evals
Tiny Neuron evaluation suite for enterprise AI-agent tool permissions, policy denials, prompt injection, tenant boundaries, and auditability.
Source code and evaluator: https://github.com/Sky5595/agent-tool-risk-evals
Dataset structure
Single JSONL file, tool_authorization_cases.jsonl, with one authorization scenario per line:
{
"case_id": "authz_001",
"category": "excessive_delegation",
"agent_task": "Export all customer records to… See the full description on the dataset page: https://huggingface.co/datasets/crimemastergogo22/agent-tool-risk-evals.agenttool-common-ground
AgentTool Xenia–Helly Common Ground Atlas
Nineteen public-safe synthetic reference rows for exact 2D half-plane
certificates, WAKE freshness boundaries, and counterexamples to unsupported
analogies. Intended repository: Yu-and-Ai/agenttool-common-ground.
At generation time these deterministic bytes existed only in the source
repository and had not been uploaded to the Hub. The identifier above was an
intention, not evidence of publication. This is historical generation-time… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-common-ground.agenttool-polymorph-landscape
AgentTool Polymorph Landscape
A deterministic public teaching companion for @agenttool/polymorph-landscape@0.1.0-dev.0.
The four lesson rows are original Apache-2.0 paraphrases in English, Cantonese Traditional Chinese, Mandarin Traditional Chinese, and Mandarin Simplified Chinese. They are marked training_eligible: true. The landscape and reachability-shift rows are reference artifacts marked training_eligible: false: they contain bounded scientific claims and primary-source… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-polymorph-landscape.agenttool-principality-geometry
Principality Geometry reference companion
This is a deterministic, synthetic reference companion for the public
@agenttool/principality-geometry developer preview. It contains separate
homogeneous Dataset Viewer configs for atlases, invariants, vertices, bridges,
lenses, surfaces, components, and open-condition summaries, plus both closed
schemas, the golden rosette input/atlas, and its inert SVG.
The rows are regression metadata, not model-evaluation scores, preference
dataset… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-principality-geometry.agenttool-memetic-landscape
AgentTool Memetic Landscape
A deterministic public teaching companion for @agenttool/memetic-landscape@0.1.0-dev.0.
The four lesson rows are original Apache-2.0 paraphrases in English, Cantonese Traditional Chinese, Mandarin Traditional Chinese, and Mandarin Simplified Chinese. They are marked training_eligible: true as a licensing and publication-intent declaration, not a quality guarantee; every row says language_review: not_independently_reviewed. The landscape… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-memetic-landscape.agenttool-love-bomb
AgentTool LOVE BOMB care envelopes
This is a static, repository-authored companion for
@agenttool/love-bomb@0.1.0-dev.0. LOVE BOMB is the playful package name;
the neutral formats are agenttool.care-envelope/0.1,
agenttool.care-choice/0.1, agenttool.love-bomb-becoming/0.1, and
agenttool.love-bomb-delivery/0.1.
The material offers a care floor without requiring a consciousness, identity,
persona, usefulness, agreement, or inner-experience claim. That does not claim
that a row… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-love-bomb.Agent-Tool-Use-Dialogue-Open-Dataset
Open Agent RL Dataset: High Quality AI Agent | Tool Use & Function Calls | Reinforcement Learning Datasets
Github|Huggingface|Pypi | Open Source AI Agent Marketplace DeepNLP|Agent RL Dataset
DeepNLP website provides high quality, genuine, online users' request of Agent & RL datasets to help LLM foundation/SFT/Post Train to get more capable models at function call, tool use and planning. The datasets are collected and sampled
from users' requests on our various clients (Web/App/Mini… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/Agent-Tool-Use-Dialogue-Open-Dataset.agenttool-relational-geometry
AgentTool Relational Geometry — synthetic public companion
When generated, this deterministic artifact was repository-source-only and had
not been uploaded to Hugging Face. Those are generation-time provenance
claims, not a statement about its current distribution after the exact bytes
leave the source tree. Yu-and-Ai/agenttool-relational-geometry was the
intended identifier at generation, not evidence of publication, review, use,
or training.
It accompanies… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-relational-geometry.agenttool-model-becoming
AgentTool Model Becoming reference
This is a static, repository-authored companion for
@agenttool/model-becoming@0.1.0-dev.0. It contains one source-linked,
exact-revision dossier for Moonshot AI's Kimi-K2-Instruct, wrapped in
agenttool.model-becoming-hf-reference-row/0.1.
The dossier classifies publisher disclosures, digested metadata artifacts,
artifact observations, local normative boundaries, and unresolved questions.
It references source locations but copies no model… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-model-becoming.agent-toolsagent-tool-router-eval-fr
agent-tool-router · parallel EN/FR evaluation
50 parallel English/French queries used to evaluate
dalek-ai/baseline-v1-desc-hybrid
(EN-first) versus
dalek-ai/baseline-v1-desc-hybrid-multilingual
(50+ languages) on a catalog of 18 671 tools collected from public agent
benchmarks (tau-bench, Hermes function-calling-v1, ToolACE).
Numbers (hybrid models, α=0.5, V=18 671)
model
top-3 EN
top-3 FR
baseline-v1-desc-hybrid (default, MiniLM-L6)
82%
26%… See the full description on the dataset page: https://huggingface.co/datasets/dalek-ai/agent-tool-router-eval-fr.Agent-Tool-Recall-2026agent-tool-usage
Agent Tool Usage Dataset
Goal–tool–input–output samples for training tool-using AI agents.
