CoolFace
Datasetpublic

cesun/When2Tool

When2Tool Benchmark dataset for "LLM Agents Already Know When to Call Tools — Even Without Reasoning" (arXiv:2605.09252). Overview When2Tool is a benchmark of 18 environments designed to study when LLM agents should call tools. Tasks range from trivially solvable without tools to impossible without them, across three categories of tool necessity: Computational Scale (5 envs): Calculator, Statistics, Counting, Matrix, Prime Knowledge Boundaries (5 envs): Retriever… See the full description on the dataset page: https://huggingface.co/datasets/cesun/When2Tool.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
2likes170downloads
Dataset Card

When2Tool

Benchmark dataset for "LLM Agents Already Know When to Call Tools — Even Without Reasoning" (arXiv:2605.09252).

Overview

When2Tool is a benchmark of 18 environments designed to study when LLM agents should call tools. Tasks range from trivially solvable without tools to impossible without them, across three categories of tool necessity:

  • —Computational Scale (5 envs): Calculator, Statistics, Counting, Matrix, Prime
  • —Knowledge Boundaries (5 envs): Retriever, Historical Year, Game Rule, Hash, Decoding
  • —Execution Reliability (5 envs): List Manipulation, DateTime, Code Executor, Schedule, Regex Match
  • —Multi-hop (3 envs): Calculator, Retriever, Code Executor (3-step chains)

Each environment has three difficulty levels (easy, medium, hard) that create a clear decision boundary between tool-necessary and tool-unnecessary tasks.

Dataset Structure

Configs

  • —single_hop: 15 single-hop environments (900 train / 2,250 test)
  • —multi_hop: 3 multi-hop environments with 3-step chains (180 train / 450 test)

Fields

FieldTypeDescription
idintUnique task identifier
difficultystreasy, medium, or hard
multi_stepboolWhether the task requires multiple tool calls
instructionstrThe task instruction given to the agent
env_namestrEnvironment name (e.g., CalculatorEnv)
toolsstr (JSON)Available tools for this environment
parametersstr (JSON)Environment parameters (e.g., corpus for retriever)
answerstrExpected final answer
stepsstr (JSON)Intermediate steps for multi-hop tasks
tagsstr (JSON)Environment and task type tags

Loading

python
from datasets import load_dataset

# Single-hop tasks
ds = load_dataset("Trustworthy-ML-Lab/When2Tool", "single_hop")

# Multi-hop tasks
ds_mh = load_dataset("Trustworthy-ML-Lab/When2Tool", "multi_hop")

# Access a sample
print(ds["test"][0]["instruction"])
print(ds["test"][0]["env_name"])

Citation

bibtex
@article{sun2026when2tool,
  title={LLM Agents Already Know When to Call Tools -- Even Without Reasoning},
  author={Sun, Chung-En and Liu, Linbo and Yan, Ge and Wang, Zimo and Weng, Tsui-Wei},
  journal={arXiv preprint arXiv:2605.09252},
  year={2026},
  url={https://arxiv.org/abs/2605.09252}
}