datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tool-selection-quality-benchmark
Tool Selection Quality Benchmark
A benchmark for evaluating whether an LLM correctly judges the quality of a
tool call / function call made by another model - i.e. given a user
request, the tools available, and the model's resulting function call (or
direct reply), did the model pick the right tool and fill it in correctly?
Each row is one turn to be judged: a message history ending in either a
function call or a direct assistant response, paired with the set of tools
that were… See the full description on the dataset page: https://huggingface.co/datasets/qualifire/tool-selection-quality-benchmark.tool-selection-architecture-resultstool-selectiontool-selection-accuracy-evalckan-tool-selectionTool_Selection_Disambiguation
🇰🇿 Kazakh Tool Selection and Disambiguation Dataset
Dataset Summary
Kazakh Tool Selection and Disambiguation Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI scenarios that require choosing the most appropriate tool from multiple available options.
The dataset focuses on tool-selection reasoning, where the assistant must understand the user’s intent, compare available tools, avoid unnecessary… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Tool_Selection_Disambiguation.phi2-tool-selection
