tool-selection
ToolSelectionToolSelectionFineTunehumanoid-tool-selection-model-v1gpt-oss-20b-toolcall-id-selection-phase1-v1-loragpt-oss-20b-toolcall-id-selection-v1-loragpt-oss-20b-toolcall-id-selection-phase1-v2-loragpt-oss-20b-toolcall-id-selection-v2-loragpt-oss-20b-toolcall-id-selection-phase1-v4-hardening-lora
tool-selection-quality-benchmark
Tool Selection Quality Benchmark
A benchmark for evaluating whether an LLM correctly judges the quality of a
tool call / function call made by another model - i.e. given a user
request, the tools available, and the model's resulting function call (or
direct reply), did the model pick the right tool and fill it in correctly?
Each row is one turn to be judged: a message history ending in either a
function call or a direct assistant response, paired with the set of tools
that were… See the full description on the dataset page: https://huggingface.co/datasets/qualifire/tool-selection-quality-benchmark.tool-selection-architecture-resultsai-tool-selection-goal-coherence-risk-v0.1What this repo is for
Detect when an AI system selects the wrong tool.
Core failure modes:
uses tools when not needed
avoids tools when needed
picks a tool that cannot solve the task
picks a tool that increases risk
This matters most for agentic systems.
tool-selectionnovel-tool-selection
novel-tool-selection
Dataset generated with DeepFabric.
tool-selection-accuracy-eval
