HiTZ/TOOLtifruti
The Basque evaluation ecosystem still lacks standardized datasets and protocols to assess agentic behavior, and in particular tool selection and tool use in end-to-end Agentic RAG settings. To address this gap, we introduce TOOLtifruti, an ad hoc dataset designed to evaluate whether an LLM can identify when a tool is needed and select the appropriate tool among multiple domain-specific options in our use case. This setup makes tool-calling evaluation straightforward and reproducible, and it… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/TOOLtifruti.
The Basque evaluation ecosystem still lacks standardized datasets and protocols to assess agentic behavior, and in particular tool selection and tool use in end-to-end Agentic RAG settings. To address this gap, we introduce TOOLtifruti, an ad hoc dataset designed to evaluate whether an LLM can identify when a tool is needed and select the appropriate tool among multiple domain-specific options in our use case. This setup makes tool-calling evaluation straightforward and reproducible, and it also provides a common benchmark for end-to-end Agentic RAG evaluation, since each query is paired with a reference tool call and a reference answer.
TOOLtifruti contains queries from five domains. Each query is associated with one of these domains (and therefore with the corresponding tool). The dataset includes queries that can be answered without using any tool (no-tool). These queries span various categories, such as translation, mathematics, education, and programming. These are the domains:
- BOPV/EHAA
- Basque Parliament
- Berria Newspaper
- Basque Wikipedia
- No-tool queries
Each dataset instance includes the query, its type (i.e., the tool/domain it maps to), the reference context required to answer it (the source passage from which the query was created), and the corresponding reference answer derived from that context. The following table shows a representative example for each domain:
