HiTZ/TOOLtifruti
The Basque evaluation ecosystem still lacks standardized datasets and protocols to assess agentic behavior, and in particular tool selection and tool use in end-to-end Agentic RAG settings. To address this gap, we introduce TOOLtifruti, an ad hoc dataset designed to evaluate whether an LLM can identify when a tool is needed and select the appropriate tool among multiple domain-specific options in our use case. This setup makes tool-calling evaluation straightforward and reproducible, and it… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/TOOLtifruti.
This repository belongs to HiTZ on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
