CoolFace
Datasetpublicgated

alucent/mirror-When2Call

When2Call 💾 Github   |    📄 Paper Dataset Description: When2Call is a benchmark designed to evaluate tool-calling decision-making for large language models (LLMs), including when to generate a tool call, when to ask follow-up questions, when to admit the question can't be answered with the tools provided, and what to do if the question seems to require tool use but a tool call can't be made. We find that state-of-the-art tool-calling LMs… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-When2Call.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes7downloads

No commit history came back for main. The revision may not exist, or the source declined the request.