CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /When2Call When2Call 💾 Github   |    📄 Paper Dataset Description: When2Call is a benchmark designed to evaluate tool-calling decision-making for large language models (LLMs), including when to generate a tool call, when to ask follow-up questions, when to admit the question can't be answered with the tools provided, and what to do if the question seems to require tool use but a tool call can't be made. We find that state-of-the-art tool-calling LMs show significant room for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/When2Call.texttext-generation10K<n<100K59 likes4.1k downloads1y agoHugging Face02interstellarninja /nvidia_hermes_when2calltext1K<n<10K0 likes57 downloads1y agoHugging Face03beyoru /when2call-GLM4.5-IItextn<1K1 likes24 downloads1y agoHugging Face04ihounie /when2call_imbalanced_request_10 when2call_imbalanced_request_10 Derived from nvidia/When2Call (train_pref, train) by downsampling the request category in chosen_response by 10% (keeping 90%). Sampling Target class: request Keep ratio: 90% Seed: 45 Counts (chosen_response) Source refusal: 2999 toolcall: 3000 request: 3001 unk: 0 Result refusal: 2999 toolcall: 3000 request: 2700 unk: 0 Final rows: 8699 1K<n<10K0 likes17 downloads6mo agoHugging Face05ihounie /when2call_imbalanced_refusal when2call_imbalanced_refusal This dataset is derived from nvidia/When2Call (train_pref, train split) by downsampling one chosen-response category to ~50% while keeping all other rows. Source Dataset: nvidia/When2Call Config: train_pref Split: train Source rows: 9000 Classification Rules (on chosen_response) Categories are assigned in this precedence order: toolcall if text contains <TOOLCALL> (case-insensitive) request if text contains ? request if text… See the full description on the dataset page: https://huggingface.co/datasets/ihounie/when2call_imbalanced_refusal.1K<n<10K0 likes14 downloads6mo agoHugging Face06langtech-languagemodeling /aliaboost_when2call_estextn<1K0 likes13 downloads6mo agoHugging Face07heegyu /When2Call-preftext1K<n<10K0 likes12 downloads1y agoHugging Face08ihounie /when2call_imbalanced_toolcall when2call_imbalanced_toolcall This dataset is derived from nvidia/When2Call (train_pref, train split) by downsampling one chosen-response category to ~50% while keeping all other rows. Source Dataset: nvidia/When2Call Config: train_pref Split: train Source rows: 9000 Classification Rules (on chosen_response) Categories are assigned in this precedence order: toolcall if text contains <TOOLCALL> (case-insensitive) request if text contains ? request if text… See the full description on the dataset page: https://huggingface.co/datasets/ihounie/when2call_imbalanced_toolcall.1K<n<10K0 likes11 downloads6mo agoHugging Face09ihounie /when2call_imbalanced_request_80 when2call_imbalanced_request_80 Derived from nvidia/When2Call (train_pref, train) by downsampling the request category in chosen_response by 80% (keeping 20%). Sampling Target class: request Keep ratio: 20% Seed: 46 Counts (chosen_response) Source refusal: 2999 toolcall: 3000 request: 3001 unk: 0 Result refusal: 2999 toolcall: 3000 request: 600 unk: 0 Final rows: 6599 1K<n<10K0 likes11 downloads6mo agoHugging Face10NotoriousH2 /instructkr-when2calltext10K<n<100K0 likes10 downloads9mo agoHugging Face11AmirhoseinGH /when2call-qwen3vl-2b-multiclasstabular10K<n<100K0 likes10 downloads6mo agoHugging Face12AlexHung29629 /When2Call_mistraltext10K<n<100K0 likes9 downloads1y agoHugging Face13MananSuri27 /when2call-test-with-schematabularn<1K0 likes9 downloads10mo agoHugging Face14MananSuri27 /when2call-grpo-argus-v2-certainty-onlytext1K<n<10K0 likes8 downloads1y agoHugging Face15beyoru /when2call-update0 likes8 downloads1y agoHugging Face16alucent /mirror-When2Callgated When2Call 💾 Github&nbsp;&nbsp; | &nbsp;&nbsp; 📄 Paper Dataset Description: When2Call is a benchmark designed to evaluate tool-calling decision-making for large language models (LLMs), including when to generate a tool call, when to ask follow-up questions, when to admit the question can't be answered with the tools provided, and what to do if the question seems to require tool use but a tool call can't be made. We find that state-of-the-art tool-calling LMs… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-When2Call.texttext-generation10K<n<100K0 likes8 downloads2mo agoHugging Face17deestudio /when2call_formattedd1 likes7 downloads1y agoHugging Face18beyoru /when2calltext1K<n<10K0 likes7 downloads1y agoHugging Face19heegyu /When2Call-Pref-RLtext1K<n<10K0 likes7 downloads10mo agoHugging Face20MananSuri27 /when2call-grpo-argustext1K<n<10K0 likes6 downloads1y agoHugging Face21beyoru /when2call-GLM4.5textn<1K0 likes5 downloads1y agoHugging Face22ihounie /when2call_imbalanced_request when2call_imbalanced_request This dataset is derived from nvidia/When2Call (train_pref, train split) by downsampling one chosen-response category to ~50% while keeping all other rows. Source Dataset: nvidia/When2Call Config: train_pref Split: train Source rows: 9000 Classification Rules (on chosen_response) Categories are assigned in this precedence order: toolcall if text contains <TOOLCALL> (case-insensitive) request if text contains ? request if text… See the full description on the dataset page: https://huggingface.co/datasets/ihounie/when2call_imbalanced_request.1K<n<10K0 likes5 downloads6mo agoHugging Face23MananSuri27 /when2call-grpotext1K<n<10K0 likes4 downloads1y agoHugging Face24AlexHung29629 /When2Call_mistral-segmentedgatedtext10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.