datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tool_callingtool_calling_shufflebfcl_shuffle_full
Berkeley Function Calling Leaderboard
The Berkeley function calling leaderboard is a live leaderboard to evaluate the ability of different LLMs to call functions (also referred to as tools).
We built this dataset from our learnings to be representative of most users' function calling use-cases, for example, in agents, as a part of enterprise workflows, etc.
To this end, our evaluation dataset spans diverse categories, and across multiple languages.
Checkout the Leaderboard at… See the full description on the dataset page: https://huggingface.co/datasets/BitAgent/bfcl_shuffle_full.ansibletool_shuffle_smallDataset for continuously updating tasks representitive of BFCL evaluation criteria. This dataset does not contain the multi-turn tasks.
tool_shuffle_small_testbfcl_shuffle_smallBFCL Tasks to evaluate
