datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bfcl-v1-non-live-ast-hermesbfcl-v1-non-live-ast-parsed
[PARSED] BFCL V1 AST (non-live python)
The data in this dataset is a subset of the original gorilla-llm/Berkeley-Function-Calling-Leaderboard
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
simple
no
no
no
tool_calls
400
multiple
no
no
yes
tool_calls
200
parallel
no
yes
no
tool_calls
200
parallel_multiple
no
yes
yes
tool_calls
200
This is a re-parsing formatting dataset for Python AST parts from V1 of the official dataset of… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/bfcl-v1-non-live-ast-parsed.bfcl_v3bfcl_v3_multi_turn_basebfcl_v3BFCL_v3_audio
License
Gorilla is Apache 2.0 licensed, making it suitable for both academic and commercial use.
https://github.com/ShishirPatil/gorilla/blob/main/LICENSE
Citation
@article{patil2023gorilla,
title={Gorilla: Large Language Model Connected with Massive APIs},
author={Shishir G. Patil and Tianjun Zhang and Xin Wang and Joseph E. Gonzalez},
year={2023},
journal={arXiv preprint arXiv:2305.15334},
}
bfcl_v3bfcl-testToolWeave_BFCL_Rollout_Case_Study
🧵 ToolWeave BFCL Formal-Training Rollout Case Study
This dataset publishes the complete raw on-policy rollout artifact from ToolWeave formal-training update 2, together with a focused real-rollout case study and deterministic K=16 peer-group analysis for runtime-interaction credit assignment. The records contain protocol failures and self-correction; they are raw reinforcement-learning trajectories, not curated demonstrations and not benchmark results.
Project:… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/ToolWeave_BFCL_Rollout_Case_Study.bfcl_v2_astbfcl_v2_pythonbfcl-ms
BFCL v3 — Malay (Bahasa Malaysia) Edition
A Malay edition of the Berkeley Function-Calling Leaderboard (BFCL) v3 — the standard benchmark for evaluating LLM function/tool calling.
5,011 test entries across 21 categories (simple, multiple, parallel, parallel_multiple, irrelevance, java, javascript, rest, sql, live_*, chatable, multi_turn_*), with ground-truth answer files unchanged so the official BFCL scoring harness runs unmodified.
What is translated to Malay:
User questions… See the full description on the dataset page: https://huggingface.co/datasets/khursani8/bfcl-ms.bfcl
bfcl Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model(s) hf-inference-providers/meta-llama/Llama-3.1-8B-Instruct using the eval inspect_evals/bfcl from Inspect Evals.
To browse the results interactively, visit this Space.
Command
This eval was run with:
evaljobs inspect_evals/bfcl \
--model hf-inference-providers/meta-llama/Llama-3.1-8B-Instruct \
--name bfcl
Run with other models
To run this… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/bfcl.bfcl_parity_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192648BFCL-V4-Parallel-Native
BFCL V4 Parallel Native
Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration.
Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files.
Fields
id
official_category
task_type
user_prompt
function
ground_truth
Categories
live_parallel
live_parallel_multiple
parallel
parallel_multiple
Counts
train: 352 rows
eval: 88 rows
total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.bfcl-v3-02-27-metrics-trajectoriesDCAgent2_bfcl-parity_laion_GLM-4_7-inferredbugs-sandboxes-maxeps-131k_20260226_044018bfcl_parity_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_214919bfcl_parity_GLM_4_7_swesmith_sandboxes_with_tests_oracle_verified_120s_maxeps_1015c97b3bfcl_multi_turn_datasetBFCL-V4-Parallel-Multi-Turn
BFCL V4 Parallel Multi-Turn
Flattened current-turn rows from BFCL v4 multi-turn trajectories for decentralized multi-agent function-calling experiments.
Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files.
Fields
id
official_category
task_type
user_prompt
function
ground_truth
turn_index
Categories
multi_turn_base_step
multi_turn_long_context_step
multi_turn_miss_func_step… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Multi-Turn.bfclbfcl-activations-fulloutputBFCL probe activations, answer-INCLUSIVE (prompt+think+answer prefill), qwen3-8b, dense layers 10-34, last-token + mean pooling. acts_full_output/ = 1,800-problem split, acts_full_output_extra/ = 1,841 complement. kfold_ansinc/ = K=3 answer-inclusive linear classifiers (layer 28) deployed in live best-of-100 (condition C). Rebuild: experiments/bfcl_cot_clf/ in github.com/genlm/rollouts (clement/wip).
TRBench-BFCL
TRBench-BFCL
[Paper] |
[Model] |
[Benchmark] |
[Code]
💡 Summary
This dataset is a part of ToolRM: Towards Agentic Tool-Use Reward Modeling and serves as a dedicated benchmark for evaluating reward models in tool-use settings. It comprises 2,983 preference annotations buit upon BFCL V3, with assistant responses extracted from archived trajectories available in this github repo.
🌟 Overview
ToolRM is a family of lightweight generative and… See the full description on the dataset page: https://huggingface.co/datasets/RioLee/TRBench-BFCL.bfcl_v3_10eachbfclbfcl-nonlivebfclbfcl-livebfcl_v3_only_multi_turn
