datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bfcl-v1-non-live-ast-hermesbfcl-v1-non-live-ast-parsed
[PARSED] BFCL V1 AST (non-live python)
The data in this dataset is a subset of the original gorilla-llm/Berkeley-Function-Calling-Leaderboard
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
simple
no
no
no
tool_calls
400
multiple
no
no
yes
tool_calls
200
parallel
no
yes
no
tool_calls
200
parallel_multiple
no
yes
yes
tool_calls
200
This is a re-parsing formatting dataset for Python AST parts from V1 of the official dataset of… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/bfcl-v1-non-live-ast-parsed.bfcl_v3bfcl_v3_multi_turn_baseBFCL
Berkeley Function Calling Leaderboard
The Berkeley function calling leaderboard is a live leaderboard to evaluate the ability of different LLMs to call functions (also referred to as tools).
We built this dataset from our learnings to be representative of most users' function calling use-cases, for example, in agents, as a part of enterprise workflows, etc.
To this end, our evaluation dataset spans diverse categories, and across multiple languages.
Checkout the Leaderboard at… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/BFCL.bfcl_v3BFCL-Hi
Dataset Description:
The BFCL-Hi (Hindi BFCL) dataset evaluates the function-calling capability of large language models (LLMs) when questions are asked in Hindi. This is the GCP-translated version of the English BFCL dataset, in which question-function-answer pairs across various domains and multiple languages are originally curated in English.
This dataset is ready for commercial/non-commercial use.
The evaluation steps are described here.
Other Hindi benchmark datasets: [… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/BFCL-Hi.BFCL_v3_audio
License
Gorilla is Apache 2.0 licensed, making it suitable for both academic and commercial use.
https://github.com/ShishirPatil/gorilla/blob/main/LICENSE
Citation
@article{patil2023gorilla,
title={Gorilla: Large Language Model Connected with Massive APIs},
author={Shishir G. Patil and Tianjun Zhang and Xin Wang and Joseph E. Gonzalez},
year={2023},
journal={arXiv preprint arXiv:2305.15334},
}
bfcl_v3ru_bfcl
Berkeley Function Calling Leaderboard
The Berkeley function calling leaderboard is a live leaderboard to evaluate the ability of different LLMs to call functions (also referred to as tools).
We built this dataset from our learnings to be representative of most users' function calling use-cases, for example, in agents, as a part of enterprise workflows, etc.
To this end, our evaluation dataset spans diverse categories, and across multiple languages.
Checkout the Leaderboard at… See the full description on the dataset page: https://huggingface.co/datasets/AvitoTech/ru_bfcl.qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3
Qwen3.6-27B GGUF quantization on a bounded BFCL V4 pilot
Q4_K_M matched Q8_0 on both tested categories: each scored 94 of 100 selected cases correct. Q5_K_M also scored 94/100; Q3_K_M scored 92/100.
Read the results page · Inspect all 400 scored rows
This is a post-result-corrected exploratory analysis of two selected non-live BFCL V4 categories, not a full leaderboard result.
Inspect the scored rows without cloning
The Hub Dataset Viewer does not render this… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative-AI/qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3.bfcl-testbfcl-activations-thinkBFCL CoT-correctness probe activations (answer-WITHHELD: prompt+think_seg prefill, last-token + mean pooling, layer list in layers_*.json). Full 3,641-problem x 100-rollout t=1 pools: acts/ = 1,800-problem split, acts_extra/ = 1,841 complement (qwen3-8b; 14B only in acts/). kfold/ = deployed K=3 fold-blind linear classifiers + routing (BFCL_LIVE_BON_SPEC.md). Rebuild: experiments/bfcl_cot_clf/ in github.com/genlm/rollouts (branch clement/wip).
BFCL_v4_information
BFCL Dataset Documentation / Tài liệu Dataset BFCL
English | Tiếng Việt
📚 Table of Contents / Mục Lục
This documentation provides a comprehensive guide to the Berkeley Function Calling Leaderboard (BFCL) dataset. Each section is available in both English and Vietnamese.
Tài liệu này cung cấp hướng dẫn toàn diện về dataset Berkeley Function Calling Leaderboard (BFCL). Mỗi phần được viết song ngữ Anh-Việt.
🚀 Getting Started / Bắt Đầu
File
Description… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/BFCL_v4_information.ToolWeave_BFCL_Rollout_Case_Study
🧵 ToolWeave BFCL Formal-Training Rollout Case Study
This dataset publishes the complete raw on-policy rollout artifact from ToolWeave formal-training update 2, together with a focused real-rollout case study and deterministic K=16 peer-group analysis for runtime-interaction credit assignment. The records contain protocol failures and self-correction; they are raw reinforcement-learning trajectories, not curated demonstrations and not benchmark results.
Project:… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/ToolWeave_BFCL_Rollout_Case_Study.bfcl-paritybfcl_v2_pythonbfcl_v2_astBFCL-V4-Parallel-Native
BFCL V4 Parallel Native
Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration.
Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files.
Fields
id
official_category
task_type
user_prompt
function
ground_truth
Categories
live_parallel
live_parallel_multiple
parallel
parallel_multiple
Counts
train: 352 rows
eval: 88 rows
total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.bfcl_shuffle_full
Berkeley Function Calling Leaderboard
The Berkeley function calling leaderboard is a live leaderboard to evaluate the ability of different LLMs to call functions (also referred to as tools).
We built this dataset from our learnings to be representative of most users' function calling use-cases, for example, in agents, as a part of enterprise workflows, etc.
To this end, our evaluation dataset spans diverse categories, and across multiple languages.
Checkout the Leaderboard at… See the full description on the dataset page: https://huggingface.co/datasets/BitAgent/bfcl_shuffle_full.bfcl_parity_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192648bfcl-ms
BFCL v3 — Malay (Bahasa Malaysia) Edition
A Malay edition of the Berkeley Function-Calling Leaderboard (BFCL) v3 — the standard benchmark for evaluating LLM function/tool calling.
5,011 test entries across 21 categories (simple, multiple, parallel, parallel_multiple, irrelevance, java, javascript, rest, sql, live_*, chatable, multi_turn_*), with ground-truth answer files unchanged so the official BFCL scoring harness runs unmodified.
What is translated to Malay:
User questions… See the full description on the dataset page: https://huggingface.co/datasets/khursani8/bfcl-ms.bfcl
bfcl Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model(s) hf-inference-providers/meta-llama/Llama-3.1-8B-Instruct using the eval inspect_evals/bfcl from Inspect Evals.
To browse the results interactively, visit this Space.
Command
This eval was run with:
evaljobs inspect_evals/bfcl \
--model hf-inference-providers/meta-llama/Llama-3.1-8B-Instruct \
--name bfcl
Run with other models
To run this… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/bfcl.BFCL_V4-trajectories
AgentSuite/BFCL_V4-trajectories
Per-model agent trajectory data for BFCL_V4 (public release).
Models: 30
Tasks per model: 5,865
One file per model: {model}.jsonl, one JSON object per line.
Fields: model_path, user_model_path, benchmark_name, task_name, sampling_params, user_sampling_params, messages, eval_result, meta.
sampling_params reflect each benchmark's own implementation; values the benchmark leaves unset are recorded as null (provider default).
Models… See the full description on the dataset page: https://huggingface.co/datasets/AgentSuite/BFCL_V4-trajectories.bfcl_full_collectiongranite-bfcl-randopt-resultsbfcl-v3-02-27-metrics-trajectoriespsd-bfcl
PSD-BFCL
Research artifacts for Privileged Self-Distillation (PSD) on the
Berkeley Function Calling Leaderboard (BFCL) multi-turn tool-use tasks.
PSD turns failed rollouts into verified local training targets: a constructor
adds privileged context (a hint), the environment verifies that the resulting
repair passes, and the hint-conditioned distribution is distilled into a
student that never receives the hint.
Code, exact run summaries, and experiment notes live in the companion… See the full description on the dataset page: https://huggingface.co/datasets/essamsleiman/psd-bfcl.DCAgent2_bfcl-parity_laion_GLM-4_7-inferredbugs-sandboxes-maxeps-131k_20260226_044018bfcl_parity_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_214919
