datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bfcl-v1-non-live-ast-parsed
[PARSED] BFCL V1 AST (non-live python)
The data in this dataset is a subset of the original gorilla-llm/Berkeley-Function-Calling-Leaderboard
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
simple
no
no
no
tool_calls
400
multiple
no
no
yes
tool_calls
200
parallel
no
yes
no
tool_calls
200
parallel_multiple
no
yes
yes
tool_calls
200
This is a re-parsing formatting dataset for Python AST parts from V1 of the official dataset of… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/bfcl-v1-non-live-ast-parsed.BFCL-Hi
Dataset Description:
The BFCL-Hi (Hindi BFCL) dataset evaluates the function-calling capability of large language models (LLMs) when questions are asked in Hindi. This is the GCP-translated version of the English BFCL dataset, in which question-function-answer pairs across various domains and multiple languages are originally curated in English.
This dataset is ready for commercial/non-commercial use.
The evaluation steps are described here.
Other Hindi benchmark datasets: [… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/BFCL-Hi.qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3
Qwen3.6-27B GGUF quantization on a bounded BFCL V4 pilot
Q4_K_M matched Q8_0 on both tested categories: each scored 94 of 100 selected cases correct. Q5_K_M also scored 94/100; Q3_K_M scored 92/100.
Read the results page · Inspect all 400 scored rows
This is a post-result-corrected exploratory analysis of two selected non-live BFCL V4 categories, not a full leaderboard result.
Inspect the scored rows without cloning
The Hub Dataset Viewer does not render this… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative-AI/qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3.ru_bfcl
Berkeley Function Calling Leaderboard
The Berkeley function calling leaderboard is a live leaderboard to evaluate the ability of different LLMs to call functions (also referred to as tools).
We built this dataset from our learnings to be representative of most users' function calling use-cases, for example, in agents, as a part of enterprise workflows, etc.
To this end, our evaluation dataset spans diverse categories, and across multiple languages.
Checkout the Leaderboard at… See the full description on the dataset page: https://huggingface.co/datasets/AvitoTech/ru_bfcl.bfcl-ms
BFCL v3 — Malay (Bahasa Malaysia) Edition
A Malay edition of the Berkeley Function-Calling Leaderboard (BFCL) v3 — the standard benchmark for evaluating LLM function/tool calling.
5,011 test entries across 21 categories (simple, multiple, parallel, parallel_multiple, irrelevance, java, javascript, rest, sql, live_*, chatable, multi_turn_*), with ground-truth answer files unchanged so the official BFCL scoring harness runs unmodified.
What is translated to Malay:
User questions… See the full description on the dataset page: https://huggingface.co/datasets/khursani8/bfcl-ms.BFCL-V4-Parallel-Native
BFCL V4 Parallel Native
Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration.
Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files.
Fields
id
official_category
task_type
user_prompt
function
ground_truth
Categories
live_parallel
live_parallel_multiple
parallel
parallel_multiple
Counts
train: 352 rows
eval: 88 rows
total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.psd-bfcl
PSD-BFCL
Research artifacts for Privileged Self-Distillation (PSD) on the
Berkeley Function Calling Leaderboard (BFCL) multi-turn tool-use tasks.
PSD turns failed rollouts into verified local training targets: a constructor
adds privileged context (a hint), the environment verifies that the resulting
repair passes, and the hint-conditioned distribution is distilled into a
student that never receives the hint.
Code, exact run summaries, and experiment notes live in the companion… See the full description on the dataset page: https://huggingface.co/datasets/essamsleiman/psd-bfcl.BFCL-V4-Parallel-Multi-Turn
BFCL V4 Parallel Multi-Turn
Flattened current-turn rows from BFCL v4 multi-turn trajectories for decentralized multi-agent function-calling experiments.
Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files.
Fields
id
official_category
task_type
user_prompt
function
ground_truth
turn_index
Categories
multi_turn_base_step
multi_turn_long_context_step
multi_turn_miss_func_step… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Multi-Turn.tw-bfcl-v4
TW-BFCL v4
A Taiwan Traditional Chinese (zh-TW) localization of the
Berkeley Function Calling Leaderboard v4 non-agentic
categories, built to measure tool-calling ability for Chinese-speaking users — specifically for a
Taiwan office phone-attendant / voice-agent setting.
Almost all function-calling benchmarks are English-only. A model that calls tools well in English may
degrade badly on Chinese input, and without a benchmark like this the regression is invisible.… See the full description on the dataset page: https://huggingface.co/datasets/Luigi/tw-bfcl-v4.
