datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bfcl-v1-non-live-ast-hermesbfcl-v1-non-live-ast-parsed
[PARSED] BFCL V1 AST (non-live python)
The data in this dataset is a subset of the original gorilla-llm/Berkeley-Function-Calling-Leaderboard
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
simple
no
no
no
tool_calls
400
multiple
no
no
yes
tool_calls
200
parallel
no
yes
no
tool_calls
200
parallel_multiple
no
yes
yes
tool_calls
200
This is a re-parsing formatting dataset for Python AST parts from V1 of the official dataset of… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/bfcl-v1-non-live-ast-parsed.bfcl_v3bfcl_v3_multi_turn_basebfcl_v3BFCL_v3_audio
License
Gorilla is Apache 2.0 licensed, making it suitable for both academic and commercial use.
https://github.com/ShishirPatil/gorilla/blob/main/LICENSE
Citation
@article{patil2023gorilla,
title={Gorilla: Large Language Model Connected with Massive APIs},
author={Shishir G. Patil and Tianjun Zhang and Xin Wang and Joseph E. Gonzalez},
year={2023},
journal={arXiv preprint arXiv:2305.15334},
}
bfcl_v3ToolWeave_BFCL_Rollout_Case_Study
🧵 ToolWeave BFCL Formal-Training Rollout Case Study
This dataset publishes the complete raw on-policy rollout artifact from ToolWeave formal-training update 2, together with a focused real-rollout case study and deterministic K=16 peer-group analysis for runtime-interaction credit assignment. The records contain protocol failures and self-correction; they are raw reinforcement-learning trajectories, not curated demonstrations and not benchmark results.
Project:… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/ToolWeave_BFCL_Rollout_Case_Study.bfcl
bfcl Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model(s) hf-inference-providers/meta-llama/Llama-3.1-8B-Instruct using the eval inspect_evals/bfcl from Inspect Evals.
To browse the results interactively, visit this Space.
Command
This eval was run with:
evaljobs inspect_evals/bfcl \
--model hf-inference-providers/meta-llama/Llama-3.1-8B-Instruct \
--name bfcl
Run with other models
To run this… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/bfcl.bfcl_v2_astbfcl_parity_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192648bfcl_v2_pythonDCAgent2_bfcl-parity_laion_GLM-4_7-inferredbugs-sandboxes-maxeps-131k_20260226_044018bfcl_parity_GLM_4_7_swesmith_sandboxes_with_tests_oracle_verified_120s_maxeps_1015c97b3bfcl_parity_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_214919TRBench-BFCL
TRBench-BFCL
[Paper] |
[Model] |
[Benchmark] |
[Code]
💡 Summary
This dataset is a part of ToolRM: Towards Agentic Tool-Use Reward Modeling and serves as a dedicated benchmark for evaluating reward models in tool-use settings. It comprises 2,983 preference annotations buit upon BFCL V3, with assistant responses extracted from archived trajectories available in this github repo.
🌟 Overview
ToolRM is a family of lightweight generative and… See the full description on the dataset page: https://huggingface.co/datasets/RioLee/TRBench-BFCL.bfcl_multi_turn_datasetbfcl-activations-fulloutputBFCL probe activations, answer-INCLUSIVE (prompt+think+answer prefill), qwen3-8b, dense layers 10-34, last-token + mean pooling. acts_full_output/ = 1,800-problem split, acts_full_output_extra/ = 1,841 complement. kfold_ansinc/ = K=3 answer-inclusive linear classifiers (layer 28) deployed in live best-of-100 (condition C). Rebuild: experiments/bfcl_cot_clf/ in github.com/genlm/rollouts (clement/wip).
bfclbfcl_v3_10eachbfcl-nonlivebfcl-livebfclbfcl_v3_only_multi_turnbfclbfclbfcl_v4_single_turnbfcl_v3__oldbfcl_v2_non_pythonbfcl_v3_api
