datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToolWeave_BFCL_Rollout_Case_Study
🧵 ToolWeave BFCL Formal-Training Rollout Case Study
This dataset publishes the complete raw on-policy rollout artifact from ToolWeave formal-training update 2, together with a focused real-rollout case study and deterministic K=16 peer-group analysis for runtime-interaction credit assignment. The records contain protocol failures and self-correction; they are raw reinforcement-learning trajectories, not curated demonstrations and not benchmark results.
Project:… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/ToolWeave_BFCL_Rollout_Case_Study.bfcl
bfcl Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model(s) hf-inference-providers/meta-llama/Llama-3.1-8B-Instruct using the eval inspect_evals/bfcl from Inspect Evals.
To browse the results interactively, visit this Space.
Command
This eval was run with:
evaljobs inspect_evals/bfcl \
--model hf-inference-providers/meta-llama/Llama-3.1-8B-Instruct \
--name bfcl
Run with other models
To run this… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/bfcl.bfcl-v3-02-27-metrics-trajectoriesTRBench-BFCL
TRBench-BFCL
[Paper] |
[Model] |
[Benchmark] |
[Code]
💡 Summary
This dataset is a part of ToolRM: Towards Agentic Tool-Use Reward Modeling and serves as a dedicated benchmark for evaluating reward models in tool-use settings. It comprises 2,983 preference annotations buit upon BFCL V3, with assistant responses extracted from archived trajectories available in this github repo.
🌟 Overview
ToolRM is a family of lightweight generative and… See the full description on the dataset page: https://huggingface.co/datasets/RioLee/TRBench-BFCL.CloudSurf-4B-FC-bfcl-results
CloudSurf-4B-FC — raw BFCL V4 result files
Raw, unmodified BFCL V4 evaluation outputs backing the leaderboard submission
PR ShishirPatil/gorilla#1357
for CloudSurf-4B-FC
(a google/gemma-4-E4B-it fine-tune, Apache-2.0).
Both sides are included: our tuned runs and the stock gemma-4-E4B-it
baselines re-measured on the identical rig, so every number in the PR can be
recomputed from primary files.
Whiskers are the min–max across the three runs on each side. Stock wins
Irrelevance… See the full description on the dataset page: https://huggingface.co/datasets/cloudsurf-software/CloudSurf-4B-FC-bfcl-results.bfcl-rollouts
BFCL v4 single-turn rollouts — Qwen3 thinking, t=1.0, n=100
728,200 completions: Qwen3-8B and Qwen3-14B (thinking mode), 100 i.i.d. samples per problem at
temperature 1.0 (Qwen3-card sampling: top_p 0.95, top_k 20, seed 42, max_tokens 16384, vLLM 0.21
float16), over the 13 single-turn categories of BFCL v4 (3,641 problems; benchmark = gorilla
bfcl_eval, snapshot matching the 2025-12-16 leaderboard release). Prompts are the exact prompt-mode
ChatML strings the BFCL harness feeds… See the full description on the dataset page: https://huggingface.co/datasets/clementepasti/bfcl-rollouts.bfcl-gpt-oss-20b-test
bfcl-gpt-oss-20b-test Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model(s) hf-inference-providers/openai/gpt-oss-20b:fastest using the eval inspect_evals/bfcl from Inspect Evals.
To browse the results interactively, visit this Space.
Command
This eval was run with:
evaljobs inspect_evals/bfcl \
--model hf-inference-providers/openai/gpt-oss-20b:fastest \
--name bfcl-gpt-oss-20b-test \
--limit 50… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/bfcl-gpt-oss-20b-test.bfcl-index
BFCL explorer index
Query index for the genlm/rollouts BFCL public-traces
explorer (docs/bfcl.html): one row per (release batch, model, benchmark question) of the
BFCL-Result archive — verdict, error type, token
counts, a preview, and the byte offset of the full record in the archive's raw JSONL (which the
explorer range-fetches from GitHub directly). Built by local/build_bfcl_index.py; published by
local/publish_bfcl_index.py.
bfcl-olmo
bfcl-olmo Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model hf-inference-providers/allenai/Olmo-3-7B-Instruct using the eval inspect_evals/bfcl from Inspect Evals.
To browse the results interactively, visit this Space.
How to Run This Eval
pip install git+https://github.com/dvsrepo/evaljobs.git
export HF_TOKEN=your_token_here
evaljobs dvilasuero/bfcl-olmo \
--model <your-model> \
--name <your-name> \
--flavor… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/bfcl-olmo.bfcl_v3_test_trace_datasetbfcl-kimi
bfcl-kimi Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model hf-inference-providers/moonshotai/Kimi-K2-Thinking:fastest using the eval inspect_evals/bfcl from Inspect Evals.
To browse the results interactively, visit this Space.
How to Run This Eval
pip install git+https://github.com/dvsrepo/evaljobs.git
export HF_TOKEN=your_token_here
evaljobs dvilasuero/bfcl-kimi \
--model <your-model> \
--name <your-name> \… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/bfcl-kimi.loom-benchmark-bfclcollabllm-multiturn-bfcl
