datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
to-tool-call-papers
📚 To-Tool-Call Papers
A curated paper library for LLM tool use, function calling, agent training, and environment synthesis
To-Tool-Call Papers is a bilingual research library for tracking papers on tool use, function calling, agent data synthesis, environment scaling, agentic RL, and tool-use benchmarks.
Quick Start ·
At a Glance ·
Files ·
Schema ·
Copyright
[!IMPORTANT]
This dataset is a research reading collection… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-papers.tool-callsTool calling master dataset
Contains the following:
Query -> Available tools (name + description + schema) -> Tool name
Sources (identified by source column):
subsets of existing tool-calling dataset sources parsed into the above format
synthetic data
Will be parsed into the following two passes:
Query -> List of tool names + descriptions -> Tool name
Tool name + tool schema -> Tool call
mcp-tool-calling-benchmark
MCP Tool-Calling Benchmark
A benchmark dataset for evaluating AI assistants' MCP (Model Context Protocol) tool-calling accuracy across 12 platforms.
Dataset Description
Contains 6,451 labeled interaction logs from systematic QA testing of Grok's MCP connectors. Each row captures a test prompt, the expected tool invocation, Grok's actual response, and the error classification.
Platforms Covered
Platform
Prompts
Tools Tested
Primary Error Pattern… See the full description on the dataset page: https://huggingface.co/datasets/brijeshvadi/mcp-tool-calling-benchmark.VStarBench_sft_w_toolcallmath-toolcall-tr-benchmark
math-toolcall-tr-benchmark
bilalabic/gemma_4_math-toolcall-tr_lora
LoRA adaptörünü temel Gemma-4 E4B modeliyle karşılaştıran benchmark sonuçları.
Bu depo yalnızca değerlendirme çıktılarını içerir. Eğitim veri seti ayrı olarak
bilalabic/math-toolcall-tr
adresinde yayımlanmaktadır.
Benchmark'lar
Benchmark
Örnek
Ölçülen davranış
Türkçe MMLU
250
Genel bilgi doğruluğu ve eğitim sonrası bilgi kaybı
Matematik Tool-Call
150
Araç seçimi, çekimserlik ve çıktı… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/math-toolcall-tr-benchmark.VStarBench_rl_wo_toolcallVStarBench_rl_w_toolcallchadgpt-tool-calling-150VStarBench_sft_wo_toolcall
