datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mcp-tool-calling-benchmark
MCP Tool-Calling Benchmark
A benchmark dataset for evaluating AI assistants' MCP (Model Context Protocol) tool-calling accuracy across 12 platforms.
Dataset Description
Contains 6,451 labeled interaction logs from systematic QA testing of Grok's MCP connectors. Each row captures a test prompt, the expected tool invocation, Grok's actual response, and the error classification.
Platforms Covered
Platform
Prompts
Tools Tested
Primary Error Pattern… See the full description on the dataset page: https://huggingface.co/datasets/brijeshvadi/mcp-tool-calling-benchmark.chadgpt-tool-calling-150
