datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LiveMCPBench
LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
Benchmarking the agent in real-world tasks within a large-scale MCP toolset.
🌐 Website |
📄 Paper |
💻 Code |
🏆 Leaderboard
|
🙏 Citation
Dataset Description
LiveMCPBench is the first comprehensive benchmark designed to evaluate LLM agents at scale across diverse Model Context Protocol (MCP) servers. It comprises 95 real-world tasks grounded in the MCP ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/ICIP/LiveMCPBench.LiveMCPBench
LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
Benchmarking the agent in real-world tasks within a large-scale MCP toolset.
🌐 Website |
📄 Paper |
💻 Code |
🏆 Leaderboard
|
🙏 Citation
Dataset Description
LiveMCPBench is the first comprehensive benchmark designed to evaluate LLM agents at scale across diverse Model Context Protocol (MCP) servers. It comprises 95 real-world tasks grounded in the MCP ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Sivachandran5710/LiveMCPBench.
