CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ServiceNow-AI /AgentJudgeBench AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling A benchmark for systematically evaluating how reliably LLM judges assess agentic tool-calling workflows across structured, dependency-driven tasks. Why this benchmark? AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.tabularquestion-answering100K<n<1M0 likes461 downloads25d agoHugging Face02Omcrec /ecommerce-ai-data-analyst-agent-benchmark E-commerce AI Data Analyst Agent Benchmark A synthetic e-commerce dataset for evaluating AI data analyst agents on realistic, multi-step business analysis, data-quality investigation, and analytical reasoning. This dataset is part of the E-commerce AI Data Analyst Agent Benchmark. Dataset summary This dataset supports evaluation of AI data analyst agents on realistic, multi-step e-commerce analysis. It contains: customers.csv products.csv orders.csv returns.csv… See the full description on the dataset page: https://huggingface.co/datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark.tabulartable-question-answering1 likes78 downloads10d agoHugging Face03anote-ai /AgenticRag AgenticRAG-FP This repository describes AgenticRAG-FP, a research dataset and evaluation suite for studying how failures propagate through agentic retrieval-augmented generation pipelines. The dataset normalizes multi-hop QA examples into a common schema, runs real or mock ReAct-style RAG agents over them, injects controlled failures at specific retrieval/reasoning hops, and records whether diagnostic methods can recover the true root cause after the failure has propagated. The… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/AgenticRag.tabularquestion-answeringn<1K0 likes27 downloads1mo agoHugging Face04aiagentkarl /mcp-server-catalog MCP Server Catalog A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more. Overview This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use. Categories Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.tabulartext-generationn<1K1 likes9 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.