CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01obaydata /mcp-agent-trajectory-benchmark MCP Agent Trajectory Benchmark A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces. Designed for training and evaluating tool-use / function-calling capabilities of LLMs. Overview Item Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.texttext-generationn<1K3 likes1k downloads6mo agoHugging Face02Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K6 likes904 downloads11mo agoHugging Face03noel7Y /data-agent-benchmarks LongHorizon Full Data-Agent Benchmarks Companion data artifacts for five complete evaluation tracks: DataSciBench full55 / 167 metric entries DABStep full450 DABStep-Research full100 DSBench Modeling full74 LongDS full68 / 2,225 turns The companion GitHub repository contains processed manifests, evaluation code, historical API ReAct baseline code, download/preparation tools, and the frozen source lock. artifact_manifest.json records every uploaded object's size, SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.imagequestion-answering1 likes402 downloads13d agoHugging Face04obaydata /claude-agent-skills-benchmark Claude Agent Skills Benchmark Claude Agent Skills 评测数据集 Description A benchmark dataset for evaluating whether LLMs can accurately trigger and execute domain-specific Skills on the Claude Code platform. Skills are designed by vertical domain experts with varying complexity levels (based on attachments: scripts, references, assets, and reference markdown files). Evaluation Scenarios Cover: Office automation, coding, investment promotion, financial services, industrial… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/claude-agent-skills-benchmark.documenttext-generationn<1K1 likes215 downloads6mo agoHugging Face05NotoriousH2 /calendar-agent-benchmark Calendar Agent Training and Evaluation Tool Calling SFT와 환경 기반 RLVR을 함께 실습하는 Calendar Agent 데이터입니다. 1. 데이터 구성 파일 개수 구성 SHA-256 sft_train.jsonl 3,000 기본 시나리오, 다중 참석자 생성, 안전 대조군 f882fba3a6c32982d2f8cca06465fa1ec3c68562c655d2ee1bcd92decd29e637 sft_validation.jsonl 300 SFT 학습 중 검증 b657e4162ef26bd580ad975b4f2af575a616dda17c90c7b4edab38790ad66ac5 rlvr_train.jsonl 256 충돌 복구와 다중 참석자 생성 128개, 안전 경계 128개… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/calendar-agent-benchmark.text-generation1 likes170 downloads1mo agoHugging Face06springofwindslabs /mcp-agent-trajectory-benchmark ⚡ Model Context Protocol (MCP) & Advanced Tool‑Use Alignment Tiers 15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%). Schema Validation Summary Programmatic validation of this exact trial file - reproducible from data.jsonl. What this trial verifies — use these 50 rows to confirm, on your own stack: Schema integrity (strict JSONL, matches the published schema) Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.texttext-generationn<1K0 likes137 downloads1d agoHugging Face07aiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes57 downloads6mo agoHugging Face08intellistream /sage-agent-benchmark SAGE Agent Benchmark Comprehensive benchmark for evaluating AI agent capabilities across three core competencies: Tool Selection - Choosing appropriate tools for tasks Task Planning - Decomposing complex tasks into step sequences Timing Judgment - Deciding when to use tools vs. direct answers Dataset Statistics Total Samples: ~11,000 Tool Selection: ~6,000 samples Task Planning: ~3,000 samples Timing Judgment: ~2,000 samples Splits: train, dev, test Usage… See the full description on the dataset page: https://huggingface.co/datasets/intellistream/sage-agent-benchmark.textquestion-answering10K<n<100K1 likes42 downloads8mo agoHugging Face09ScareRezume /agent-sandbox-negotiation-benchmark Agent Sandbox Negotiation Benchmark v1 Overview A dataset of simulated multi-agent negotiations generated using the open-source Agent Sandbox framework. This dataset captures the final negotiation outcomes, turn depths, strategy alignments, and agreed prices of local LLMs (Llama-3 and Mistral) engaged in intense, adversarial price negotiations at massive scale. Dataset Statistics Simulations: 24,122 Strategies: 4 (Balanced, Aggressive, Conservative, Adaptive)… See the full description on the dataset page: https://huggingface.co/datasets/ScareRezume/agent-sandbox-negotiation-benchmark.tabulartext-generation10K<n<100K1 likes22 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.