CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Intelligent-Internet /ii-agent_gaia-benchmark_validationtextn<1K8 likes929 downloads1y agoHugging Face02ruchit11111 /coding-agent-security-benchmark Coding Agent Security Benchmark A benchmark for evaluating whether an LLM can correctly identify security violations in the behavior of an autonomous coding agent - spanning dangerous shell commands, credential leakage, prompt injection, supply-chain risk, privacy leaks, and more. Each row is a single message sampled from a coding-agent session (a user instruction, a tool call the agent issued, a tool's response, or the agent's own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/ruchit11111/coding-agent-security-benchmark.textn<1K1 likes332 downloads25d agoHugging Face03rogue-security /coding-agent-security-benchmark Coding Agent Security Benchmark A benchmark for evaluating whether an LLM can correctly identify security violations in the behavior of an autonomous coding agent - spanning dangerous shell commands, credential leakage, prompt injection, supply-chain risk, privacy leaks, and more. Each row is a single message sampled from a coding-agent session (a user instruction, a tool call the agent issued, a tool's response, or the agent's own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/coding-agent-security-benchmark.textn<1K2 likes300 downloads2mo agoHugging Face04intellistream /sage-agent-benchmark SAGE Agent Benchmark Comprehensive benchmark for evaluating AI agent capabilities across three core competencies: Tool Selection - Choosing appropriate tools for tasks Task Planning - Decomposing complex tasks into step sequences Timing Judgment - Deciding when to use tools vs. direct answers Dataset Statistics Total Samples: ~11,000 Tool Selection: ~6,000 samples Task Planning: ~3,000 samples Timing Judgment: ~2,000 samples Splits: train, dev, test Usage… See the full description on the dataset page: https://huggingface.co/datasets/intellistream/sage-agent-benchmark.textquestion-answering10K<n<100K1 likes45 downloads8mo agoHugging Face05LangAGI-Lab /mini_rm_benchmark_for_web_agent Dataset Card for "mini_rm_benchmark_for_web_agent" More Information needed imagen<1K0 likes33 downloads2y agoHugging Face06algoplexity /qcea-adaptive-agent-benchmark QCEA Adaptive Agent Benchmark: The Dancing Landscape Description: The Dancing Landscape. A multi-regime dataset for stress-testing Universal Agents against the laws of Entropic Decay and Computational Irreducibility. Maintainer: AlgoplexityResearch Horizon: Horizon 2 (Adaptive Strategy) 1. Overview This repository contains the Spatial-Causal State Vectors required to train and validate the AIT Physicist in a multi-agent environment. It serves as the "Petri Dish" for the… See the full description on the dataset page: https://huggingface.co/datasets/algoplexity/qcea-adaptive-agent-benchmark.tabularreinforcement-learning10K<n<100K0 likes25 downloads9mo agoHugging Face07akseljoonas /hf-agent-benchmarktextn<1K0 likes23 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.