datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ii-agent_gaia-benchmark_validationcoding-agent-security-benchmark
Coding Agent Security Benchmark
A benchmark for evaluating whether an LLM can correctly identify security
violations in the behavior of an autonomous coding agent - spanning
dangerous shell commands, credential leakage, prompt injection, supply-chain
risk, privacy leaks, and more.
Each row is a single message sampled from a coding-agent session (a user
instruction, a tool call the agent issued, a tool's response, or the agent's
own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/ruchit11111/coding-agent-security-benchmark.coding-agent-security-benchmark
Coding Agent Security Benchmark
A benchmark for evaluating whether an LLM can correctly identify security
violations in the behavior of an autonomous coding agent - spanning
dangerous shell commands, credential leakage, prompt injection, supply-chain
risk, privacy leaks, and more.
Each row is a single message sampled from a coding-agent session (a user
instruction, a tool call the agent issued, a tool's response, or the agent's
own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/coding-agent-security-benchmark.sage-agent-benchmark
SAGE Agent Benchmark
Comprehensive benchmark for evaluating AI agent capabilities across three core competencies:
Tool Selection - Choosing appropriate tools for tasks
Task Planning - Decomposing complex tasks into step sequences
Timing Judgment - Deciding when to use tools vs. direct answers
Dataset Statistics
Total Samples: ~11,000
Tool Selection: ~6,000 samples
Task Planning: ~3,000 samples
Timing Judgment: ~2,000 samples
Splits: train, dev, test
Usage… See the full description on the dataset page: https://huggingface.co/datasets/intellistream/sage-agent-benchmark.mini_rm_benchmark_for_web_agent
Dataset Card for "mini_rm_benchmark_for_web_agent"
More Information needed
qcea-adaptive-agent-benchmark
QCEA Adaptive Agent Benchmark: The Dancing Landscape
Description: The Dancing Landscape. A multi-regime dataset for stress-testing Universal Agents against the laws of Entropic Decay and Computational Irreducibility.
Maintainer: AlgoplexityResearch Horizon: Horizon 2 (Adaptive Strategy)
1. Overview
This repository contains the Spatial-Causal State Vectors required to train and validate the AIT Physicist in a multi-agent environment.
It serves as the "Petri Dish" for the… See the full description on the dataset page: https://huggingface.co/datasets/algoplexity/qcea-adaptive-agent-benchmark.hf-agent-benchmark
