agentic-benchmark
agentic-redteam-benchmark
agentic-redteam-benchmark
v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented.
A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful.
📦 Code, eval harness & issues: github.com/Alkur123/agentic-redteam-benchmark · 📄 Paper: A Per-Step Trajectory Benchmark for AI-Agent Governance Verifiers and a Corrected Catch-at-Drift Metric (Aegis AI, 2026)… See the full description on the dataset page: https://huggingface.co/datasets/jash-ai/agentic-redteam-benchmark.ai-research-berkeley-agentic-verification-benchmarkInternal documentation
Analysis and results
Consolidated English analysis · Raw data · Figures
The report separates evaluation tasks and exact group memberships. Each plot appears above its data table.
Current trained MV results cover 12 settings with 99 groups each; published baselines use a separate 99-group cohort.
A matched MV vs. fixed-baseline comparison uses the same 65 DeepSWE groups. The full 98-group intersection is documented; baseline episode records are needed to… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-agentic-verification-benchmark.repro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces
Agent traces
Agent sessions published from a Trackio Logbook.
agentic-task-benchmark
YouMind Agentic Task Benchmark
v0.1.1 · Experimental
YouMindInc/agentic-task-benchmark is a small experimental benchmark for
evaluating creative and research task outcomes against selected references.
This release contains four task descriptions, a standard outcome format, and
an offline text scorer. Task definitions and evaluation protocols may change.
Tasks
Configuration
Task
Status
image_generation
Riso portrait
Defined; reference image supplied… See the full description on the dataset page: https://huggingface.co/datasets/YouMindInc/agentic-task-benchmark.agentic-reasoning-benchmark
Agentic & Reasoning Benchmark (ARB) – Expanded
Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning.
Überblick
Eigenschaft
Wert
Anzahl Beispiele
2.550
Kategorien
8
Schwierigkeitsgrade
easy / medium / hard
Formate
CSV + JSON
Reproduzierbarkeit
Generator-Skript (seed=42) enthalten
Lizenz
CC-BY-4.0
Kategorien
Kategorie
Anzahl
Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.smoke24-agentic-benchmarks
Smoke24 Agentic Benchmarks
Public, reproducible Terminal-Bench 2.0 Smoke24 benchmark artifacts for local RTX 3090-class agentic model serving.
Why Smoke24
I created the Smoke24 subset because I needed a relatively quick benchmark that
could run locally on an RTX 3090-class machine in a couple of hours. The goal is
to get a practical read on model performance and stability under a real agentic
Terminal-Bench workload, without paying the turnaround cost of a much… See the full description on the dataset page: https://huggingface.co/datasets/XReyRobert/smoke24-agentic-benchmarks.
