CoolFace
11 results

agentic-benchmark

jash-ai /agentic-redteam-benchmark agentic-redteam-benchmark v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented. A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful. 📦 Code, eval harness & issues: github.com/Alkur123/agentic-redteam-benchmark · 📄 Paper: A Per-Step Trajectory Benchmark for AI-Agent Governance Verifiers and a Corrected Catch-at-Drift Metric (Aegis AI, 2026)… See the full description on the dataset page: https://huggingface.co/datasets/jash-ai/agentic-redteam-benchmark.texttext-classification1K<n<10K2 likes1.3k downloads23d agoHugging FaceMinjaeLee-FuriosaAI-Ext /ai-research-berkeley-agentic-verification-benchmarkgatedInternal documentation Analysis and results Consolidated English analysis · Raw data · Figures The report separates evaluation tasks and exact group memberships. Each plot appears above its data table. Current trained MV results cover 12 settings with 99 groups each; published baselines use a separate 99-group cohort. A matched MV vs. fixed-baseline comparison uses the same 65 DeepSWE groups. The full 98-group intersection is documented; baseline episode records are needed to… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-agentic-verification-benchmark.0 likes273 downloads6h agoHugging Faceretroam /repro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes110 downloads2mo agoHugging FaceYouMindInc /agentic-task-benchmark YouMind Agentic Task Benchmark v0.1.1 · Experimental YouMindInc/agentic-task-benchmark is a small experimental benchmark for evaluating creative and research task outcomes against selected references. This release contains four task descriptions, a standard outcome format, and an offline text scorer. Task definitions and evaluation protocols may change. Tasks Configuration Task Status image_generation Riso portrait Defined; reference image supplied… See the full description on the dataset page: https://huggingface.co/datasets/YouMindInc/agentic-task-benchmark.texttext-generationn<1K0 likes106 downloads14d agoHugging Faceroskosmos19 /agentic-reasoning-benchmark Agentic & Reasoning Benchmark (ARB) – Expanded Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning. Überblick Eigenschaft Wert Anzahl Beispiele 2.550 Kategorien 8 Schwierigkeitsgrade easy / medium / hard Formate CSV + JSON Reproduzierbarkeit Generator-Skript (seed=42) enthalten Lizenz CC-BY-4.0 Kategorien Kategorie Anzahl Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.textquestion-answering1K<n<10K1 likes102 downloads20d agoHugging FaceXReyRobert /smoke24-agentic-benchmarks Smoke24 Agentic Benchmarks Public, reproducible Terminal-Bench 2.0 Smoke24 benchmark artifacts for local RTX 3090-class agentic model serving. Why Smoke24 I created the Smoke24 subset because I needed a relatively quick benchmark that could run locally on an RTX 3090-class machine in a couple of hours. The goal is to get a practical read on model performance and stability under a real agentic Terminal-Bench workload, without paying the turnaround cost of a much… See the full description on the dataset page: https://huggingface.co/datasets/XReyRobert/smoke24-agentic-benchmarks.text-generation0 likes96 downloads26d agoHugging Face