datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tenacious-bench-v0.1
Tenacious-Bench v0.1
A style-compliance evaluation benchmark for B2B sales AI agents
Author: Gashaw Bekele | gashaw@10academy.org
Built for TRP1 Week 11 — Sales Agent Evaluation Bench challenge
Code: https://github.com/gashawbekele06/tenacious-bench
What This Is
Tenacious-Bench evaluates AI sales agents on failure modes that public benchmarks
(τ²-Bench, AgentBench) miss: tone preservation, hiring-signal grounding, bench
commitment accuracy, and discovery-call… See the full description on the dataset page: https://huggingface.co/datasets/gashawbekele/tenacious-bench-v0.1.tenacious-bench-v0.1
Tenacious-Bench v0.1
A domain-specific evaluation benchmark for B2B sales agents, testing failure modes
that general-purpose benchmarks (τ²-Bench, AgentBench) do not measure.
Dataset Summary
218 tasks across 3 splits, covering 10 Tenacious-specific failure categories:
Split
Tasks
train
109 (50%)
dev
65 (30%)
held_out
44 (20%)
Source Mode
Count
%
Programmatic
108
49.5%
Multi-LLM Synthesis
55
25.2%
Trace-derived
35
16.1%
Hand-authored… See the full description on the dataset page: https://huggingface.co/datasets/mamaru13/tenacious-bench-v0.1.tenacious-bench-v0.1Tenacious-Bench-v0.1tenacious-bench-v0.1tenacious-bench-v0.1
