tenacious-bench
tenacious-bench
Tenacious-Bench v0.1
A domain-specific evaluation benchmark for B2B sales outreach agents, targeting failure modes that τ²-Bench retail cannot grade.
Dataset Summary
Tenacious-Bench v0.1 is a 293-task evaluation dataset for testing AI sales agents on Tenacious Intelligence Corporation-style B2B outreach. Every task is machine-verifiable: a script reads a task plus an agent output and returns a numerical score with no human in the loop.
Key Numbers
Metric… See the full description on the dataset page: https://huggingface.co/datasets/yakobd/tenacious-bench.tenacious-bench-v0.1
Datasheet — Tenacious-Bench v0.1
A datasheet for the Tenacious-Bench v0.1 evaluation dataset, following Gebru et al. (2021) Datasheets for Datasets and extending with Pushkarna et al. (2022) Data Cards layered detail.
Layered detail (Pushkarna):
Telescopic — what this dataset is and who should care, in three sentences.
Periscopic — composition, collection, and intended use at the depth a reviewer needs to decide whether to use it.
Microscopic — exact field semantics, scoring… See the full description on the dataset page: https://huggingface.co/datasets/nahdes/tenacious-bench-v0.1.tenacious-bench-v0.1
Tenacious-Bench v0.1
A domain-specific evaluation benchmark for B2B outbound sales agents. Grades five failure modes that standard benchmarks (τ²-Bench retail) cannot measure.
300 tasks · 5 rubric dimensions · 4 authoring modes · CC-BY-4.0
Why This Benchmark Exists
τ²-Bench retail tests cooperative airline-service policy transactions. B2B outbound sales requires:
Dimension
Trigger Rate (Week 10 probes)
Commercial Risk
signal_grounding_fidelity
35%
CTO… See the full description on the dataset page: https://huggingface.co/datasets/rafiakedir/tenacious-bench-v0.1.tenacious-bench-v0.1
Tenacious-Bench v0.1 — B2B Sales Agent Evaluation Benchmark
A domain-specific evaluation benchmark for B2B sales agents, grounded in
Tenacious Intelligence Corporation's ICP segments, signal enrichment pipeline,
and tone requirements.
Built on top of Week 10: github.com/Meseretbolled/conversion-engine
What This Is
τ²-Bench retail cannot grade Tenacious-specific failure modes — it scores retail
transaction completion. It has no concept of signal confidence thresholds, ICP… See the full description on the dataset page: https://huggingface.co/datasets/meseretbolled/tenacious-bench-v0.1.tenacious-bench
Tenacious-Bench v0.1
Tenacious-specific evaluation benchmark for B2B sales-agent output quality.
Version v0.1.0-interim — 44 tasks; v0.1.0 final scales to 240.
Headline result (Delta A): +0.2188 lift on held-out (95 % CI [+0.1177, +0.3198], p < 0.0001).
Companion model adapter: bonneyjr/tenacious-judge-lora-v0.1
Code + reproduction scripts: github.com/atnabon/sales-eval-bench
What this bench grades
τ²-Bench retail tells you whether a sales agent works in… See the full description on the dataset page: https://huggingface.co/datasets/bonneyjr/tenacious-bench.tenacious-bench
Tenacious-Bench v0.1
Sales agent evaluation benchmark for B2B outbound honesty and tone compliance.
274 tasks (250 programmatic + 24 hand-authored). Machine-verifiable rubric.
Author: Eyobed Feleke
License: CC-BY-4.0
Model: eyobed7b/tenacious-bench-simpo-judge-v1
