datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tenacious-bench-v0.1
Tenacious-Bench v0.1
A domain-specific evaluation benchmark for B2B outbound sales agents. Grades five failure modes that standard benchmarks (τ²-Bench retail) cannot measure.
300 tasks · 5 rubric dimensions · 4 authoring modes · CC-BY-4.0
Why This Benchmark Exists
τ²-Bench retail tests cooperative airline-service policy transactions. B2B outbound sales requires:
Dimension
Trigger Rate (Week 10 probes)
Commercial Risk
signal_grounding_fidelity
35%
CTO… See the full description on the dataset page: https://huggingface.co/datasets/rafiakedir/tenacious-bench-v0.1.tenacious_bench_v0.1
Tenacious-Bench v0.1
A domain-specific evaluation benchmark for B2B sales AI agents in the technical staffing vertical. Tenacious-Bench tests whether an AI outreach agent is honest, proportionate, and professionally grounded — not just whether it completed a task.
Quickstart
git clone https://huggingface.co/datasets/Yohannesdn/tenacious_bench_v0.1
pip install jsonschema
python scoring_evaluator.py --schema schema.json --example 1
Run python scoring_evaluator.py --schema… See the full description on the dataset page: https://huggingface.co/datasets/Yohannesdn/tenacious_bench_v0.1.tenacious_bench_v0.1
Tenacious-Bench v0.1
A 266-task evaluation benchmark for B2B sales-outreach agents, grounded in the
Tenacious (B2B engineering-outsourcing) workflow. 10 failure dimensions, 5
authoring source modes, mechanically-gradable rubric (no human in the loop).
Held-out partition is sealed — see held_out/SEALED.md for details.
Dataset summary
Total tasks (train + dev): 191
Held-out (sealed): 75
Source modes: 5
Failure dimensions: 10
License: CC-BY-4.0
Quickstart
from… See the full description on the dataset page: https://huggingface.co/datasets/eyorata/tenacious_bench_v0.1.tenacious-bench
📊 Tenacious-Bench: B2B Sales Outreach Judge Preference Dataset
Version: v0.1Author: Bethelhem Abay · 10 Academy TRP1Date: 2026-05-02License: MIT
A curated preference dataset of 323 (chosen, rejected) pairs for training and evaluating a pre-send judge that blocks unsafe B2B sales outreach before it reaches the wrong people.
🔗 Quick Links
Resource
Link
📦 Dataset (this page)
bethelhem21/tenacious-bench
🤖 Trained Judge Model… See the full description on the dataset page: https://huggingface.co/datasets/bethelhem21/tenacious-bench.tenacious-bench-v0.1
Tenacious-Bench v0.1
Tenacious-Bench v0.1 is a small, contamination-checked benchmark for Tenacious-style B2B sales outreach. It measures whether an agent can ground outreach in public hiring signals, stay within truthful capacity and pricing boundaries, preserve the Tenacious tone markers, and avoid brand-damaging failure modes that generic retail-style agent benchmarks do not measure.
What is in the dataset
train split: task authoring and training-pair construction
dev… See the full description on the dataset page: https://huggingface.co/datasets/abdulaziz0111/tenacious-bench-v0.1.tenacious-bench
Tenacious-Bench v0.1
A specialized benchmark for B2B sales agent evaluation
Tenacious-Bench measures five critical dimensions that existing benchmarks (τ²-Bench retail, WebArena, BrowseComp) do not capture: signal grounding, capacity honesty, tone preservation, consent-first coordination, and gap framing. These dimensions map directly to the highest-cost failure modes observed in the Tenacious Conversion Engine (Week 10 evidence).
⚙️ Setup
Requirements… See the full description on the dataset page: https://huggingface.co/datasets/samuellachisa/tenacious-bench.tenacious-bench
Tenacious-Bench v0.2
A domain-specific evaluation benchmark for B2B sales outreach agents. Measures compliance with structured business rules that generic benchmarks (τ²-Bench retail, etc.) cannot grade: bench capacity constraints, ICP segment routing, signal confidence calibration, and tone compliance.
Why this benchmark exists
The Week 10 Tenacious agent failed every bench-capacity probe (40/40, 100%) and misrouted 54% of ICP classification tasks — while producing… See the full description on the dataset page: https://huggingface.co/datasets/Natnaela/tenacious-bench.tenacious-bench-v0.1
Tenacious-Bench v0.1
A style-compliance evaluation benchmark for B2B sales AI agents
Author: Gashaw Bekele | gashaw@10academy.org
Built for TRP1 Week 11 — Sales Agent Evaluation Bench challenge
Code: https://github.com/gashawbekele06/tenacious-bench
What This Is
Tenacious-Bench evaluates AI sales agents on failure modes that public benchmarks
(τ²-Bench, AgentBench) miss: tone preservation, hiring-signal grounding, bench
commitment accuracy, and discovery-call… See the full description on the dataset page: https://huggingface.co/datasets/gashawbekele/tenacious-bench-v0.1.tenacious-bench-v0.1
Tenacious-Bench v0.1
A domain-specific evaluation benchmark for B2B sales agents, testing failure modes
that general-purpose benchmarks (τ²-Bench, AgentBench) do not measure.
Dataset Summary
218 tasks across 3 splits, covering 10 Tenacious-specific failure categories:
Split
Tasks
train
109 (50%)
dev
65 (30%)
held_out
44 (20%)
Source Mode
Count
%
Programmatic
108
49.5%
Multi-LLM Synthesis
55
25.2%
Trace-derived
35
16.1%
Hand-authored… See the full description on the dataset page: https://huggingface.co/datasets/mamaru13/tenacious-bench-v0.1.tenacious_bench_v0.1
TenaciousBench v0.1
A 220-task evaluation benchmark for B2B outbound sales agents.
TenaciousBench v0.1 evaluates LLM-based B2B outreach systems across ten failure dimensions: ICP targeting precision, confidence-aware phrasing, signal grounding fidelity, tone safety, hallucination avoidance, CTA behavior, competitor gap reasoning, pricing discipline, multi-turn objection handling, and thread continuation coherence.
Why This Benchmark Exists
Existing LLM benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/eyorg/tenacious_bench_v0.1.tenacious-bench-path-b-preference
Tenacious Bench Path B Preference
1. Motivation
This dataset exists because generic assistant benchmarks do not reliably measure the failure modes that matter in Tenacious-style B2B outbound work. The goal here is not broad conversational quality; it is grounded business behavior under uncertainty.
The preference pairs focus on:
grounded language
weak-confidence handling
over-claiming avoidance
pricing handoff safety
qualification correctness
channel routing… See the full description on the dataset page: https://huggingface.co/datasets/ephorata/tenacious-bench-path-b-preference.tenacious-bench-v0.1
Tenacious-Bench v0.1
A 200-task evaluation benchmark for B2B sales outreach agents. Measures five failure dimensions that general-purpose benchmarks (τ²-Bench retail) cannot detect.
Why this benchmark exists
The Tenacious AI sales agent achieves 38.7% pass@1 on τ²-Bench retail — but the failures are invisible to that rubric. τ²-Bench has no bench inventory model, no style guide, and no signal confidence grading. This benchmark was built specifically to catch what τ²-Bench… See the full description on the dataset page: https://huggingface.co/datasets/ketewodros41/tenacious-bench-v0.1.
