CoolFace
15 results

tenacious-bench

yakobd /tenacious-bench Tenacious-Bench v0.1 A domain-specific evaluation benchmark for B2B sales outreach agents, targeting failure modes that τ²-Bench retail cannot grade. Dataset Summary Tenacious-Bench v0.1 is a 293-task evaluation dataset for testing AI sales agents on Tenacious Intelligence Corporation-style B2B outreach. Every task is machine-verifiable: a script reads a task plus an agent output and returns a numerical score with no human in the loop. Key Numbers Metric… See the full description on the dataset page: https://huggingface.co/datasets/yakobd/tenacious-bench.text-generationn<1K0 likes117 downloads5mo agoHugging Facenahdes /tenacious-bench-v0.1 Datasheet — Tenacious-Bench v0.1 A datasheet for the Tenacious-Bench v0.1 evaluation dataset, following Gebru et al. (2021) Datasheets for Datasets and extending with Pushkarna et al. (2022) Data Cards layered detail. Layered detail (Pushkarna): Telescopic — what this dataset is and who should care, in three sentences. Periscopic — composition, collection, and intended use at the depth a reviewer needs to decide whether to use it. Microscopic — exact field semantics, scoring… See the full description on the dataset page: https://huggingface.co/datasets/nahdes/tenacious-bench-v0.1.text-generationn<1K0 likes90 downloads5mo agoHugging Facerafiakedir /tenacious-bench-v0.1 Tenacious-Bench v0.1 A domain-specific evaluation benchmark for B2B outbound sales agents. Grades five failure modes that standard benchmarks (τ²-Bench retail) cannot measure. 300 tasks · 5 rubric dimensions · 4 authoring modes · CC-BY-4.0 Why This Benchmark Exists τ²-Bench retail tests cooperative airline-service policy transactions. B2B outbound sales requires: Dimension Trigger Rate (Week 10 probes) Commercial Risk signal_grounding_fidelity 35% CTO… See the full description on the dataset page: https://huggingface.co/datasets/rafiakedir/tenacious-bench-v0.1.texttext-generationn<1K0 likes85 downloads5mo agoHugging Facemeseretbolled /tenacious-bench-v0.1 Tenacious-Bench v0.1 — B2B Sales Agent Evaluation Benchmark A domain-specific evaluation benchmark for B2B sales agents, grounded in Tenacious Intelligence Corporation's ICP segments, signal enrichment pipeline, and tone requirements. Built on top of Week 10: github.com/Meseretbolled/conversion-engine What This Is τ²-Bench retail cannot grade Tenacious-specific failure modes — it scores retail transaction completion. It has no concept of signal confidence thresholds, ICP… See the full description on the dataset page: https://huggingface.co/datasets/meseretbolled/tenacious-bench-v0.1.0 likes58 downloads5mo agoHugging Facebonneyjr /tenacious-bench Tenacious-Bench v0.1 Tenacious-specific evaluation benchmark for B2B sales-agent output quality. Version v0.1.0-interim — 44 tasks; v0.1.0 final scales to 240. Headline result (Delta A): +0.2188 lift on held-out (95 % CI [+0.1177, +0.3198], p < 0.0001). Companion model adapter: bonneyjr/tenacious-judge-lora-v0.1 Code + reproduction scripts: github.com/atnabon/sales-eval-bench What this bench grades τ²-Bench retail tells you whether a sales agent works in… See the full description on the dataset page: https://huggingface.co/datasets/bonneyjr/tenacious-bench.text-generationn<1K0 likes50 downloads5mo agoHugging Faceeyobed7b /tenacious-bench Tenacious-Bench v0.1 Sales agent evaluation benchmark for B2B outbound honesty and tone compliance. 274 tasks (250 programmatic + 24 hand-authored). Machine-verifiable rubric. Author: Eyobed Feleke License: CC-BY-4.0 Model: eyobed7b/tenacious-bench-simpo-judge-v1 texttext-classificationn<1K0 likes38 downloads5mo agoHugging Face