datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grocery-bench
Grocery Bench
30-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a grocery ordering assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a grocery ordering assistant helping a customer build, modify, and finalize an order. The conversation is designed around 15 difficulty enhancements that… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/grocery-bench.product-bench
Product Bench
31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a laptop comparison shopping assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a laptop comparison shopping assistant helping a customer evaluate, compare, and order laptops. The conversation features multi-intent turns… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/product-bench.humans-benchmark
HUMANS Benchmark Dataset
Authors: Woody Haosheng Gan¹, William Held²'³, Diyi Yang²
¹University of Southern California, ²Stanford University, ³OpenAthena
This dataset is part of the Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment paper.
HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark is designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned… See the full description on the dataset page: https://huggingface.co/datasets/woodygan/humans-benchmark.appointment-bench
Appointment Bench
25-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a dental office receptionist handling appointment scheduling.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a dental office receptionist scheduling appointments for two patients with confusable names (Daniel and Danielle Nolan)… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/appointment-bench.conversation-bench
Conversation Bench
75-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a conference assistant for the AI Engineer World's Fair.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a conference assistant for the AI Engineer World's Fair, handling session registration, schedule queries, speaker lookups, and… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/conversation-bench.assistant-bench
Assistant Bench
31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a personal assistant handling flights, email, calendar, and reminders.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a personal assistant managing flight bookings, email composition, calendar events, and reminders. Turns include dual… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/assistant-bench.humans-benchmark
HUMANS Benchmark Dataset (Anonymous, Under Review)
This dataset is part of the HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark, designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned regression weights.
Installation
Install the HUMANS evaluation package from GitHub (our anonymous repo):
# Option 1: Install via pip
pip install… See the full description on the dataset page: https://huggingface.co/datasets/HUMANSBenchmark/humans-benchmark.event-bench
Event Bench
29-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as an event planning assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as an event planning assistant managing venue bookings, catering, and guest logistics. The conversation features cascading changes — a venue switch triggers catering… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/event-bench.
