datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ContextASR-Bench
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.ambient-acoustic-context
Dataset Card for Ambient Acoustic Context
The Ambient Acoustic Context dataset contains 1-second segments for activities that occur in a workplace setting. Each segment is associated with speaker_id.
Dataset Details
Using Amazin Mechanical Turk, crowd workers were asked to listen to 1-second segments and choose the right label. To ensure the quality of the annotations, audio segments that did not reach majority agreement among the turkers were excluded.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/ambient-acoustic-context.acoustic_context_switchingthai-contextasr-bench
Thai Contextual-Biasing ASR Benchmark
TL;DR
Does your Thai ASR system actually use the context you give it (e.g., a list of names, custom words from your own dictionary) — and does it hallucinate when the context is irrelevant?
Each utterance comes with a bias list: entity strings (brands, person names,
places) that may or may not be spoken in the audio, written the way a real Thai user
would write them — one list, mixed Thai and Latin script. A good system does… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-contextasr-bench.MM-ContextASR-Bench
MM-ContextASR Bench
Metadata and evaluation splits for Multimodal Conversational Context for
LLM-Based ASR: Data Construction, Training, and Benchmark.
Dataset summary
Config
Examples
Audio
Context
Primary metric
mm_contextasr
1,250 (250 current utterances × 5 histories)
1,439 WAV files included
Controlled user-assistant dialogue
entity Recall
kespeech
19,212
Source ID only
Same-speaker speech and transcript
CER, SER, entity Recall
cv_yue
3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.contextual-earnings22earnings22-keywords
Annotation of this dataset is still in progress. Argmax will publish a conference paper on the annotation process later in 2025.
The license for this dataset is the same as that of the original dataset.
ePark_qing_jing_zu_yu_contextual_indigenous_language
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language.expresso-contextual
The Expresso Dataset
[paper] [demo samples] [Original repository]
Introduction
The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised).
Improvised Dialogues
This HuggingFace dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Zackh/expresso-contextual.ContextTTS_dataset
ContextTTS Evaluation Dataset
This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form
Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks.
Dataset Summary
The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.contextual_asr_benchmark
Synthetic Contextual ASR Benchmark (Indic)
Dataset Summary
This dataset is a Synthetic Contextual Automatic Speech Recognition (ASR) benchmark designed to evaluate and improve speech recognition systems in voice bot scenarios. It focuses on context-aware transcription, where the ASR model can leverage conversation history and agent prompts to better transcribe user responses.
The dataset covers the top 10 Indian languages, providing a diverse linguistic landscape for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/contextual_asr_benchmark.ContextDialog
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
🎉 We are excited to announce that our paper has been accepted to the Findings of ACL 2025!
Demo Page: https://contextdialog.github.io
arXiv: https://arxiv.org/abs/2502.19759
ContextDialog is a comprehensive benchmark designed to evaluate a voice interaction model’s ability to engage in, retain, and leverage relevant information throughout multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/ContextDialog/ContextDialog.Sample-Voice-Context-Data
Sample Voice Context Data
A small synthetic dataset containing LLM-generated context information simulating a job seeker narrating their career trajectory.
Purpose
This dataset was created to test a voice-to-vector-database RAG pipeline. The workflow being evaluated involves:
Voice data (MP3 recordings) transcribed to text
Transcriptions reformatted as structured context data
Text data upserted into a vector database (Pinecone or Ragie)
Retrieval accuracy tested by… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Sample-Voice-Context-Data.asr-context-induced-leakage
When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR
Overview
SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts, fine-tune on proprietary recordings, or both. We identify and systematically investigate an overlooked privacy risk of such customisation: a model adapted to recognise domain-specific terminology can be nudged into transcribing a phonetically similar… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/asr-context-induced-leakage.ContextASR-Bench
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/bsmu666/ContextASR-Bench.violence_contextturntaking-contextual-ttscontextual_asr_benchmark
Synthetic Contextual ASR Benchmark (Indic)
Dataset Summary
This dataset is a Synthetic Contextual Automatic Speech Recognition (ASR) benchmark designed to evaluate and improve speech recognition systems in voice bot scenarios. It focuses on context-aware transcription, where the ASR model can leverage conversation history and agent prompts to better transcribe user responses.
The dataset covers the top 10 Indian languages, providing a diverse linguistic landscape for… See the full description on the dataset page: https://huggingface.co/datasets/mygitphase/contextual_asr_benchmark.contextual_asr_benchmark
Synthetic Contextual ASR Benchmark (Indic)
Dataset Summary
This dataset is a Synthetic Contextual Automatic Speech Recognition (ASR) benchmark designed to evaluate and improve speech recognition systems in voice bot scenarios. It focuses on context-aware transcription, where the ASR model can leverage conversation history and agent prompts to better transcribe user responses.
The dataset covers the top 10 Indian languages, providing a diverse linguistic landscape for… See the full description on the dataset page: https://huggingface.co/datasets/benkum/contextual_asr_benchmark.libri_clean_long_contextin_context_QA_ASR_TTS_finetune_3-2-11B_rank64_ls960_replay_v4speech-accent-context2138 English speakers from a variety of language backgrounds (pulled from the Speech Accent Archive) saying the sentence
"Please call Stella. Ask her to bring these things with her from the store: Six spoons of fresh snow peas, five thick slabs of blue cheese, and maybe a snack for her brother Bob. We also need a small plastic snake and a big toy frog for the kids. She can scoop these things into three red bags, and we will go meet her Wednesday at the train station."
We focus on three… See the full description on the dataset page: https://huggingface.co/datasets/KoelLabs/speech-accent-context.ambient-acoustic-context-smallContextConsistencyChecks
Audio Consistency Checks Dataset
This dataset contains audio-based ambiguity-resolution tasks across prosody categories.
Columns:
row_id: unique row id
group_id: _
category: category name
audio: relative path to the single combined input audio per row. Order: target separator + target, then spoken separators and three completion audios in randomized A/B/C order.
correct_completion: one of "Completion A", "Completion B", "Completion C"
input: text of the input (target) sentence… See the full description on the dataset page: https://huggingface.co/datasets/StevenDillmann/ContextConsistencyChecks.ContextDialog
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
🎉 We are excited to announce that our paper has been accepted to the Findings of ACL 2025!
Demo Page: https://contextdialog.github.io
arXiv: https://arxiv.org/abs/2502.19759
ContextDialog is a comprehensive benchmark designed to evaluate a voice interaction model’s ability to engage in, retain, and leverage relevant information throughout multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/wangyueyiiiiiii/ContextDialog.in_context_ASR_TTS_finetune_3-2-11B_rank64_ls960_replay_v5Contextual_Audio_cuesambient-acoustic-context-small
