CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MrSupW /ContextASR-Bench ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.textautomatic-speech-recognition10K<n<100K38 likes1.7k downloads1y agoHugging Face02flwrlabs /ambient-acoustic-context Dataset Card for Ambient Acoustic Context The Ambient Acoustic Context dataset contains 1-second segments for activities that occur in a workplace setting. Each segment is associated with speaker_id. Dataset Details Using Amazin Mechanical Turk, crowd workers were asked to listen to 1-second segments and choose the right label. To ensure the quality of the annotations, audio segments that did not reach majority agreement among the turkers were excluded. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/ambient-acoustic-context.audioaudio-classification10K<n<100K5 likes471 downloads2y agoHugging Face03taejinp /acoustic_context_switchingaudio1K<n<10K1 likes368 downloads1y agoHugging Face04wayu-ai /thai-contextasr-bench Thai Contextual-Biasing ASR Benchmark TL;DR Does your Thai ASR system actually use the context you give it (e.g., a list of names, custom words from your own dictionary) — and does it hallucinate when the context is irrelevant? Each utterance comes with a bias list: entity strings (brands, person names, places) that may or may not be spoken in the audio, written the way a real Thai user would write them — one list, mixed Thai and Latin script. A good system does… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-contextasr-bench.audioautomatic-speech-recognition1K<n<10K1 likes337 downloads2mo agoHugging Face05lilonghao /MM-ContextASR-Bench MM-ContextASR Bench Metadata and evaluation splits for Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark. Dataset summary Config Examples Audio Context Primary metric mm_contextasr 1,250 (250 current utterances × 5 histories) 1,439 WAV files included Controlled user-assistant dialogue entity Recall kespeech 19,212 Source ID only Same-speaker speech and transcript CER, SER, entity Recall cv_yue 3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.audioautomatic-speech-recognition10K<n<100K1 likes196 downloads6d agoHugging Face06argmaxinc /contextual-earnings22earnings22-keywords Annotation of this dataset is still in progress. Argmax will publish a conference paper on the annotation process later in 2025. The license for this dataset is the same as that of the original dataset. audion<1K3 likes176 downloads3mo agoHugging Face07FormosanBank /ePark_qing_jing_zu_yu_contextual_indigenous_language FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language.audioautomatic-speech-recognition10K<n<100K0 likes137 downloads2mo agoHugging Face08Zackh /expresso-contextual The Expresso Dataset [paper] [demo samples] [Original repository] Introduction The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). Improvised Dialogues This HuggingFace dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Zackh/expresso-contextual.audiotext-to-speechn<1K3 likes127 downloads4mo agoHugging Face09rodenhhh /ContextTTS_dataset ContextTTS Evaluation Dataset This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks. Dataset Summary The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.audiotext-to-speechn<1K0 likes109 downloads6mo agoHugging Face10sarvamai /contextual_asr_benchmark Synthetic Contextual ASR Benchmark (Indic) Dataset Summary This dataset is a Synthetic Contextual Automatic Speech Recognition (ASR) benchmark designed to evaluate and improve speech recognition systems in voice bot scenarios. It focuses on context-aware transcription, where the ASR model can leverage conversation history and agent prompts to better transcribe user responses. The dataset covers the top 10 Indian languages, providing a diverse linguistic landscape for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/contextual_asr_benchmark.audio1K<n<10K4 likes93 downloads8mo agoHugging Face11ContextDialog /ContextDialog Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models 🎉 We are excited to announce that our paper has been accepted to the Findings of ACL 2025! Demo Page: https://contextdialog.github.io arXiv: https://arxiv.org/abs/2502.19759 ContextDialog is a comprehensive benchmark designed to evaluate a voice interaction model’s ability to engage in, retain, and leverage relevant information throughout multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/ContextDialog/ContextDialog.audioaudio-to-audio1K<n<10K2 likes68 downloads1y agoHugging Face12danielrosehill /Sample-Voice-Context-Data Sample Voice Context Data A small synthetic dataset containing LLM-generated context information simulating a job seeker narrating their career trajectory. Purpose This dataset was created to test a voice-to-vector-database RAG pipeline. The workflow being evaluated involves: Voice data (MP3 recordings) transcribed to text Transcriptions reformatted as structured context data Text data upserted into a vector database (Pinecone or Ragie) Retrieval accuracy tested by… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Sample-Voice-Context-Data.audion<1K0 likes63 downloads10mo agoHugging Face13maikezu /asr-context-induced-leakage When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR Overview SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts, fine-tune on proprietary recordings, or both. We identify and systematically investigate an overlooked privacy risk of such customisation: a model adapted to recognise domain-specific terminology can be nudged into transcribing a phonetically similar… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/asr-context-induced-leakage.audioautomatic-speech-recognition1K<n<10K0 likes60 downloads4mo agoHugging Face14bsmu666 /ContextASR-Bench ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/bsmu666/ContextASR-Bench.textautomatic-speech-recognition10K<n<100K0 likes54 downloads9mo agoHugging Face15carlicode /violence_contextaudion<1K1 likes39 downloads3y agoHugging Face16fixie-ai /turntaking-contextual-ttsaudion<1K4 likes29 downloads1y agoHugging Face17mygitphase /contextual_asr_benchmark Synthetic Contextual ASR Benchmark (Indic) Dataset Summary This dataset is a Synthetic Contextual Automatic Speech Recognition (ASR) benchmark designed to evaluate and improve speech recognition systems in voice bot scenarios. It focuses on context-aware transcription, where the ASR model can leverage conversation history and agent prompts to better transcribe user responses. The dataset covers the top 10 Indian languages, providing a diverse linguistic landscape for… See the full description on the dataset page: https://huggingface.co/datasets/mygitphase/contextual_asr_benchmark.audio1K<n<10K0 likes24 downloads5mo agoHugging Face18benkum /contextual_asr_benchmark Synthetic Contextual ASR Benchmark (Indic) Dataset Summary This dataset is a Synthetic Contextual Automatic Speech Recognition (ASR) benchmark designed to evaluate and improve speech recognition systems in voice bot scenarios. It focuses on context-aware transcription, where the ASR model can leverage conversation history and agent prompts to better transcribe user responses. The dataset covers the top 10 Indian languages, providing a diverse linguistic landscape for… See the full description on the dataset page: https://huggingface.co/datasets/benkum/contextual_asr_benchmark.audio1K<n<10K0 likes21 downloads5mo agoHugging Face19nc33 /libri_clean_long_contextaudio1K<n<10K0 likes16 downloads2y agoHugging Face20chiyuanhsiao /in_context_QA_ASR_TTS_finetune_3-2-11B_rank64_ls960_replay_v4audion<1K0 likes15 downloads2y agoHugging Face21KoelLabs /speech-accent-contextgated2138 English speakers from a variety of language backgrounds (pulled from the Speech Accent Archive) saying the sentence "Please call Stella. Ask her to bring these things with her from the store: Six spoons of fresh snow peas, five thick slabs of blue cheese, and maybe a snack for her brother Bob. We also need a small plastic snake and a big toy frog for the kids. She can scoop these things into three red bags, and we will go meet her Wednesday at the train station." We focus on three… See the full description on the dataset page: https://huggingface.co/datasets/KoelLabs/speech-accent-context.audio1K<n<10K0 likes9 downloads2y agoHugging Face22mteb /ambient-acoustic-context-smallaudio1K<n<10K0 likes9 downloads1y agoHugging Face23StevenDillmann /ContextConsistencyChecks Audio Consistency Checks Dataset This dataset contains audio-based ambiguity-resolution tasks across prosody categories. Columns: row_id: unique row id group_id: _ category: category name audio: relative path to the single combined input audio per row. Order: target separator + target, then spoken separators and three completion audios in randomized A/B/C order. correct_completion: one of "Completion A", "Completion B", "Completion C" input: text of the input (target) sentence… See the full description on the dataset page: https://huggingface.co/datasets/StevenDillmann/ContextConsistencyChecks.audion<1K0 likes8 downloads1y agoHugging Face24wangyueyiiiiiii /ContextDialog Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models 🎉 We are excited to announce that our paper has been accepted to the Findings of ACL 2025! Demo Page: https://contextdialog.github.io arXiv: https://arxiv.org/abs/2502.19759 ContextDialog is a comprehensive benchmark designed to evaluate a voice interaction model’s ability to engage in, retain, and leverage relevant information throughout multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/wangyueyiiiiiii/ContextDialog.audioaudio-to-audio1K<n<10K0 likes8 downloads5mo agoHugging Face25chiyuanhsiao /in_context_ASR_TTS_finetune_3-2-11B_rank64_ls960_replay_v5audion<1K0 likes7 downloads2y agoHugging Face26Yudong03 /Contextual_Audio_cuesaudion<1K0 likes6 downloads11mo agoHugging Face27AdnanElAssadi /ambient-acoustic-context-smallaudio1K<n<10K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.