datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-contextasr-bench
Thai Contextual-Biasing ASR Benchmark
TL;DR
Does your Thai ASR system actually use the context you give it (e.g., a list of names, custom words from your own dictionary) — and does it hallucinate when the context is irrelevant?
Each utterance comes with a bias list: entity strings (brands, person names,
places) that may or may not be spoken in the audio, written the way a real Thai user
would write them — one list, mixed Thai and Latin script. A good system does… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-contextasr-bench.asr-context-induced-leakage
When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR
Overview
SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts, fine-tune on proprietary recordings, or both. We identify and systematically investigate an overlooked privacy risk of such customisation: a model adapted to recognise domain-specific terminology can be nudged into transcribing a phonetically similar… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/asr-context-induced-leakage.
