datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
asr-semantic-probe-eng
ASR Semantic Probing Dataset (English)
Synthetic English audio dataset for probing whether ASR encoder representations
encode semantic category information beyond acoustic features. Constructed for
mechanistic interpretability studies of speech recognition models.
Splits
This dataset is released as a single unsplit collection. Downstream users are
expected to define their own train/test splits based on the experimental design.
For probing experiments where speaker… See the full description on the dataset page: https://huggingface.co/datasets/soaring0616/asr-semantic-probe-eng.mondegreenbench
MondegreensEval
Audio companion dataset for MondegreensEval: A Phonetic Benchmark for Measuring
Language-Model Bias in Automatic Speech Recognition (ICML ML for Audio Workshop, 2026).
Code, evaluation pipeline, and per-model transcription/metric outputs:
https://github.com/soarhigh/mondegreenbench
Mondegreens — phonetically near-identical phrase pairs with distinct meanings — expose a
measurable failure mode in decoder-based ASR: the model's internal language-model prior
can… See the full description on the dataset page: https://huggingface.co/datasets/soarhigh/mondegreenbench.
