zbrunner/speakeroverlap_multiseg
MultiSeg Dataset for ASR Hallucinations Description MultiSeg is a perturbed and altered version of the TEDLIUM3 dataset, specifically created for evaluating the robustness of Automatic Speech Recognition (ASR) systems. This dataset is derived from the 'speakeroverlap' subset, which consists of held-back training data from TEDLIUM3. Purpose The primary purpose of the MultiSeg dataset is to: Elicit hallucinations from ASR systems Evaluate ASR… See the full description on the dataset page: https://huggingface.co/datasets/zbrunner/speakeroverlap_multiseg.
07
Update README.md
Update README.md
Update README.md
Upload folder using huggingface_hub
initial commit
