datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_and_noise_interleaved
common_voice_and_noise_interleaved
This dataset was built for:
measuring timestamps drift with forced aligners relying on audio emissions and a backtracking algorithm (method used in whisperX for instance). Indeed, periods of noise may alter the backtracking algorithm
testing hallucination resistance: long periods without speech increase the risk.
Curation Method
This dataset is synthetically built by interleaving real speech and background noise segments. Speech… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/common_voice_and_noise_interleaved.interleaved-dataset
Interleaved-Dataset
This dataset contains word-level aligned audio segments corpus. Generated using WhisperX alignment.
