datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ESpeech-webinars2
Webinar Audio Dataset
Dataset Description
This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Task: TTS, ASR, Quality Asessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON metadata
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.ESpeech-buldjat
Buldjat YouTube Audio Dataset
Dataset Description
This dataset contains 54 hours of processed audio segments extracted from the "Buldjat" YouTube channel with corresponding metadata. Each audio file represents a segment from the channel's videos and content, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Task: TTS, ASR, Quality Assessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON metadata
Source:… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-buldjat.
