datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
toronto-tv-ukrainian
Toronto TV Ukrainian Speech Dataset
For educational purposes only.
All rights to the original video and audio content belong to Телебачення Торонто (YouTube channel).
Dataset Summary
A Ukrainian-language speech dataset parsed from the Телебачення Торонто YouTube channel. Each sample consists of a short audio clip and its corresponding Ukrainian subtitle text, intended for use in automatic speech recognition (ASR) research and education.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yuriilaba/toronto-tv-ukrainian.ukrainian-tts-audiobooks-24khz
Ukrainian Audiobook TTS Dataset (24 kHz)
Description
Ukrainian speech dataset for TTS and ASR tasks.
Source Dataset
https://huggingface.co/datasets/Yehor/audiobooks-xxl
Processing Pipeline
MusicDetection filtering — removed samples with background music/noise
Audio processing (Sidon) — resampled 16 kHz → 24 kHz, converted to mono
Transcription — generated with nvidia/canary-1b-v2
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Mikhailo/ukrainian-tts-audiobooks-24khz.Ukrainian-Speech-Dataset
Ukrainian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Ukrainian (uk)
🏷️ Tags
Audio, Speech, Speech Recognition, ML, Machine, Machine Learning
📦 Size Category
n < 1K
