datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
toronto-tv-ukrainian
Toronto TV Ukrainian Speech Dataset
For educational purposes only.
All rights to the original video and audio content belong to Телебачення Торонто (YouTube channel).
Dataset Summary
A Ukrainian-language speech dataset parsed from the Телебачення Торонто YouTube channel. Each sample consists of a short audio clip and its corresponding Ukrainian subtitle text, intended for use in automatic speech recognition (ASR) research and education.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yuriilaba/toronto-tv-ukrainian.ukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples
validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.ukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples… See the full description on the dataset page: https://huggingface.co/datasets/vladsfa/ukr-dialects-audio-dataset.ukrainian-tts-audiobooks-24khz
Ukrainian Audiobook TTS Dataset (24 kHz)
Description
Ukrainian speech dataset for TTS and ASR tasks.
Source Dataset
https://huggingface.co/datasets/Yehor/audiobooks-xxl
Processing Pipeline
MusicDetection filtering — removed samples with background music/noise
Audio processing (Sidon) — resampled 16 kHz → 24 kHz, converted to mono
Transcription — generated with nvidia/canary-1b-v2
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Mikhailo/ukrainian-tts-audiobooks-24khz.Ukrainian-Speech-Dataset
Ukrainian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Ukrainian (uk)
🏷️ Tags
Audio, Speech, Speech Recognition, ML, Machine, Machine Learning
📦 Size Category
n < 1K
17-minute-world-languages_ukrainien
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/ukrainienne/
Site à scrapper
