CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yuriilaba /toronto-tv-ukrainian Toronto TV Ukrainian Speech Dataset For educational purposes only. All rights to the original video and audio content belong to Телебачення Торонто (YouTube channel). Dataset Summary A Ukrainian-language speech dataset parsed from the Телебачення Торонто YouTube channel. Each sample consists of a short audio clip and its corresponding Ukrainian subtitle text, intended for use in automatic speech recognition (ASR) research and education. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yuriilaba/toronto-tv-ukrainian.audioautomatic-speech-recognition10K<n<100K2 likes145 downloads4mo agoHugging Face02robinhad /VOA-ukrgatedaudio10K<n<100K0 likes82 downloads2y agoHugging Face03KSE-RESEARCH-Group /ukr-dialects-audio-dataset Ukrainian Dialects Audio Dataset Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits. Dataset Description This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets: NaUKMA-Audio-Dataset Ivanna-Stefiuk-Audio-Dataset Larysa-Irodenko-Audio-Dataset Hutsulendia-Audio-Dataset Dido-Yvanchyk-Audio-Dataset-v2 Dataset Structure train: 27,675 samples validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.audioautomatic-speech-recognition10K<n<100K1 likes80 downloads7mo agoHugging Face04vladsfa /ukr-dialects-audio-dataset Ukrainian Dialects Audio Dataset Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits. Dataset Description This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets: NaUKMA-Audio-Dataset Ivanna-Stefiuk-Audio-Dataset Larysa-Irodenko-Audio-Dataset Hutsulendia-Audio-Dataset Dido-Yvanchyk-Audio-Dataset-v2 Dataset Structure train: 27,675 samples… See the full description on the dataset page: https://huggingface.co/datasets/vladsfa/ukr-dialects-audio-dataset.audioautomatic-speech-recognition10K<n<100K0 likes80 downloads16d agoHugging Face05Mikhailo /ukrainian-tts-audiobooks-24khz Ukrainian Audiobook TTS Dataset (24 kHz) Description Ukrainian speech dataset for TTS and ASR tasks. Source Dataset https://huggingface.co/datasets/Yehor/audiobooks-xxl Processing Pipeline MusicDetection filtering — removed samples with background music/noise Audio processing (Sidon) — resampled 16 kHz → 24 kHz, converted to mono Transcription — generated with nvidia/canary-1b-v2 Dataset Structure Column Type… See the full description on the dataset page: https://huggingface.co/datasets/Mikhailo/ukrainian-tts-audiobooks-24khz.audioautomatic-speech-recognition1M<n<10M3 likes72 downloads3mo agoHugging Face06vbrydik /ukr-male-speaker-0-v0audio1K<n<10K1 likes30 downloads3y agoHugging Face07shunyalabs /ukrainian-speech-datasetaudio1K<n<10K0 likes28 downloads1y agoHugging Face08rishchen /ukrainian-tts-audiobook-pani-nina-parquet Ukrainian TTS audiobook dataset Pani Nina (Parquet) Segmented Ukrainian audiobook speech with aligned text, prepared for training and evaluating Text-to-Speech (TTS) models. The dataset is published as Hugging Face-compatible Parquet shards so the Hub Dataset Preview can render an audio column. The dataset was prepared using whisper and ffmpeg: Whisper was used for transcription and approximate segment timing. FFmpeg was used to slice audio into short utterances (roughly 2-10… See the full description on the dataset page: https://huggingface.co/datasets/rishchen/ukrainian-tts-audiobook-pani-nina-parquet.audiotext-to-speech100K<n<1M0 likes15 downloads5mo agoHugging Face09Speech-data /Ukrainian-Speech-Dataset Ukrainian Dataset Metadata Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Ukrainian (uk) 🏷️ Tags Audio, Speech, Speech Recognition, ML, Machine, Machine Learning 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes15 downloads6mo agoHugging Face10vbrydik /ukr-male-speaker-0-v0-vitsaudio1K<n<10K0 likes14 downloads3y agoHugging Face11vbrydik /ukr-voice-vbaudion<1K0 likes11 downloads2y agoHugging Face12vbrydik /ukr-cv-speakers-v0audio10K<n<100K1 likes10 downloads2y agoHugging Face13Saidakbar01 /ukraina_urushaudion<1K0 likes9 downloads4mo agoHugging Face14rishchen /ukrainian-tts-audiobook-pani-nina-parquet-old Ukrainian TTS audiobook dataset (Parquet) Segmented Ukrainian audiobook speech with aligned text, prepared for training and evaluating Text-to-Speech (TTS) models. The dataset is published as Hugging Face-compatible Parquet shards so the Hub Dataset Preview can render an audio column. That was mabe by using whisper (https://github.com/openai/whisper) and ffmpeg (https://www.ffmpeg.org/), where with whisper we set start and end of voices + transcribe it and using ffmpeg slice into… See the full description on the dataset page: https://huggingface.co/datasets/rishchen/ukrainian-tts-audiobook-pani-nina-parquet-old.audiotext-to-speech100K<n<1M0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.