CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Peacockery /georgian-asr-corpus-v0 georgian-asr-corpus-v0 145.3 hours of Georgian ASR training data: 92,185 clips across FLEURS ka_ge and Common Voice Georgian (scripted 25.0 and spontaneous 3.0, via the Mozilla Data Collective). Splits: train 64,633 / dev 13,456 / test 14,096. Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=kat_Geor/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and audio_size (sample count).… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/georgian-asr-corpus-v0.textautomatic-speech-recognition10K<n<100K0 likes23 downloads4mo agoHugging Face02Speech-data /Georgian-Speech-Dataset Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Georgian (ka) 🏷️ Tags Audio, Speech, Speech Recognition, Georgian, ML, Machine, Machine Learning 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes20 downloads6mo agoHugging Face03akalandia /ka-geo-voice-male-v1 Dataset Card for Georgian Male Voice Dataset v1 Intended Use Primary Use: Training and fine-tuning TTS models for Georgian language synthesis, including microsoft/speecht5_tts. Secondary Use: Research in speech synthesis, voice conversion, or linguistic analysis. SpeechT5 Compatibility This dataset is specifically formatted to be compatible with microsoft/speecht5_tts fine-tuning. The dataset includes: Audio: 22,050 Hz mono WAV files (matching SpeechT5… See the full description on the dataset page: https://huggingface.co/datasets/akalandia/ka-geo-voice-male-v1.audiotext-to-speechn<1K0 likes9 downloads9mo agoHugging Face04GeoPoll /dataset-20250728_102101-swgated GeoPoll Swahili Speech Dataset This dataset contains speech recognition data for Swahili (sw) collected and processed by GeoPoll. Dataset Summary This dataset is designed for fine-tuning speech recognition models on Swahili audio data. It includes high-quality audio segments with corresponding transcriptions. Dataset Statistics Total samples: 11814 Total duration: 20.45 hours Average duration: 6.23 seconds per sample Number of speakers: 6 Language: Swahili… See the full description on the dataset page: https://huggingface.co/datasets/GeoPoll/dataset-20250728_102101-sw.audioautomatic-speech-recognition10K<n<100K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.