CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01winsweb /Akan_non_standardspeechaudio1K<n<10K0 likes44 downloads1y agoHugging Face02louisbertson /moore-audio-standardized --- pretty_name: louisbertson/moore-audio-standardized language: - mos tags: - audio - moore - self-supervised-learning - speech size_categories: - n<1K --- # Mooré Standardized Audio Dataset This dataset was exported from the preprocessing pipeline in this repository. It keeps the repository's canonical split manifests and uses standardized WAV audio so the same files work in local training, Google Colab, and Hugging Face Hub uploads.… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/moore-audio-standardized.audion<1K1 likes31 downloads7mo agoHugging Face03nlewins /standard_dataset_nonsynthetic_sorted_validation_set Dataset Card for "standard_dataset_nonsynthetic_sorted" More Information needed audio1K<n<10K0 likes16 downloads3y agoHugging Face04cdli /ghanian_ga_standard_speech_v1.0gatedThis dataset provides 8.8 hours of Ga standard speech recordings (16,028 samples) from 81 Ga speakers with standard speech. This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or phrase level. Two passes of transcription have been made to ensure correctness. For this dataset, all recordings with low confidence transcriptions or disagreement between transcribers have been removed. Additionally, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/cdli/ghanian_ga_standard_speech_v1.0.audio10K<n<100K1 likes8 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.