datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Akan_non_standardspeechmoore-audio-standardized ---
pretty_name: louisbertson/moore-audio-standardized
language:
- mos
tags:
- audio
- moore
- self-supervised-learning
- speech
size_categories:
- n<1K
---
# Mooré Standardized Audio Dataset
This dataset was exported from the preprocessing pipeline in this repository. It keeps the repository's canonical split manifests and uses standardized WAV audio so the same files work in local training, Google Colab, and Hugging Face Hub uploads.… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/moore-audio-standardized.standard_dataset_nonsynthetic_sorted_validation_set
Dataset Card for "standard_dataset_nonsynthetic_sorted"
More Information needed
ghanian_ga_standard_speech_v1.0This dataset provides 8.8 hours of Ga standard speech recordings (16,028 samples) from 81 Ga speakers with standard speech.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or phrase level. Two passes of transcription have been made to ensure correctness. For this dataset, all recordings with low confidence transcriptions or disagreement between transcribers have been removed.
Additionally, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/cdli/ghanian_ga_standard_speech_v1.0.
