datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenBible_Swahili_book_splitAfrivoice_Swahili_ASRAfrivoice_Swahili_ASROpenBible_Swahili_cleanswahili-speech-400hrswahiliSwahilidata_77swahiliSwahilidata_11swahiliSwahilidata_88swahiliSwahilidata_55swahili-speech-400hrswahiliSwahilidata_2swahiliSwahilidata_22swahili_MED_otherSwahilidata_44swahiliSwahilidata_33swahili_mediumSwahilidata_77swahili_mediumSwahilidata_11swahiliSwahilidata_4swahiliSwahilidata_66swahiliSwahilidata_44swahiliSwahilidata_3swahili_mediumSwahilidata_33swahili_mediumSwahilidata_66swahili_mediumSwahilidata_88swahili_mediumSwahilidata_22swahili_MED_otherSwahilidata_11mozilla_commonvoice_Swahili_preprocessed_train_batch_1swahili-speech
Swahili Speech-to-Text Dataset
This dataset contains paired audio and text data for training and evaluating speech-to-text models in Swahili. The audio files have been processed to remove silence, converted to 44.1kHz mono FLAC format, and are paired with corresponding transcriptions.
Structure
audio_*.flac: Audio files in FLAC format, named by their corresponding text corpus ID.
metadata.jsonl: JSON Lines file with metadata for each audio-text pair. Each line is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/stem-content-ai-project/swahili-speech.mozilla_commonvoice_Swahili_preprocessed_train_batch_3swahiliSwahilidata_1swahili_MED_otherSwahilidata_22
