datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mozilla_commonvoice_hackathon_preprocessed_train_batch_3
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_2
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_5
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_1
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_4
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_6
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6"
More Information needed
tat_hackathon_asr
Hackathon Tatar ASR
Dataset Summary
Hackathon Tatar ASR is a speech dataset distributed during the "Татар.Бу Хакатон" (Tatar.Bu Hackathon) held in Tatarstan in May 2024. This dataset likely consists of newly collected crowdsourced recordings created after the last release of TatSC (Tatar Speech Corpus), although some intersections with TatSC might be present. While TatSC contains 269.1 hours of transcribed speech with 271,914 utterances, this hackathon dataset comprises… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tat_hackathon_asr.kinyarwanda-speech-hackathon
📚 Kinyarwanda ASR Dataset
This dataset contains transcribed Kinyarwanda audio, designed to support training and evaluation of Automatic Speech Recognition (ASR) systems. It is part of a study on how varying training data volumes affect model performance using Whisper-large-v3.
📂 Data Overview
The full dataset consists of approximately 263,000 audio samples covering 5 key domains:
🏥 Health
🏛️ Government
💰 Financial Services
🎓 Education
🌾 Agriculture
To… See the full description on the dataset page: https://huggingface.co/datasets/evie-8/kinyarwanda-speech-hackathon.kinyarwanda-hackathonkinyarwanda-speech-hackathon
