datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio-Transcription-Models-Comparison-PT-BR
Audio Transcription Models Comparison
A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese.
About the Dataset
This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering:
Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.mig-burmese-audio-transcription
👨💻 Burmese Audio Transcription Dataset
Myanmar (Burmese) audio transcription အတွက် ပြုစုထားသော dataset ဖြစ်ပါတယ်။
Speech to Text, Text to Speech (TTS) နဲ့ ASR လုပ်ငန်းစဉ်များအတွက် တစ်ထောင့်တစ်နေရာက အထောက်အကူပြုနိုင်လိမ့်မယ်လို့ မျှော်လင့်မိပါတယ်။
Samples ပေါင်း 2822 ဝန်းကျင်ခန့် ရှိတာကြောင့် project အသေးလေးတွေအတွက် စမ်းကြည့်နေလို့ ရပါပြီ။
နောက်ပိုင်းမှာလည်း တတ်နိုင်သလောက် ဖြည့်စွတ်ပေးသွားပါမယ်။
Audio ဖိုင်တွေကိုတော့ Ramblings by Hein, Knowledge Worm နဲ့ youtube audio book များမှ… See the full description on the dataset page: https://huggingface.co/datasets/Ko-Yin-Maung/mig-burmese-audio-transcription.CHiME6_formatted_transcriptionsmig-burmese-audio-transcription
👨💻 Burmese Audio Transcription Dataset
Myanmar (Burmese) audio transcription အတွက် ပြုစုထားသော dataset ဖြစ်ပါတယ်။
Speech to Text, Text to Speech (TTS) နဲ့ ASR လုပ်ငန်းစဉ်များအတွက် တစ်ထောင့်တစ်နေရာက အထောက်အကူပြုနိုင်လိမ့်မယ်လို့ မျှော်လင့်မိပါတယ်။
Samples ပေါင်း 2822 ဝန်းကျင်ခန့် ရှိတာကြောင့် project အသေးလေးတွေအတွက် စမ်းကြည့်နေလို့ ရပါပြီ။
နောက်ပိုင်းမှာလည်း တတ်နိုင်သလောက် ဖြည့်စွတ်ပေးသွားပါမယ်။
Audio ဖိုင်တွေကိုတော့ Ramblings by Hein, Knowledge Worm နဲ့ youtube audio book… See the full description on the dataset page: https://huggingface.co/datasets/hackerlim7/mig-burmese-audio-transcription.mig-burmese-audio-transcription
👨💻 Burmese Audio Transcription Dataset
Myanmar (Burmese) audio transcription အတွက် ပြုစုထားသော dataset ဖြစ်ပါတယ်။
Speech to Text, Text to Speech (TTS) နဲ့ ASR လုပ်ငန်းစဉ်များအတွက် တစ်ထောင့်တစ်နေရာက အထောက်အကူပြုနိုင်လိမ့်မယ်လို့ မျှော်လင့်မိပါတယ်။
Samples ပေါင်း 2822 ဝန်းကျင်ခန့် ရှိတာကြောင့် project အသေးလေးတွေအတွက် စမ်းကြည့်နေလို့ ရပါပြီ။
နောက်ပိုင်းမှာလည်း တတ်နိုင်သလောက် ဖြည့်စွတ်ပေးသွားပါမယ်။
Audio ဖိုင်တွေကိုတော့ Ramblings by Hein, Knowledge Worm နဲ့ youtube audio book… See the full description on the dataset page: https://huggingface.co/datasets/Nawsantki/mig-burmese-audio-transcription.masri_audio_transcriptionArgentinian-audio-transcriptions
Argentinian audio transcriptions dataset
The dataset choosen contains audio recordings of Argentinian speaker along with their corresponding transcription. It was obtain from this link.
It consists of 3,921 recordings from female speakers and 1,818 recordings from male speakers.
These recordings were downsampled to a 16 kHz sample rate.
The spectrogram folder contains the spectrograms of the recordings, as well as spectrograms of augmented versions of those recordings.
For more… See the full description on the dataset page: https://huggingface.co/datasets/LautaroOcho/Argentinian-audio-transcriptions.longform-audio-transcriptionLIGHT_transcriptionsaudio-transcription-sample4-yas
Dataset Card for "audio-transcription-sample4-yas"
More Information needed
audio-transcription-sample3-yas
Dataset Card for "audio-transcription-sample3-yas"
More Information needed
New-transcription-with-audioArabic_Audio_Transcription_WhisperAudioTranscriptionDatasetaudio-transcription-dataset
Dataset Card for "audio-transcription-dataset"
More Information needed
arabic_audio_transcriptionsynthetic_transcription_audioaudio-transcription-sample1
Dataset Card for "audio-transcription-sample1"
More Information needed
sample-audio-transcriptionaudio-transcription-sample1
Dataset Card for "audio-transcription-sample1"
More Information needed
audio-transcription-sample2
Dataset Card for "audio-transcription-sample2"
More Information needed
audio-transcription-aya_2
Dataset Card for "audio-transcription-aya_2"
More Information needed
audio-transcription-test
Dataset Card for "audio-transcription-test"
More Information needed
