datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tigrinya-asr-merged
tigrinya-asr-merged
A merged Tigrinya speech-recognition dataset, combining and deduplicating:
badrex/tigrinya-speech (train pool)
google/WaxalNLP config tir_asr (train pool)
UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set)
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.juba-arabic-audio-translation
Juba Arabic Audio to English Translation Dataset
Language Pair
Source Audio
Target Text
Total Samples
Total Duration
Juba Arabic (pga) $\rightarrow$ English (en)
Juba Arabic Spoken Audio (MP3)
English Story Translation
40
~55 minutes
📌 Dataset Summary
This dataset pairs Juba Arabic (عربي جوبا / Sudanese Creole Arabic), the primary lingua franca spoken across South Sudan, with aligned English translations.
The dataset consists of 40 narrated… See the full description on the dataset page: https://huggingface.co/datasets/harikc456/juba-arabic-audio-translation.amharic-asr-merged
amharic-asr-merged
A merged Amharic speech-recognition dataset, combining and deduplicating:
badrex/amharic-speech
chappM/amharic-bdu-asr
beimnet777/amharic-asr
snapwre/amharic-speech
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed
Re-split into train (90%) / validation (5%) / test (5%), ignoring original source… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/amharic-asr-merged.eka-hard
EKA Hard — Medical ASR Benchmark
Entity-aware medical ASR benchmark — 50 hard rows from Indian-accented clinical speech.
Prepared by Trelis Research. Watch more on Youtube or inquire about our custom voice AI (ASR/TTS) services here.
Source
Derived from ekacare/eka-medical-asr-evaluation-dataset (3,619 EN rows, MIT license). Real clinical speech from 57 speakers across 4 Indian medical colleges, 16kHz mono.
Preparation
Filter: audio ≥ 2s, text ≥ 20 chars… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eka-hard.multimed-hard
MultiMed Hard — Medical ASR Benchmark
Entity-aware medical ASR benchmark — 50 hard rows from medical lectures and interviews.
Prepared by Trelis Research. Watch more on Youtube or inquire about our custom voice AI (ASR/TTS) services here.
Source
Derived from leduckhai/MultiMed EN test split (4,751 rows, MIT license). YouTube medical channels — lectures, interviews, podcasts, documentaries. Transcripts are human-reviewed.
Preparation
Filter: audio ≥ 2s, ≤ 29s… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/multimed-hard.omniscribe_corpus
OmniScribe Corpus
A multilingual speech transcription corpus designed for fine-tuning ASR models on Indian medical and general-domain speech. It covers Hindi, Marathi, and Indian English, with a focus on clinical and healthcare contexts.
Overview
Split
Rows (after oversampling)
Approx. Duration
train
~30750
~230 hrs
benchmark
~4,089
~25 hrs
Audio samples average 20–30 seconds each. All samples are at least 5 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Harshkmr/omniscribe_corpus.
