datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tigrinya-asr-merged
tigrinya-asr-merged
A merged Tigrinya speech-recognition dataset, combining and deduplicating:
badrex/tigrinya-speech (train pool)
google/WaxalNLP config tir_asr (train pool)
UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set)
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.juba-arabic-audio-translation
Juba Arabic Audio to English Translation Dataset
Language Pair
Source Audio
Target Text
Total Samples
Total Duration
Juba Arabic (pga) $\rightarrow$ English (en)
Juba Arabic Spoken Audio (MP3)
English Story Translation
40
~55 minutes
📌 Dataset Summary
This dataset pairs Juba Arabic (عربي جوبا / Sudanese Creole Arabic), the primary lingua franca spoken across South Sudan, with aligned English translations.
The dataset consists of 40 narrated… See the full description on the dataset page: https://huggingface.co/datasets/harikc456/juba-arabic-audio-translation.seeed-local-voice-perf-corpus
Seeed Local Voice — Perf Test Corpus
Fixed 20-file audio corpus used to benchmark
Seeed-Projects/seeed-local-voice
across Jetson, Rockchip, and Raspberry Pi deployments.
The same .wav bytes are pulled by every device, so RTF / latency deltas
between devices are pure compute — not input variation.
Contents
5× zh short (1.5 – 4.0 s)
5× zh long (10 – 16 s)
5× en short (1.5 – 4.0 s)
5× en long (10 – 12 s)
Audio spec: 16 kHz mono 16-bit WAV.
Each file's SHA256 + transcript… See the full description on the dataset page: https://huggingface.co/datasets/harvestsu/seeed-local-voice-perf-corpus.amharic-asr-merged
amharic-asr-merged
A merged Amharic speech-recognition dataset, combining and deduplicating:
badrex/amharic-speech
chappM/amharic-bdu-asr
beimnet777/amharic-asr
snapwre/amharic-speech
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed
Re-split into train (90%) / validation (5%) / test (5%), ignoring original source… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/amharic-asr-merged.DFToolBench-A-500
DFToolBench-A-500
A 500-query benchmark for evaluating audio tool-use agents on deepfake-related forensic tasks.
Each query is a multi-turn ReAct-style dialog in which an assistant invokes audio analysis tools
(e.g. speaker_verification, nisqa, silero_vad, deepfake_audio, language_id, muq,
calculator) over a single audio file and produces a final verdict.
Files
dataset.json — pretty-printed list of 500 records.
dataset.jsonl — one record per line (preferred for… See the full description on the dataset page: https://huggingface.co/datasets/hardiksharma6555/DFToolBench-A-500.eka-hard
EKA Hard — Medical ASR Benchmark
Entity-aware medical ASR benchmark — 50 hard rows from Indian-accented clinical speech.
Prepared by Trelis Research. Watch more on Youtube or inquire about our custom voice AI (ASR/TTS) services here.
Source
Derived from ekacare/eka-medical-asr-evaluation-dataset (3,619 EN rows, MIT license). Real clinical speech from 57 speakers across 4 Indian medical colleges, 16kHz mono.
Preparation
Filter: audio ≥ 2s, text ≥ 20 chars… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eka-hard.asrdata
Dataset Card for harshk04/h
This dataset contains bilingual Hindi and English audio-text pairs formatted for automatic speech recognition (ASR) fine-tuning.
Dataset Overview
Languages: Hindi and English
Samples: 400
Split: train
License: MIT
Columns
Column
Type
Description
audio
Audio
Path to the audio file
language
string
Primary language of speech
languagesKnown
string
List of other known languages
gender
string
Speaker gender
state… See the full description on the dataset page: https://huggingface.co/datasets/harshk04/asrdata.multimed-hard
MultiMed Hard — Medical ASR Benchmark
Entity-aware medical ASR benchmark — 50 hard rows from medical lectures and interviews.
Prepared by Trelis Research. Watch more on Youtube or inquire about our custom voice AI (ASR/TTS) services here.
Source
Derived from leduckhai/MultiMed EN test split (4,751 rows, MIT license). YouTube medical channels — lectures, interviews, podcasts, documentaries. Transcripts are human-reviewed.
Preparation
Filter: audio ≥ 2s, ≤ 29s… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/multimed-hard.omniscribe_corpus
OmniScribe Corpus
A multilingual speech transcription corpus designed for fine-tuning ASR models on Indian medical and general-domain speech. It covers Hindi, Marathi, and Indian English, with a focus on clinical and healthcare contexts.
Overview
Split
Rows (after oversampling)
Approx. Duration
train
~30750
~230 hrs
benchmark
~4,089
~25 hrs
Audio samples average 20–30 seconds each. All samples are at least 5 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Harshkmr/omniscribe_corpus.h
Dataset Card for harshk04/h
This dataset contains bilingual Hindi and English audio-text pairs for automatic speech recognition (ASR) fine-tuning.
Dataset Overview
Languages: Hindi and English
Samples: 400
Split: train
License: MIT
Sources: Hindi (agent-laxmi), English (downloaded)
Columns
Column
Type
Description
audio_files
Audio
Path to the audio file
transcripts
string
Text transcript
language
string
'hindi' or 'english'
source… See the full description on the dataset page: https://huggingface.co/datasets/harshk04/h.
