datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wjbmattingly_xhosa_merged_audio
Xhosa Merged Audio
This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below).
The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/ilyes25/wjbmattingly_xhosa_merged_audio.muscat-merged-samples
MUSCAT — Merged Long-Form Samples
This dataset is a merged, long-form reformatting of
goodpiku/muscat-eval
(MUSCAT: A Multi-Device Dataset for Code-Switching ASR and Segmentation
Evaluation).
The original MUSCAT release stores each conversation as many short,
single-language segments. Here those segments are concatenated back into one
continuous recording per conversation, so each row is a single long-form
code-switching audio with inline language/timing markers. The layout… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/muscat-merged-samples.xhosa_merged_audio
Xhosa Merged Audio
This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below).
The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/wjbmattingly/xhosa_merged_audio.tigrinya-asr-merged
tigrinya-asr-merged
A merged Tigrinya speech-recognition dataset, combining and deduplicating:
badrex/tigrinya-speech (train pool)
google/WaxalNLP config tir_asr (train pool)
UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set)
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.xhosa_merged_audio
Xhosa Merged Audio
This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below).
The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/xhosa_merged_audio.amharic-asr-merged
amharic-asr-merged
A merged Amharic speech-recognition dataset, combining and deduplicating:
badrex/amharic-speech
chappM/amharic-bdu-asr
beimnet777/amharic-asr
snapwre/amharic-speech
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed
Re-split into train (90%) / validation (5%) / test (5%), ignoring original source… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/amharic-asr-merged.english-en-x-code-switching-main-lang-samples-merged
English EN-X Code-Switching Main-Language Merged Samples
This dataset contains contiguous same-language segments from the paired mixed dataset.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples-merged.asr-merged-fr-20260814-144411
Test_Merged_datasets_3
Dataset ASR fusionné : audio + texte matérialisés en Parquet sur le Hub
(colonnes audio, text, source_dataset).
Nombre d'exemples poussés : 346.
Sources
BrunoHays/Accueil_UBS (config=default, split=test, audio=audio, text=sentence, max_samples=500)
eustlb/french-long-form-test (config=default, split=test, audio=audio, text=sentence, max_samples=500)
Normalisation texte
Mode : training_v3 (training_v3 | legacy | none)… See the full description on the dataset page: https://huggingface.co/datasets/Zeldeo/asr-merged-fr-20260814-144411.waxal-orm-tts-merged
Waxal Oromo TTS Merged
This dataset merges the human-labeled Oromo ASR split from
google/WaxalNLP with the autolabeled Oromo split from
israel/waxal-autolabled.
For TTS use, the leading [ORM] language tag has been removed from
autolabeled transcriptions. Rows include both text and transcription
with the same cleaned value.
Target repo: b1n1yam/waxal-orm-tts-merged
