datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Indic-total-New-TTS-Merge
Indic Total TTS Merge
Merged TTS dataset with 13 Indic languages. All audio clips are >= 3.0 seconds duration.
Languages
assamese, bengali, english, gujarati, hindi, kannada, malayalam, marathi, nepali, odia, punjabi, tamil, telugu
Columns
audio: Audio data
text: Transcript text
duration: Duration in seconds (all >= 3.0s)
language: Language name
merged_speech_datasetsyspin_hindi_mergedanv_data_ke_kikuyu_mergedhindi_karya_mergedmerged-data-250924Test test Hello 1234
indo-merged-dataset-2
Dataset Card for "indo-merged-dataset-2"
More Information needed
merged-bambara-dioula-datasetmerged-hindi-audio-datasetpersian_tts_merged
Merged Persian TTS Dataset
Overview
This dataset is a comprehensive collection of Persian speech data, merged from several high-quality sources to facilitate Text-to-Speech (TTS) research and development for the Persian language. It combines audio recordings with their corresponding transcriptions, providing a rich resource for training and evaluating Persian TTS systems.
Dataset Details
Language: Persian (Farsi)
Total Samples: [Insert total number of samples… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian_tts_merged.darija-merged-asroct21_merged_finalShrutilipi_Hindi_resampled_44100_merged_10wjbmattingly_xhosa_merged_audio
Xhosa Merged Audio
This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below).
The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/ilyes25/wjbmattingly_xhosa_merged_audio.muscat-merged-samples
MUSCAT — Merged Long-Form Samples
This dataset is a merged, long-form reformatting of
goodpiku/muscat-eval
(MUSCAT: A Multi-Device Dataset for Code-Switching ASR and Segmentation
Evaluation).
The original MUSCAT release stores each conversation as many short,
single-language segments. Here those segments are concatenated back into one
continuous recording per conversation, so each row is a single long-form
code-switching audio with inline language/timing markers. The layout… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/muscat-merged-samples.dysarthria-asr-mergedTurkish-Podcast-Merge-v1merged-bambara-dioula-datasetS2T_Korean_Merge_2_fixed4Shrutilipi_Hindi_resampled_44100_merged_1merged_english_accent_dataset241022_merged_data_fixed_distributionxhosa_merged_audio
Xhosa Merged Audio
This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below).
The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/wjbmattingly/xhosa_merged_audio.syspin_merged_with_descMerged_Luo_DatasetTTS_merge-ties_ls960-testIndicVoices_Hindi_audio_44100_mergedTTS_merge-dare_ls960-testViMD_Dataset_Merged2026tigrinya-asr-merged
tigrinya-asr-merged
A merged Tigrinya speech-recognition dataset, combining and deduplicating:
badrex/tigrinya-speech (train pool)
google/WaxalNLP config tir_asr (train pool)
UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set)
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.
