datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stream-data-newCodemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
Emotion_new_collected_datasetgdpval_preference_rubricsdaily-bio-newsnew-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.Indic-total-New-TTS-Merge
Indic Total TTS Merge
Merged TTS dataset with 13 Indic languages. All audio clips are >= 3.0 seconds duration.
Languages
assamese, bengali, english, gujarati, hindi, kannada, malayalam, marathi, nepali, odia, punjabi, tamil, telugu
Columns
audio: Audio data
text: Transcript text
duration: Duration in seconds (all >= 3.0s)
language: Language name
voice_ds_newglobal-news-radio-30s
Global News Radio Dataset
Multilingual news radio recordings from 51 languages across 42 countries.
Recordings
51
Total audio
1500 min (25.0 h)
Format
MP3 16kHz mono 64kbps
Parquet shards
11
Languages
51
Countries
42
Size
687 MB
Languages
Amharic, Arabic, Bashkir, Basque, Belarusian, Bengali, Brazilian Portuguese,Portugues Do Brasil,Português Brasil, Catalan, Croatian, Czech, Danish, Dutch, English, Estonian, Faroese, Finnish, Flemish… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-30s.new-twi-tts-aligned
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi TTS Dataset
A speech dataset of Twi (Akan) extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Text-To-Speech (TTS) models.
📂 Dataset Structure
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned.news-segmentationnews_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/islomov/news_youtube_uzbek_speech_dataset.new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/new-twi-tts-aligned-ipa.new_trainbible-new-testamentnew-rl-vitavoice_ds_new_200news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/azimislom/news_youtube_uzbek_speech_dataset.news_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/BoburAmirov/news_youtube_uzbek_speech_dataset.tibetan-speech-english-text-dataset-new-updatednews_youtube_uzbek_speech_dataset
News Youtube Uzbek Speech Dataset
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with differenent dialects. The data was collected from publicly available news videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Kunuz, Qalampir YouTube channels. The data was transcribed using Gemini 2.5 Pro and was intelligently… See the full description on the dataset page: https://huggingface.co/datasets/hostbot77/news_youtube_uzbek_speech_dataset.voice_medical_newglobal-news-radio-debug
Global News Radio Dataset (1 hour per station)
Every news radio station from the Radio Browser API, recorded for 1 hour each.
Attempted
3037
Successful
2553
Failed
484
Total audio
21 hours
Parquet shards
256
Size
0.6 GB
Format
MP3 16kHz mono 64kbps
Usage
from datasets import load_dataset
ds = load_dataset("NathanRoll/global-news-radio-debug", streaming=True)
for sample in ds["train"]:
print(sample["station_name"], sample["language"]… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-debug.global-news-radio-fanoutNews
AnonymousContinuousBench — News
A news-grounded QA benchmark built from Common Crawl News (CC-NEWS) articles
crawled in September 2025. QAs are generated by Gemini 2.5 from clusters
of related articles, then filtered for answerability and grounded with a
retrieval-based set of supporting articles drawn from the corpus.
What's inside
Config
Splits
Size
What it's for
qa (default)
val (1,189), test (1,415)
233 MB
Evaluate QA on news, post-event
corpus_large… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousContinuousBench/News.kannada_new_datamunch-1-latent-NEW-parquet
🎙️ Urdu TTS Latent Dataset — munch-1-latent-NEW-parquet
Pre-computed DACVAE latent representations for 51,021 Urdu utterances, ready for TTS model training. No audio decoding required at training time — load the dataset, reshape the binary blob, and train.
Source
Field
Value
Source audio
Humair332/Urdu-munch-1
Codec
Aratako/Semantic-DACVAE-Japanese-32dim
Codec sample rate
48,000 Hz
Encoder hop size
1,920 samples
Latent frame rate
25.0 Hz
Latent dim… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/munch-1-latent-NEW-parquet.new1000dtsnew-twi-tts-aligned_normalised
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
BZNSYP_new
