datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/google/WaxalNLP.laions_got_talent
LAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel.
Dataset Composition
The dataset includes:
Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.laions_got_talent_raw0-9up_google_speech_commands_augmented_raw
Dataset Card for "google_speech_commands_augmented_raw_fixed"
More Information needed
laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.arc-voicesamples-generatedarc-speeches-refinedgooshkon-chunks
Gooshkon MP3 chunks
A single-column Hugging Face Audio dataset. Every row contains embedded,
playable MP3 bytes in the audio column. Chunks target approximately 30
seconds and are selected at detected quiet intervals with a 12-48 second
safety range.
The source recordings are not transcript-aligned; this release uses acoustic
silence boundaries to avoid cutting through words whenever a usable pause is
available.
Dataset runtime
The dataset contains approximately 2… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/gooshkon-chunks.laions_got_talent_enhanced_no_metadatalaions_got_talent_german_bicodecgoogletime
googletime
Audio validation set organized as validation/audio/* plus validation/metadata.jsonl. The metadata contains only file_name and transcription; transcriptions include timestamp and speaker markers.
sberdevices_golos_10h_crowd
Dataset Card for sberdevices_golos_10h_crowd
Dataset Summary
Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated.
Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.Taiwan-Tongues-ASR-CE-dataset-hokkien
Taiwan-Tongues-ASR-CE-dataset-hokkien
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.golos_mfa_punctuation
Golos MFA Punctuation
Расширенная версия датасета Golos —
русскоязычного корпуса речи с краудсорс и студийными записями.
Датасет дополнен пунктуацией и word-level временными метками (MFA alignment).
Опубликовано и поддерживается Jeti Labs.
Описание
Параметр
Значение
Язык
Русский (ru)
Записей
970,597
Аудио
~1,044 часов
Частота дискретизации
16,000 Hz
Формат
WAV, mono, 16-bit
Что добавлено по сравнению с оригинальным Golos… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/golos_mfa_punctuation.laions_got_talent_previewsberdevices_golos_100h_farfield
Dataset Card for sberdevices_golos_100h_farfield
Dataset Summary
Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated.
Authors divide all dataset into train and test… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield.QUT-Event-VTR-Dataset
Event-Based Visual Teach-and-Repeat via Fast Fourier-Domain Cross-Correlation
Welcome to the official QUT-Event-VTR-Dataset dataset repository attached to the paper Event-Based Visual Teach-and-Repeat via Fast Fourier-Domain Cross-Correlation, to be presented at the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026).
File Structure
Cite us at
Event-Based Visual Teach-and-Repeat via Fast Fourier-Domain… See the full description on the dataset page: https://huggingface.co/datasets/gokulbnr/QUT-Event-VTR-Dataset.google-chilean-spanish
Dataset Card for Tamil Speech
Dataset Summary
This dataset consists of 7 hours of transcribed high-quality audio of Chilean Spanish sentences recorded by 31 volunteers. The dataset is intended for speech technologies.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
Supported Tasks
text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS).
automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-chilean-spanish.govvox-100h-v3
GovVox-100h-v3 — 9 tỉnh đặc trưng × 3 miền đều 33.3h
thời lượng
95.37 giờ (chia đều ~33.3h/miền)
số đoạn
38,744
người nói
1,056 (2,343 id diarization theo phiên)
tỉnh
9 (mỗi miền đúng 3 tỉnh đặc trưng nhất)
bản ghi nguồn
267
9 tỉnh được chọn — theo kết quả thí nghiệm thực tế
Miền
Tỉnh
Nguồn khả dụng
Accuracy mô hình (3 lớp / 5 lớp)
Bắc
Hà Nội
21.97h
92–100%
Bắc
Hải Phòng
30.19h
84–86%
Bắc
Ninh Bình
58.25h
89–92%
Trung
Hà… See the full description on the dataset page: https://huggingface.co/datasets/vnpost-ai/govvox-100h-v3.japanese-casual-conversational-speech-golden-dataset-preview
Japanese Casual Conversational Speech Golden Dataset (Preview)
💼 Commercial License & Full Access
This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning.
To purchase the full dataset, please contact us:
👉 Email: info@hth-inc.com
👉 Website: https://hth-inc.com/business
🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.leyu-amharic-gonder-dialect
Leyu Amharic - Gonder Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.leyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.google-colombian-spanish
Dataset Card for "google-colombian-spanish"
More Information needed
google-la-voices
Dataset Card for "google-la-voices"
Speaker Durations
Speaker
Duration (seconds)
00295
1606.144
00610
7026.261
01208
3284.907
01523
6309.888
02121
4687.445
02436
4654.080
02484
9379.925
02485
130.219
03034
5186.048
03349
5143.381
03397
7852.203
03398
118.101
03853
638.037
04310
8260.437
04311
105.472
04766
590.165
05223
8257.773
05679
846.251
0613610207.707
06592
863.659
07049
7580.715
07060
575.659
07505
1743.531… See the full description on the dataset page: https://huggingface.co/datasets/ittailup/google-la-voices.simplified_google_speech_commands_wav2vec2_960hgerman-golden-audio_speech-IPA
🌟 German Golden Speech & IPA Corpus (FLEURS + Multilingual TEDx)
An ultra-clean, high-standard curated German speech dataset combining Google FLEURS (de_de) and Multilingual TEDx German (mTEDx), fully embedded with 16kHz WAV audio bytes, normalized orthographic text, and pre-computed International Phonetic Alphabet (IPA) transcriptions.
📊 Dataset Summary
Total Samples: 1,354 high-quality audio recordings.
Total Size: ~419 MB (Compressed Parquet format).
Audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-golden-audio_speech-IPA.google-tamil
Dataset Card for Tamil Speech
Dataset Summary
This dataset consists of 7 hours of transcribed high-quality audio of Tamil sentences recorded by 50 volunteers. The dataset is intended for speech technologies.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
Supported Tasks
text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS).
automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-tamil.Taiwan-Tongues-ASR-CE-dataset-zhtw
Taiwan-Tongues-ASR-CE-dataset-zhtw
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.visualears-golden-6669
🗂️ visualears-golden-6669
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Held-out VisualEars6669 / Golden6669 evaluation dataset.
مجموعهٔ ارزیابی نگهداشتهشدهٔ Golden6669 با شرایط پاک، دورمیدان و مسدود برای سنجش واقعگرایانهٔ ASR فارسی.
🧩 Role
evaluation and benchmarking asset
مصنوع ارزیابی و بنچمارک
📦 Snapshot
13 files; approximately 916.66 MB
13 فایل؛ حدود… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-golden-6669.
