datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.WaxalNLP
Waxal Datasets
The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.
Dataset Description
The Waxal project provides datasets for both Automated Speech Recognition (ASR)
and Text-to-Speech (TTS) for African languages. The goal of this dataset's
creation and release is to facilitate research that improves the accuracy and
fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/google/WaxalNLP.0-9up_google_speech_commands_augmented_raw
Dataset Card for "google_speech_commands_augmented_raw_fixed"
More Information needed
googletime
googletime
Audio validation set organized as validation/audio/* plus validation/metadata.jsonl. The metadata contains only file_name and transcription; transcriptions include timestamp and speaker markers.
google-chilean-spanish
Dataset Card for Tamil Speech
Dataset Summary
This dataset consists of 7 hours of transcribed high-quality audio of Chilean Spanish sentences recorded by 31 volunteers. The dataset is intended for speech technologies.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
Supported Tasks
text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS).
automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-chilean-spanish.google-colombian-spanish
Dataset Card for "google-colombian-spanish"
More Information needed
google-la-voices
Dataset Card for "google-la-voices"
Speaker Durations
Speaker
Duration (seconds)
00295
1606.144
00610
7026.261
01208
3284.907
01523
6309.888
02121
4687.445
02436
4654.080
02484
9379.925
02485
130.219
03034
5186.048
03349
5143.381
03397
7852.203
03398
118.101
03853
638.037
04310
8260.437
04311
105.472
04766
590.165
05223
8257.773
05679
846.251
0613610207.707
06592
863.659
07049
7580.715
07060
575.659
07505
1743.531… See the full description on the dataset page: https://huggingface.co/datasets/ittailup/google-la-voices.simplified_google_speech_commands_wav2vec2_960hgoogle-tamil
Dataset Card for Tamil Speech
Dataset Summary
This dataset consists of 7 hours of transcribed high-quality audio of Tamil sentences recorded by 50 volunteers. The dataset is intended for speech technologies.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
Supported Tasks
text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS).
automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-tamil.google-argentinian-spanish
Dataset Card for "google-argentinian-spanish"
More Information needed
google-gujarati
Dataset Card for "google-gujarati"
More Information needed
google-speech-commands-wav2vec2-960hgoogle-britain-irelandDownloaded from OpenSLR
This data set contains transcribed high-quality audio of English sentences recorded by volunteers speaking different dialects of the language. The data set consists of wave files, and a TSV file (line_index.tsv). The file line_index.csv contains a line id, an anonymized FileID and the transcription of audio in the file. The recordings from the Welsh English speakers were collected in collaboration with Cardiff University. The data set contains the following number of… See the full description on the dataset page: https://huggingface.co/datasets/kth-tmh/google-britain-ireland.google-cloud-voice-mixkhmer-speech-large-english-google-translations
Dataset Card for khmer-speech-large-english-google-translation
Audio recordings of khmer speech with varying speakers and background noises.
English transcriptions were transcribed from the Khmer labels using Google Translate.
Based off of seanghay/khmer-speech-large.
Dataset Details
Dataset Sources
Huggingface: seanghay/khmer-speech-large
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/djsamseng/khmer-speech-large-english-google-translations.google-latam-spanish-boundary-normalized
Google LATAM Spanish Boundary-Normalized Audio
Female Spanish speech from the following upstream datasets:
Argentina: ylacombe/google-argentinian-spanish
Chile: ylacombe/google-chilean-spanish
Colombia: ylacombe/google-colombian-spanish
Attribution and Thanks
Many thanks to ylacombe for publishing
and maintaining the original Argentinian, Chilean, and Colombian Spanish
datasets. The recordings, transcripts, speaker labels, and original dataset
structure come… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-boundary-normalized.simplified-google-speech-commands-wav2vec2-960hgoogle-latam-spanish-uniform-vad
Google LATAM Spanish — Uniform VAD and 0.5 s Edges
This public derivative contains 15,016 Latin American Spanish utterances from
the Google crowdsourced TTS datasets packaged by ylacombe. Female and male
audio were freshly exported from the same pinned upstream revisions and passed
through exactly the same processing pipeline.
Configurations
Configuration
Train
Validation
Total
Hours including edge padding
argentina-female
3,542
379
3,921
4.181… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-uniform-vad.myanmar-speech-dataset-google-fleursPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Speech Dataset (Google Fleurs)
This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual Google Fleurs dataset.
For the complete multilingual dataset and additional information, please visit the original dataset repository
of Google Fleurs HuggingFace page.
Original Source
Fleurs is the speech version of the FLoRes machine translation benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-google-fleurs.SpeechCommandRecognition_GoogleSpeechCommandsV1
Dataset Card for "SpeechCommandRecognition_GoogleSpeechCommandsV1"
More Information needed
SpeechCommandRecognition_GoogleSpeechCommandsV1google_fleurs_plus_common_voice_11_ar
Dataset Card for "google_fleurs_plus_common_voice_11_ar"
More Information needed
google_waxal_dsgoogle-fleurs-te-romanizedgoogle-marathi
Dataset Card for "google-marathi"
More Information needed
google_myanmar_asr_voices
Google Myanmar ASR Dataset (WebDataset Version)
This repository provides a clean, user-friendly, and robust version of the Google Myanmar ASR Dataset, which is derived from the OpenSLR-80 Burmese Speech Corpus.
This version has been carefully re-processed into the WebDataset format. Each sample consists of a .wav audio file and a clean .json metadata file, packaged into sharded .tar archives. This format is highly efficient for large-scale training of ASR models.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/google_myanmar_asr_voices.hindi_google_fleursfork-google-openslr-javanesegoogle_waxal_ds_filteredSpeechCommandRecognition_GoogleSpeechCommandsV1_TTS
