CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.audioautomatic-speech-recognition100K<n<1M468 likes100k downloads4mo agoHugging Face02google /svq Simple Voice Questions Simple Voice Questions (SVQ) is a set of short audio questions recorded in 26 locales across 17 languages under multiple audio conditions. It serves as a core evaluation componenet for Massive Sound Embedding Benchmark (MSEB). Technical Specifications Feature Details Locales 26 Languages 17 Total Speakers ~700 (Capped at 250 recordings per speaker) Audio Conditions Clean, Background Speech, Media, Traffic Noise Gender… See the full description on the dataset page: https://huggingface.co/datasets/google/svq.audioquestion-answering1M<n<10M60 likes71k downloads2d agoHugging Face03google /WaxalNLP Waxal Datasets The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus. Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/google/WaxalNLP.audioautomatic-speech-recognition1M<n<10M286 likes13k downloads24d agoHugging Face04mazkooleg /0-9up_google_speech_commands_augmented_raw Dataset Card for "google_speech_commands_augmented_raw_fixed" More Information needed audio1M<n<10M0 likes3.4k downloads4y agoHugging Face05zhaochenyang20 /googletime googletime Audio validation set organized as validation/audio/* plus validation/metadata.jsonl. The metadata contains only file_name and transcription; transcriptions include timestamp and speaker markers. audion<1K0 likes862 downloads2mo agoHugging Face06ylacombe /google-chilean-spanish Dataset Card for Tamil Speech Dataset Summary This dataset consists of 7 hours of transcribed high-quality audio of Chilean Spanish sentences recorded by 31 volunteers. The dataset is intended for speech technologies. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. Supported Tasks text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS). automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-chilean-spanish.audiotext-to-speech1K<n<10K24 likes309 downloads3y agoHugging Face07Hunzla /simplified_google_speech_commands_wav2vec2_960haudio10K<n<100K1 likes255 downloads3y agoHugging Face08ylacombe /google-tamil Dataset Card for Tamil Speech Dataset Summary This dataset consists of 7 hours of transcribed high-quality audio of Tamil sentences recorded by 50 volunteers. The dataset is intended for speech technologies. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. Supported Tasks text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS). automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-tamil.audiotext-to-speech1K<n<10K7 likes251 downloads3y agoHugging Face09ylacombe /google-colombian-spanish Dataset Card for "google-colombian-spanish" More Information needed audio1K<n<10K14 likes247 downloads3y agoHugging Face10ittailup /google-la-voices Dataset Card for "google-la-voices" Speaker Durations Speaker Duration (seconds) 00295 1606.144 00610 7026.261 01208 3284.907 01523 6309.888 02121 4687.445 02436 4654.080 02484 9379.925 02485 130.219 03034 5186.048 03349 5143.381 03397 7852.203 03398 118.101 03853 638.037 04310 8260.437 04311 105.472 04766 590.165 05223 8257.773 05679 846.251 0613610207.707 06592 863.659 07049 7580.715 07060 575.659 07505 1743.531… See the full description on the dataset page: https://huggingface.co/datasets/ittailup/google-la-voices.audio1K<n<10K0 likes221 downloads2y agoHugging Face11ylacombe /google-argentinian-spanish Dataset Card for "google-argentinian-spanish" More Information needed audio1K<n<10K19 likes179 downloads3y agoHugging Face12ylacombe /google-gujarati Dataset Card for "google-gujarati" More Information needed audio1K<n<10K2 likes121 downloads3y agoHugging Face13Hunzla /google-speech-commands-wav2vec2-960haudio10K<n<100K1 likes103 downloads3y agoHugging Face14djsamseng /khmer-speech-large-english-google-translations Dataset Card for khmer-speech-large-english-google-translation Audio recordings of khmer speech with varying speakers and background noises. English transcriptions were transcribed from the Khmer labels using Google Translate. Based off of seanghay/khmer-speech-large. Dataset Details Dataset Sources Huggingface: seanghay/khmer-speech-large Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/djsamseng/khmer-speech-large-english-google-translations.audioautomatic-speech-recognition10K<n<100K5 likes100 downloads6mo agoHugging Face15groxaxo /google-latam-spanish-uniform-vad Google LATAM Spanish — Uniform VAD and 0.5 s Edges This public derivative contains 15,016 Latin American Spanish utterances from the Google crowdsourced TTS datasets packaged by ylacombe. Female and male audio were freshly exported from the same pinned upstream revisions and passed through exactly the same processing pipeline. Configurations Configuration Train Validation Total Hours including edge padding argentina-female 3,542 379 3,921 4.181… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-uniform-vad.audiotext-to-speech10K<n<100K0 likes98 downloads5d agoHugging Face16unlimitedbytes /google-cloud-voice-mixaudio1K<n<10K3 likes94 downloads1y agoHugging Face17chuuhtetnaing /myanmar-speech-dataset-google-fleursPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Speech Dataset (Google Fleurs) This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual Google Fleurs dataset. For the complete multilingual dataset and additional information, please visit the original dataset repository of Google Fleurs HuggingFace page. Original Source Fleurs is the speech version of the FLoRes machine translation benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-google-fleurs.audiotext-to-speech1K<n<10K0 likes87 downloads1y agoHugging Face18groxaxo /google-latam-spanish-boundary-normalized Google LATAM Spanish Boundary-Normalized Audio Female Spanish speech from the following upstream datasets: Argentina: ylacombe/google-argentinian-spanish Chile: ylacombe/google-chilean-spanish Colombia: ylacombe/google-colombian-spanish Attribution and Thanks Many thanks to ylacombe for publishing and maintaining the original Argentinian, Chilean, and Colombian Spanish datasets. The recordings, transcripts, speaker labels, and original dataset structure come… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-boundary-normalized.audiotext-to-speech1K<n<10K1 likes79 downloads3mo agoHugging Face19kth-tmh /google-britain-irelandDownloaded from OpenSLR This data set contains transcribed high-quality audio of English sentences recorded by volunteers speaking different dialects of the language. The data set consists of wave files, and a TSV file (line_index.tsv). The file line_index.csv contains a line id, an anonymized FileID and the transcription of audio in the file. The recordings from the Welsh English speakers were collected in collaboration with Cardiff University. The data set contains the following number of… See the full description on the dataset page: https://huggingface.co/datasets/kth-tmh/google-britain-ireland.audio10K<n<100K0 likes74 downloads7mo agoHugging Face20Hunzla /simplified-google-speech-commands-wav2vec2-960haudio10K<n<100K0 likes72 downloads3y agoHugging Face21MohammadJamalaldeen /google_fleurs_plus_common_voice_11_ar Dataset Card for "google_fleurs_plus_common_voice_11_ar" More Information needed audio10K<n<100K0 likes45 downloads4y agoHugging Face22DynamicSuperb /SpeechCommandRecognition_GoogleSpeechCommandsV1 Dataset Card for "SpeechCommandRecognition_GoogleSpeechCommandsV1" More Information needed audion<1K0 likes44 downloads3y agoHugging Face23AhunInteligence /google_waxal_dsaudio10K<n<100K0 likes33 downloads7mo agoHugging Face24ylacombe /google-marathi Dataset Card for "google-marathi" More Information needed audio1K<n<10K3 likes30 downloads3y agoHugging Face25freococo /google_myanmar_asr_voices Google Myanmar ASR Dataset (WebDataset Version) This repository provides a clean, user-friendly, and robust version of the Google Myanmar ASR Dataset, which is derived from the OpenSLR-80 Burmese Speech Corpus. This version has been carefully re-processed into the WebDataset format. Each sample consists of a .wav audio file and a clean .json metadata file, packaged into sharded .tar archives. This format is highly efficient for large-scale training of ASR models.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/google_myanmar_asr_voices.audioautomatic-speech-recognition1K<n<10K0 likes28 downloads1y agoHugging Face26jayasuryajsk /google-fleurs-te-romanizedaudio1K<n<10K0 likes26 downloads2y agoHugging Face27octava /fork-google-openslr-javaneseaudio1K<n<10K0 likes20 downloads2y agoHugging Face28Ritwika03 /hindi_google_fleursaudio1K<n<10K0 likes18 downloads1y agoHugging Face29DynamicSuperbPrivate /SpeechCommandRecognition_GoogleSpeechCommandsV1_TTSaudion<1K0 likes14 downloads2y agoHugging Face30yayossd /google-chilean-spanish Dataset Card for Tamil Speech Dataset Summary This dataset consists of 7 hours of transcribed high-quality audio of Chilean Spanish sentences recorded by 31 volunteers. The dataset is intended for speech technologies. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. Supported Tasks text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS).… See the full description on the dataset page: https://huggingface.co/datasets/yayossd/google-chilean-spanish.audiotext-to-speech1K<n<10K0 likes12 downloads1d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.