CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TigreGotico /FalaBracarense_splitsdataset website: projectofalabracarense Licence CC - BY - NC - ND Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives audioautomatic-speech-recognition100K<n<1M0 likes1.6k downloads1y agoHugging Face02beshiribrahim /tigre-hubert-dataaudio1K<n<10K0 likes953 downloads2mo agoHugging Face03TigreGotico /not-wake-words-speech-en not-wake-words-speech-en Negative (non-wake-word) speech clips, used to measure false accepts for OVOS wake-word plugins. Derived from the Multilingual Spoken Words Corpus (MLCommons), which is built from Mozilla Common Voice and licensed CC-BY-4.0. This derivative keeps the same licence and attribution requirement. Produced with support from the NGI0 Commons Fund. audioaudio-classification10K<n<100K0 likes595 downloads19d agoHugging Face04TigreGotico /synthetic-wakeword-hey_computer synthetic-wakeword-hey_computer Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "hey computer". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.audioaudio-classification1K<n<10K0 likes404 downloads6d agoHugging Face05TigreGotico /ESC-50 ESC-50: Dataset for Environmental Sound Classification Overview | Download | Results | Repository content | License | Citing | Caveats | Changelog       The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification. The dataset consists of 5-second-long recordings organized into 50 semantical classes (with 40 examples per class) loosely arranged into 5 major categories:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/ESC-50.audio1K<n<10K0 likes390 downloads11mo agoHugging Face06TigreGotico /ambient_noisesaudio1K<n<10K0 likes191 downloads11mo agoHugging Face07TigreGotico /NAR NAR Dataset Author : Maxime Janvier maxime.janvier@gmail.com Lab : INRIA Rhones-Alpes Perception team (https://team.inria.fr/perception) Download link : https://team.inria.fr/perception/nard The NAR dataset is a set of audio recordings made with the humanoid robot Nao in real world conditions for supervised sound classification applications. All the recordings have been produced using the robot’s auditory devices and thus have the following characteristics : * recorded with… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/NAR.audion<1K0 likes184 downloads2y agoHugging Face08TigreGotico /FMA_3secs3 second clips extracted from https://github.com/mdeff/fma audion<1K0 likes147 downloads2y agoHugging Face09TigreGotico /synthetic-wakeword-hey_mycroft synthetic-wakeword-hey_mycroft Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "hey mycroft". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_mycroft.audioaudio-classification1K<n<10K0 likes144 downloads6d agoHugging Face10TigreGotico /public_domain_sounds_3secs Public Domain Sounds This is a backup of the 635 copyright-free sound recordings submitted to pdsounds.org before April 2009. all files split into 3 seconds chunks LICENSE NOTICE pdsounds.org - all sounds archive March 2, 2009 ALL 635 SOUNDS in this archive are entirely public domain and copyright free. No rights are reserved. The sounds were recorded by volunteers and uploaded to pdsounds.org with the condition that henceforth they are license-free and public domain.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/public_domain_sounds_3secs.audion<1K0 likes109 downloads2y agoHugging Face11TigreGotico /synthetic-wakeword-hey_siri synthetic-wakeword-hey_siri Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "hey siri". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on, including for… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_siri.audioaudio-classificationn<1K0 likes103 downloads19d agoHugging Face12TigreGotico /synthetic-wakeword-wake_up synthetic-wakeword-wake_up Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "wake up". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on, including for… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-wake_up.audioaudio-classification1K<n<10K0 likes96 downloads6d agoHugging Face13TigreGotico /synthetic-wakeword-voice_assistant synthetic-wakeword-voice_assistant Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "voice assistant". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-voice_assistant.audioaudio-classification1K<n<10K0 likes94 downloads19d agoHugging Face14TigreGotico /synthetic-wakeword-home_assistant synthetic-wakeword-home_assistant Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "home assistant". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-home_assistant.audioaudio-classification1K<n<10K0 likes88 downloads19d agoHugging Face15TigreGotico /synthetic-wakeword-alexa synthetic-wakeword-alexa Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "alexa". Every clip is machine-generated: text-to-speech synthesis followed by voice conversion to simulate multiple speakers. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on, including for model… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-alexa.audioaudio-classification1K<n<10K0 likes86 downloads19d agoHugging Face16TigreGotico /synapseoriginal data: http://corpora.ugr.es/synapse audio1K<n<10K0 likes76 downloads8mo agoHugging Face17Harbidel /tigrinya-asr-mergedgated tigrinya-asr-merged A merged Tigrinya speech-recognition dataset, combining and deduplicating: badrex/tigrinya-speech (train pool) google/WaxalNLP config tir_asr (train pool) UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set) Processing Standardized to audio (16kHz mono) and text columns, with a source column tracking origin Unicode NFC-normalized transcripts, empty transcripts dropped Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.audioautomatic-speech-recognition10K<n<100K0 likes73 downloads22d agoHugging Face18TigreGotico /synthetic-wakeword-hey_jarvis synthetic-wakeword-hey_jarvis Synthetic wake-word audio for training and benchmarking OVOS wake-word plugins, covering the phrase "hey jarvis". Every clip is machine-generated text-to-speech. No human recording is included, and no natural voice is reproduced. Machine-generated audio carries no copyright of its own, so this dataset is published CC-BY-4.0 and is free to use, redistribute and build on, including for model training. Produced with support from the NGI0 Commons Fund.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_jarvis.audioaudio-classification1K<n<10K0 likes71 downloads6d agoHugging Face19TigreGotico /ArquivoDialetalCLUPdataset info: https://cl.up.pt/arquivo/ license CC BY-NC-ND audioautomatic-speech-recognitionn<1K0 likes54 downloads1y agoHugging Face20badrex /tigrinya-speechaudio10K<n<100K0 likes53 downloads11mo agoHugging Face21TigreGotico /SpokenPortugueseGeographicalSocialVarieties Spoken Portuguese - Geographical and Social Varieties dataset source: https://www.clul.ulisboa.pt (1995-1997 - European Commission DGXXII, Programme LINGUA/SOCRATES) The project is concluded and the materials are published in CD-ROM, with the exclusive publishing support of Instituto Camões, under the title Português Falado - Documentos Autênticos: Gravações áudio com transcrição alinhada. Its distribution outside of Portugal is ensured by Instituto Camões and in Portugal by CLUL.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/SpokenPortugueseGeographicalSocialVarieties.audioautomatic-speech-recognitionn<1K0 likes50 downloads1y agoHugging Face22TigreGotico /building_106_kitchen_3secsMirror of https://www.csc.kth.se/~jastork/pages/datasets.html all sounds split in 3 seconds chunks meant for usage as background noise or environmental event detection audioaudio-classification1K<n<10K0 likes37 downloads2y agoHugging Face23TigreGotico /madisonsource: http://teitok.clul.ul.pt/madison/pt/index.php? audion<1K0 likes33 downloads8mo agoHugging Face24TigreGotico /Terreiro-de-la-LhenguaAudio: Terreiro de la Lhéngua 25 (podcast) Text: La fala Screbida (ebook) intro/outro has been cropped for this dataset audion<1K0 likes31 downloads10mo agoHugging Face25TigreGotico /VocativesEuropeanPortuguesedataset from https://www.clul.ulisboa.pt/en/recurso/vocatives-european-portuguese This corpus was originally a corpus created for a study concerning with vocatives in European Portuguese. The main goal of this study was to analyze some prosodic features of the vocative in European Portuguese and their relation with the syntactic distribution (initial, medial, final) of these constituents. The corpus has 432 audio files. This number results from the recording of 108 sentences (54 target… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/VocativesEuropeanPortuguese.audioautomatic-speech-recognitionn<1K1 likes27 downloads1y agoHugging Face26TigreGotico /locallingua_ptRecordings from Portugal downloaded from https://localingual.com audioautomatic-speech-recognitionn<1K0 likes26 downloads2y agoHugging Face27TigreGotico /InstitutoCamoesdownloaded from https://www.instituto-camoes.pt audioautomatic-speech-recognitionn<1K1 likes23 downloads1y agoHugging Face28TigreGotico /compare-accents-ptsmall dataset of multiple portuguese speakers from various dialects speaking the same sentence "Dom Sebastião I era o décimo-sexto Rei de Portugal, e sétimo da Dinastia de Avis. Era neto do rei João III, tornou-se herdeiro do trono depois da morte do seu pai, o príncipe João de Portugal duas semanas antes do seu nascimento, e rei com apenas três anos, em 1557. Em virtude de ser um herdeiro tão esperado para dar continuidade à Dinastia de Avis, ficou conhecido como O Desejado; alternativamente… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/compare-accents-pt.audioautomatic-speech-recognitionn<1K1 likes22 downloads2y agoHugging Face29TigreGotico /MdMvaudion<1K0 likes19 downloads1y agoHugging Face30TigreGotico /speech_MASSIVE_pt-PTpt-PT subset from FBK-MT/Speech-MASSIVE audioautomatic-speech-recognition1K<n<10K1 likes18 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.