datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FalaBracarense_splitsdataset website: projectofalabracarense
Licence
CC - BY - NC - ND
Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives
tigre-hubert-datanot-wake-words-speech-en
not-wake-words-speech-en
Negative (non-wake-word) speech clips, used to measure false accepts for OVOS
wake-word plugins.
Derived from the Multilingual Spoken Words Corpus
(MLCommons), which is built from Mozilla Common Voice and licensed CC-BY-4.0.
This derivative keeps the same licence and attribution requirement.
Produced with support from the NGI0 Commons Fund.
synthetic-wakeword-hey_computer
synthetic-wakeword-hey_computer
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey computer".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_computer.ESC-50
ESC-50: Dataset for Environmental Sound Classification
Overview | Download | Results | Repository content | License | Citing | Caveats | Changelog
The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification.
The dataset consists of 5-second-long recordings organized into 50 semantical classes (with 40 examples per class) loosely arranged into 5 major categories:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/ESC-50.ambient_noisesNAR
NAR Dataset
Author : Maxime Janvier maxime.janvier@gmail.com
Lab : INRIA Rhones-Alpes Perception team (https://team.inria.fr/perception)
Download link : https://team.inria.fr/perception/nard
The NAR dataset is a set of audio recordings made with the humanoid robot Nao in real world conditions for supervised sound classification applications. All the recordings have been produced using the robot’s auditory devices and thus have the following characteristics :
* recorded with… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/NAR.FMA_3secs3 second clips extracted from https://github.com/mdeff/fma
synthetic-wakeword-hey_mycroft
synthetic-wakeword-hey_mycroft
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey mycroft".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_mycroft.public_domain_sounds_3secs
Public Domain Sounds
This is a backup of the 635 copyright-free sound recordings submitted to pdsounds.org before April 2009.
all files split into 3 seconds chunks
LICENSE NOTICE
pdsounds.org - all sounds archive
March 2, 2009
ALL 635 SOUNDS in this archive are entirely public domain and copyright free.
No rights are reserved.
The sounds were recorded by volunteers and uploaded to pdsounds.org with the
condition that henceforth they are license-free and public domain.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/public_domain_sounds_3secs.synthetic-wakeword-hey_siri
synthetic-wakeword-hey_siri
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey siri".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on, including for… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_siri.synthetic-wakeword-wake_up
synthetic-wakeword-wake_up
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "wake up".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on, including for… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-wake_up.synthetic-wakeword-voice_assistant
synthetic-wakeword-voice_assistant
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "voice assistant".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-voice_assistant.synthetic-wakeword-home_assistant
synthetic-wakeword-home_assistant
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "home assistant".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-home_assistant.synthetic-wakeword-alexa
synthetic-wakeword-alexa
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "alexa".
Every clip is machine-generated: text-to-speech synthesis followed by voice
conversion to simulate multiple speakers. No human recording is included, and
no natural voice is reproduced. Machine-generated audio carries no copyright
of its own, so this dataset is published CC-BY-4.0 and is free to use,
redistribute and build on, including for model… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-alexa.synapseoriginal data: http://corpora.ugr.es/synapse
tigrinya-asr-merged
tigrinya-asr-merged
A merged Tigrinya speech-recognition dataset, combining and deduplicating:
badrex/tigrinya-speech (train pool)
google/WaxalNLP config tir_asr (train pool)
UBC-NLP/SimbaBench_dataset config asr_test_tir (held-out benchmark test set)
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed from the train… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/tigrinya-asr-merged.synthetic-wakeword-hey_jarvis
synthetic-wakeword-hey_jarvis
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey jarvis".
Every clip is machine-generated text-to-speech. No human recording is
included, and no natural voice is reproduced. Machine-generated audio carries
no copyright of its own, so this dataset is published CC-BY-4.0 and is free
to use, redistribute and build on, including for model training.
Produced with support from the NGI0 Commons Fund.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_jarvis.ArquivoDialetalCLUPdataset info: https://cl.up.pt/arquivo/
license CC BY-NC-ND
tigrinya-speechSpokenPortugueseGeographicalSocialVarieties
Spoken Portuguese - Geographical and Social Varieties
dataset source: https://www.clul.ulisboa.pt
(1995-1997 - European Commission DGXXII, Programme LINGUA/SOCRATES)
The project is concluded and the materials are published in CD-ROM, with the exclusive publishing support of Instituto Camões, under the title Português Falado - Documentos Autênticos: Gravações áudio com transcrição alinhada. Its distribution outside of Portugal is ensured by Instituto Camões and in Portugal by CLUL.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/SpokenPortugueseGeographicalSocialVarieties.building_106_kitchen_3secsMirror of https://www.csc.kth.se/~jastork/pages/datasets.html
all sounds split in 3 seconds chunks
meant for usage as background noise or environmental event detection
madisonsource: http://teitok.clul.ul.pt/madison/pt/index.php?
Terreiro-de-la-LhenguaAudio: Terreiro de la Lhéngua 25 (podcast)
Text: La fala Screbida (ebook)
intro/outro has been cropped for this dataset
VocativesEuropeanPortuguesedataset from https://www.clul.ulisboa.pt/en/recurso/vocatives-european-portuguese
This corpus was originally a corpus created for a study concerning with vocatives in European Portuguese. The main goal of this study was to analyze some prosodic features of the vocative in European Portuguese and their relation with the syntactic distribution (initial, medial, final) of these constituents.
The corpus has 432 audio files. This number results from the recording of 108 sentences (54 target… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/VocativesEuropeanPortuguese.locallingua_ptRecordings from Portugal downloaded from https://localingual.com
InstitutoCamoesdownloaded from https://www.instituto-camoes.pt
compare-accents-ptsmall dataset of multiple portuguese speakers from various dialects speaking the same sentence
"Dom Sebastião I era o décimo-sexto Rei de Portugal, e sétimo da Dinastia de Avis. Era neto do rei João III, tornou-se herdeiro do trono depois da morte do seu pai, o príncipe João de Portugal duas semanas antes do seu nascimento, e rei com apenas três anos, em 1557. Em virtude de ser um herdeiro tão esperado para dar continuidade à Dinastia de Avis, ficou conhecido como O Desejado; alternativamente… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/compare-accents-pt.MdMvspeech_MASSIVE_pt-PTpt-PT subset from FBK-MT/Speech-MASSIVE
