datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fleurscommon_voice_21_0_minifsdkaggle2019-parquet
FSDKaggle2019
FSDKaggle2019[1] is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology.
FSDKaggle2019 has been used for the DCASE Challenge 2019 Task 2, which was run as a Kaggle competition titled Freesound Audio Tagging 2019.
All audio clips are provided as uncompressed PCM 16 bit, 44.1 kHz, mono audio files.
This version of database could be found and downloaded from here.
Data Split Statistics
Curated
Noisy
Test… See the full description on the dataset page: https://huggingface.co/datasets/mteb/fsdkaggle2019-parquet.crema-dVoxPopuliAccentPairClassificationFilteredsib-fleurs-multilingual-miniClothoBirdSetGTZANAudioRerankingjam-alt-linesUrbansound8K_t2a
Dataset Card for "Urbansound8K_t2a"
More Information needed
OmniCVRiemocapSpeechCommandsZeroshotv0.02
SpeechCommandsZeroshotv0.02
An MTEB dataset
Massive Text Embedding Benchmark
Sound Classification/Keyword Spotting Dataset. This is a set of one-second audio clips containing a single spoken English word or background noise. These words are from a small set of commands such as 'yes', 'no', and 'stop' spoken by various speakers. With a total of 10 labels/commands for keyword spotting and a total of 30 labels for other auxiliary tasks
Task category
a2t
Domains
Spoken… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SpeechCommandsZeroshotv0.02.CREMADPairClassificationMACS_t2a
Dataset Card for "MACS_t2a"
More Information needed
gigaspeech_t2acommonlanguage-age-miniVehicle_sounds_classification_datasetbirdclef25-miniNMSQAPairClassificationspoken-squad-t2agtzan-genrebeijing-operavoxceleb-sentimentminds14-multilingualmridingham-tonicRavdessZeroshot
RavdessZeroshot
An MTEB dataset
Massive Text Embedding Benchmark
Emotion classification Dataset. RAVDESS contains 24 professional actors (12 female, 12 male), vocalizing two lexically-matched statements in a neutral North American accent. Speech emotions includes neutral,calm, happy, sad, angry, fearful, surprise, and disgust expressions. These 8 emtoions also serve as labels for the dataset.
Task category
a2t
Domains
Spoken
Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/RavdessZeroshot.voxpopuli-minimini-voxpopuli
