datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fma-labeled
FMA Labeled — Multi-Attribute Music Dataset
🏆 Submitted to the Uncharted Data Challenge
hosted by Adaption Labs — credit to
Adaptive Data by Adaption for organizing the hackathon.
A large-scale labeled music dataset built on top of the Creative-Commons
subset of the Free Music Archive (FMA). Every
track has been automatically annotated with lyrics, genre, mood, instruments,
tempo, key, and more using Google Gemini (gemini-flash-latest).
Intended for training and evaluating music… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/fma-labeled.common_voice_13_0_dv_preprocessed
Dataset Card for Common Voice Corpus 13.0
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 27141 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 17689 validated hours in 108 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/fmagot01/common_voice_13_0_dv_preprocessed.
