datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test_librispeech_parquetnsynth-parquetfsdkaggle2019-parquet
FSDKaggle2019
FSDKaggle2019[1] is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology.
FSDKaggle2019 has been used for the DCASE Challenge 2019 Task 2, which was run as a Kaggle competition titled Freesound Audio Tagging 2019.
All audio clips are provided as uncompressed PCM 16 bit, 44.1 kHz, mono audio files.
This version of database could be found and downloaded from here.
Data Split Statistics
Curated
Noisy
Test… See the full description on the dataset page: https://huggingface.co/datasets/mteb/fsdkaggle2019-parquet.esc50-parquetfsdkaggle2019-parquet
FSDKaggle2019
FSDKaggle2019[1] is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology.
FSDKaggle2019 has been used for the DCASE Challenge 2019 Task 2, which was run as a Kaggle competition titled Freesound Audio Tagging 2019.
All audio clips are provided as uncompressed PCM 16 bit, 44.1 kHz, mono audio files.
This version of database could be found and downloaded from here.
Data Split Statistics
Curated
Noisy
Test… See the full description on the dataset page: https://huggingface.co/datasets/confit/fsdkaggle2019-parquet.librispeech_parquetDDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translateDisclaimer: The original dataset can be found here.
It is published by Digital Divide Data Cambodia (DDD-Cambodia).
License:
Khmer ASR Cultural Dataset's license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0).
Please attribute Digital Divide Data if you use this dataset in any way.
Objective of this dataset
Add English translation: a new column en_translate is added to the original dataset (only from parquet 000 to 159 of the original… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translate.cremad-parquetsimchoir-parquet
FastMSS synthetic multi-speaker meetings - parquet edition
Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring.
Subsets and splits
debug — splits: train — 1 mixtures, 1.6 min total, 6 unique speakers… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/simchoir-parquet.gtzan-parquet
GTZAN Music Genre Classification
GTZAN consists of 100 30-second recording excerpts in each of 10 categories, and is the most-used public dataset in music information retrieval (MIR) research.
Following Kereliuk et al. (2015), we use the "fault-filtered" partitioning version of GTZAN, which is constructed by hand to include 443/197/290 excerpts.
This version of database could be found and downloaded from here.
Citations
@article{kereliuk2015deep,
title={Deep… See the full description on the dataset page: https://huggingface.co/datasets/confit/gtzan-parquet.ravdess-parquetvctk-corpus-en-parquetmunch-1-latent-NEW-parquet
🎙️ Urdu TTS Latent Dataset — munch-1-latent-NEW-parquet
Pre-computed DACVAE latent representations for 51,021 Urdu utterances, ready for TTS model training. No audio decoding required at training time — load the dataset, reshape the binary blob, and train.
Source
Field
Value
Source audio
Humair332/Urdu-munch-1
Codec
Aratako/Semantic-DACVAE-Japanese-32dim
Codec sample rate
48,000 Hz
Encoder hop size
1,920 samples
Latent frame rate
25.0 Hz
Latent dim… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/munch-1-latent-NEW-parquet.mswc-parquetcommon_voice_22_0_toki_pona_parquet
Common Voice 22.0 - Toki Pona Subset!
My own Parquet conversion of Toki Pona's subset of Fsicoli's reupload of Common Voice 22 so we don't have to downgrade to Datasets 3.6 anymore!
Why?
Because the original dataset required Hugging Face Datasets 3.6 or older because it has Python code and it's in TAR shards.
This is in Parquet and works with any recent version of Hugging Face Datasets!
Details
Dataset Structure
DatasetDict({… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/common_voice_22_0_toki_pona_parquet.wmms-parquet
Watkins Marine Mammal Sound (WMMS) Database
Sound files on this website are free to download for personal or academic (not commercial) use.
Sound files and associated metadata are credited as follows: "Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
Database could be found and downloaded from here.
In this database version, the audio archive includes sounds of 32 species:
Atlantic_Spotted_Dolphin
Bearded_Seal
Beluga… See the full description on the dataset page: https://huggingface.co/datasets/confit/wmms-parquet.pianos-parquet
Pianos Sound Quality Dataset
This version of dataset comprises seven models of pianos:
Kawai upright piano
Kawai grand piano
Young Change upright piano
Hsinghai upright piano
Grand Theatre Steinway piano
Steinway grand piano
Pearl River upright piano.
Note: the paper (Zhou et al., 2023) only uses the first 7 piano classes in the dataset, its future work has finished the 8-class evaluation.
License
MIT License
Copyright (c) CCMUSIC
Permission is hereby granted… See the full description on the dataset page: https://huggingface.co/datasets/confit/pianos-parquet.one_voice_FACEBOOK_PARQUET
Artificial Omnivoice Hungarian Speaker Dataset
Ez egy teljesen szintetikus magyar nyelvű beszédadatbázis, amely kiváló minőségű szövegfelolvasó (TTS) és beszédfelismerő (ASR) modellek tanításához és finomhangolásához készült.
Adatforrás és Referencia Hang
A dataset alapjául szolgáló referencia hang (speaker identity) egy 20 másodperces részlet az alábbi YouTube videóból:
Forrás: Hogyan legyél tökéletes magyar várvédő tutorial
Licenc: A videó CC (Creative Commons)… See the full description on the dataset page: https://huggingface.co/datasets/fablevi/one_voice_FACEBOOK_PARQUET.my_parquet_dataset_2afri-temp-data5-parquetlibrispeech-sid-parquet
LibriSpeech Speaker Identification
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey.
The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
However, although LibriSpeech is very popular in ASR tasks, we use LibriSpeech database as a speaker identification task.
We follow SincNet paper official split for training and evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/confit/librispeech-sid-parquet.wmms-parquet
Watkins Marine Mammal Sound (WMMS) Database
Sound files on this website are free to download for personal or academic (not commercial) use.
Sound files and associated metadata are credited as follows: "Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
Database could be found and downloaded from here.
In this database version, the audio archive includes sounds of 32 species:
Atlantic_Spotted_Dolphin
Bearded_Seal
Beluga… See the full description on the dataset page: https://huggingface.co/datasets/qdskipper/wmms-parquet.emodb-parquet
EmoDB
The EmoDB is the freely available German emotional database, containing a total of 535 utterances.
It comprises of seven emotions: 1) anger; 2) boredom; 3) anxiety; 4) happiness; 5) sadness; 6) disgust; and 7) neutral.
The data was recorded at a 48-kHz sampling rate and then down-sampled to 16-kHz.
We follow the unofficial speaker-independent train/test split from here.
Citations
@inproceedings{burkhardt2005database,
title={A database of German emotional… See the full description on the dataset page: https://huggingface.co/datasets/confit/emodb-parquet.dhivehi-javaabu-speech-parquetmy_parquet_dataset_13FeruzaSpeech_parquet_dataset
Dataset Card for Dataset Name
FeruzaSpeech is a read speech dataset
of the Uzbek language, transcribed in both Cyrillic
and Latin alphabets, freely available for academic research pur-
poses. It includes 60 hours of high-quality recordings
from a single native female speaker from Tashkent, Uzbekistan.
ICNLSPConference: https://www.youtube.com/watch?v=9whj9yzI_s4&ab_channel=ICNLSPConference
Paper: https://arxiv.org/abs/2410.00035
python veiw.py… See the full description on the dataset page: https://huggingface.co/datasets/k2speech/FeruzaSpeech_parquet_dataset.librispeech_asr_parquetwhiser_parquet
Dataset Description
Motivation
Enable both categorical emotion recognition and dimensional affect regression on speech clips with multiple crowd-sourced annotations per instance.
Study annotator variability, label aggregation, and uncertainty.
Composition
Audio: WAV files in wavs/
Annotations:
Raw: WHiSER/Labels/labels.txt (per-file summary + multiple worker lines)
Aggregated: WHiSER/Labels/labels_consensus.{csv,json}
Per-annotator:… See the full description on the dataset page: https://huggingface.co/datasets/Lab-MSP/whiser_parquet.prueba_parquetThis is an example of a repository with parquet files only.
gtzan-parquet
GTZAN Music Genre Classification
GTZAN consists of 100 30-second recording excerpts in each of 10 categories, and is the most-used public dataset in music information retrieval (MIR) research.
Following Kereliuk et al. (2015), we use the "fault-filtered" partitioning version of GTZAN, which is constructed by hand to include 443/197/290 excerpts.
This version of database could be found and downloaded from here.
Citations
@article{kereliuk2015deep,
title={Deep… See the full description on the dataset page: https://huggingface.co/datasets/Dora73/gtzan-parquet.
