datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NaturalVoices_VC_0.1 NaturalVoices VC 10%
A large voice conversion (VC) dataset curated from spontaneous, in-the-wild podcast speech as part of the NaturalVoices project in collaboration with 🤗MSP Lab at CMU LTI. This release provides the 10% subset uniformly sampled from 870-hour VC dataset and subsets mainly intended for training and evaluating emotion-aware voice conversion systems but not limited to VC tasks.
📄 Paper: NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice… See the full description on the dataset page: https://huggingface.co/datasets/JHU-SmileLab/NaturalVoices_VC_0.1.SMIIP-NV
SMIIP-NV: A Multi-Annotation Non-Verbal Expressive Speech Corpus
Dataset Description
SMIIP-NV is a multi-annotated non-verbal expressive speech corpus designed for training and evaluating LLM-based text-to-speech systems. It contains a diverse set of non-verbal sounds (e.g., laughter, crying, coughing) along with emotion labels (happy, sad, neutral, angry, surprised), enabling the synthesis of natural and expressive speech.
Key Features:
Multi-dimensional… See the full description on the dataset page: https://huggingface.co/datasets/xunyi/SMIIP-NV.AISHELL6-Whisper
🗣️ AISHELL6-Whisper
AISHELL6-Whisper is a large-scale open-source Chinese Mandarin audio-visual whisper speech dataset,containing 30 hours each of whisper and parallel normal speech, with synchronized frontal RGB facial videos.
📘 Dataset Summary
Property
Description
Language
Chinese (Mandarin, ZH)
License
CC BY-NC-SA 4.0
Duration
~60 hours total (30 h whisper + 30 h normal)
Speakers
167 total (121 with RGB-D, 46 audio-only)
Environment
Controlled… See the full description on the dataset page: https://huggingface.co/datasets/SMIIP-lab/AISHELL6-Whisper.OILThe Online Indonesian Learning (OIL) Dataset
The Online Indonesian Learning (OIL) dataset or corpus currently contains lessons from three Indonesian teachers who have posted content on YouTube.
For further details please see Zara Maxwell-Smith and Ben Foley, (forthcoming), Automated speech recognition of Indonesian-English language lessons on YouTube using transfer learning, Field Matters Workshop, EACL 2023
How to cite this dataset.
Please use the following .bib to reference this work.… See the full description on the dataset page: https://huggingface.co/datasets/ZMaxwell-Smith/OIL.AISHELL8-RealScene
📘 Dataset Summary
Property
Description
Language
Chinese (Mandarin, ZH)
License
CC BY-NC-SA 4.0
Total Duration
102.19 hours
Speakers
171 foreground speakers
Scenes
5 real-world locations
Recording Style
Conversational speech
Audio
Near-field + far-field
Video
Multi-view RGB facial video
Sampling Rate
16 kHz
Far-field Audio
8-channel
Video Resolution
256×256 @ 25 fps
🎙️ Dataset Description
AISHELL8-RealScene is a public… See the full description on the dataset page: https://huggingface.co/datasets/SMIIP-lab/AISHELL8-RealScene.ami-1s-ftbook_of_genesis_DIA
Dataset Card for Dataset Name
Dataset Details
Dataset Description
This dataset was created using the English translation of the Book of Genesis via the Vatican's website.
The audio was generated using Nari Lab's DIA TTS model, where the text that was used was complete sentences from the Book of Genesis that roughly equaled 140-560 characters in length, or roughly 5 - 20 seconds.
Enjoy!
Curated by: Tyler Smith
Language(s) (NLP): English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SmittyB00p/book_of_genesis_DIA.ami-ft-over3s30c-incl-fwSMILE
