datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vctk
Dataset Card for "vctk"
More Information needed
vctk
VCTK
This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning.
The original dataset is available at: https://datashare.ed.ac.uk/handle/10283/3443.
Reproducing
This repository notably lacks a requirements.txt file. There's likely a missing dependency or two, but roughly:
pydub
tqdm
torch
torchaudio… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/vctk.VCTKVCB-Bench
VCB-Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
Introduction
Voice Chat Bot Bench (VCB Bench) is a high-quality Chinese benchmark built entirely on real human speech. It evaluates large audio language models (LALMs) along three complementary dimensions:
(1) Instruction following: Text Instruction Following (TIF), Speech Instruction Following (SIF), English Text Instruction Following (TIF-En)… See the full description on the dataset page: https://huggingface.co/datasets/tencent/VCB-Bench.vctk-fullvctk_resampled_16k_balancednoisy_vctk_16k_synth
Dataset Card for "noisy_vctk_16k_synth"
More Information needed
vctk-48khz
Dataset Card for VCTK (48kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.vctk
VCTK
This is a mirror of the VCTK Corpus.
The original files were converted from FLAC to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
train_mic1: 90 speakers, 33.6 hours, 35987 utterances
train_mic2: 90 speakers, 33.6 hours, 35987 utterances
val_mic1: 10 random speakers unseen during training: p238, p244, p254, p263, p265, p272, p288, p294, p305, and p335. 4.0 hours, 4179 utterances.
val_mic2: Same speakers as val_mic1.… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/vctk.vctk-ttsvctk-16khz
Dataset Card for VCTK (16kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.VCDB-Core-AudioVideo
VCDB Core Audio-Video Retrieval
This repository packages synchronized video and extracted audio from the
528-video core set of VCDB as a symmetric video+audio-to-video+audio
retrieval task for MTEB/MOEB. The separate 100,000-video background collection
is not included.
Terms and provenance
The source dataset is provided by Fudan University for research purposes
only. The source authors and Fudan University make no warranties about the
dataset, including… See the full description on the dataset page: https://huggingface.co/datasets/pranitchawla/VCDB-Core-AudioVideo.Evaluation-Multilingual-VC
Evaluation-Multilingual-VC
We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon,
Filter languages that support by Whisper Large V3 to evaluate WER automatically,
Only take test set, sort by up votes.
Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows.
Only build first 500 rows for each language
Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.vctk_p225_allcolsvctk
Dataset Card for VCTK
Dataset Summary
This CSTR VCTK Corpus includes around 44-hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out about 400 sentences, which were selected from a newspaper, the rainbow passage and an elicitation paragraph used for the speech accent archive.
Supported Tasks
automatic-speech-recognition, speaker-identification: The dataset can be used to train a model for Automatic Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/kth-tmh/vctk.mixed-language-detection-english-accented-vc
Mixed-Language Speech Detection Pilot
This dataset is a 6,000-clip binary audio-classification pilot for detecting
whether an utterance contains one language (label = 0) or more than one
language (label = 1). It covers Turkish (tur), Northern Kurdish/Kurmanji
(kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and
English (eng).
Dataset composition
Construction
Mixed
Monolingual
Total
Single-call OmniVoice
500
500
1,000
Segment-level… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-english-accented-vc.VCTKresampled_16KHrz_vctk_speakers_splitvctk-corpus-en-parquetvctkiemocap-vc-wavlm-large-layer-9-temporalvctk
Disclaimer
This dataset is not mine and I do not accept any legal responsibility for its use. This dataset is simply being reuploaded for easier accessibility.
VCapAV
Dataset Card for VCapAV
VCapAV is a large-scale audio-visual deepfake detection dataset focused on non-speech environmental sounds. It introduces new multimodal deepfake scenarios using both Text-to-Audio (TTA) and Video-to-Audio (V2A) pipelines, together with Text-to-Video (TTV) synthesis.The dataset contains 90,990 clips, totaling 252.75 hours, and supports audio-only, visual-only, and audio-visual detection tasks.
Dataset Description
VCapAV addresses the lack of… See the full description on the dataset page: https://huggingface.co/datasets/WailyWang/VCapAV.iemocap-vc-wavlm-layer-6-temporalNoiseDetection_VCTK_MUSAN-Music
Dataset Card for "NoiseDetectionmusic_VCTKMusan"
More Information needed
merged-vctk-cmuarctic-gbiv-corpus-19.0-2024-09-13-mnVoiceConversion_VCTK
Dataset Card for "VoiceConversion_VCTK"
More Information needed
NaturalVoices_VC_0.1 NaturalVoices VC 10%
A large voice conversion (VC) dataset curated from spontaneous, in-the-wild podcast speech as part of the NaturalVoices project in collaboration with 🤗MSP Lab at CMU LTI. This release provides the 10% subset uniformly sampled from 870-hour VC dataset and subsets mainly intended for training and evaluating emotion-aware voice conversion systems but not limited to VC tasks.
📄 Paper: NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice… See the full description on the dataset page: https://huggingface.co/datasets/JHU-SmileLab/NaturalVoices_VC_0.1.echo-embeddings-vctk-tar
VCTK Speaker Embeddings (tarred)
Items: 109
This dataset ships as a single tar at the repo root. Members preserve paths like
VCTK/<id>/audio.mp3 and VCTK/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the CSTR VCTK Corpus. Distributed under CC BY 4.0; attribution required.
