CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sanchit-gandhi /vctk Dataset Card for "vctk" More Information needed audio10K<n<100K2 likes2.7k downloads3y agoHugging Face02jspaulsen /vctk VCTK This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning. The original dataset is available at: https://datashare.ed.ac.uk/handle/10283/3443. Reproducing This repository notably lacks a requirements.txt file. There's likely a missing dependency or two, but roughly: pydub tqdm torch torchaudio… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/vctk.audiotext-to-speech10K<n<100K1 likes1k downloads1y agoHugging Face03badayvedat /VCTKaudio10K<n<100K6 likes1k downloads2y agoHugging Face04tencent /VCB-Bench VCB-Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents Introduction Voice Chat Bot Bench (VCB Bench) is a high-quality Chinese benchmark built entirely on real human speech. It evaluates large audio language models (LALMs) along three complementary dimensions: (1) Instruction following: Text Instruction Following (TIF), Speech Instruction Following (SIF), English Text Instruction Following (TIF-En)… See the full description on the dataset page: https://huggingface.co/datasets/tencent/VCB-Bench.audio9 likes808 downloads9mo agoHugging Face05confit /vctk-fullaudio0 likes383 downloads2y agoHugging Face06Milana /vctk_resampled_16k_balancedaudio10K<n<100K0 likes381 downloads2y agoHugging Face07Codec-SUPERB /noisy_vctk_16k_synth Dataset Card for "noisy_vctk_16k_synth" More Information needed audio100K<n<1M0 likes352 downloads3y agoHugging Face08saeedzou /vctk-48khzgated Dataset Card for VCTK (48kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.audioautomatic-speech-recognition10K<n<100K1 likes332 downloads2mo agoHugging Face09philgzl /vctk VCTK This is a mirror of the VCTK Corpus. The original files were converted from FLAC to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz Channels: 1 Format: Opus Splits: train_mic1: 90 speakers, 33.6 hours, 35987 utterances train_mic2: 90 speakers, 33.6 hours, 35987 utterances val_mic1: 10 random speakers unseen during training: p238, p244, p254, p263, p265, p272, p288, p294, p305, and p335. 4.0 hours, 4179 utterances. val_mic2: Same speakers as val_mic1.… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/vctk.audio10K<n<100K0 likes327 downloads4mo agoHugging Face10psusac /vctk-ttsaudio100K<n<1M0 likes264 downloads3mo agoHugging Face11saeedzou /vctk-16khz Dataset Card for VCTK (16kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.audioautomatic-speech-recognition10K<n<100K0 likes264 downloads2mo agoHugging Face12pranitchawla /VCDB-Core-AudioVideo VCDB Core Audio-Video Retrieval This repository packages synchronized video and extracted audio from the 528-video core set of VCDB as a symmetric video+audio-to-video+audio retrieval task for MTEB/MOEB. The separate 100,000-video background collection is not included. Terms and provenance The source dataset is provided by Fudan University for research purposes only. The source authors and Fudan University make no warranties about the dataset, including… See the full description on the dataset page: https://huggingface.co/datasets/pranitchawla/VCDB-Core-AudioVideo.audioother10K<n<100K0 likes263 downloads28d agoHugging Face13Scicom-intl /Evaluation-Multilingual-VC Evaluation-Multilingual-VC We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon, Filter languages that support by Whisper Large V3 to evaluate WER automatically, Only take test set, sort by up votes. Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows. Only build first 500 rows for each language Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.audio10K<n<100K0 likes252 downloads6mo agoHugging Face14tevien /vctk_p225_allcolsaudion<1K0 likes251 downloads2y agoHugging Face15kth-tmh /vctk Dataset Card for VCTK Dataset Summary This CSTR VCTK Corpus includes around 44-hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out about 400 sentences, which were selected from a newspaper, the rainbow passage and an elicitation paragraph used for the speech accent archive. Supported Tasks automatic-speech-recognition, speaker-identification: The dataset can be used to train a model for Automatic Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/kth-tmh/vctk.audioautomatic-speech-recognition10K<n<100K0 likes190 downloads7mo agoHugging Face16TartarusXXX /mixed-language-detection-english-accented-vc Mixed-Language Speech Detection Pilot This dataset is a 6,000-clip binary audio-classification pilot for detecting whether an utterance contains one language (label = 0) or more than one language (label = 1). It covers Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and English (eng). Dataset composition Construction Mixed Monolingual Total Single-call OmniVoice 500 500 1,000 Segment-level… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-english-accented-vc.audioaudio-classification1K<n<10K0 likes180 downloads1mo agoHugging Face17julianyi1 /VCTKaudio10K<n<100K0 likes153 downloads5mo agoHugging Face18Milana /resampled_16KHrz_vctk_speakers_splitaudio10K<n<100K0 likes148 downloads2y agoHugging Face19mimba /vctk-corpus-en-parquetgatedaudio10K<n<100K0 likes145 downloads4mo agoHugging Face20yfyeung /vctkaudion<1K0 likes135 downloads1y agoHugging Face21saeedzou /iemocap-vc-wavlm-large-layer-9-temporalaudio1K<n<10K0 likes130 downloads2mo agoHugging Face22Razer112 /vctk Disclaimer This dataset is not mine and I do not accept any legal responsibility for its use. This dataset is simply being reuploaded for easier accessibility. audion<1K0 likes126 downloads2mo agoHugging Face23WailyWang /VCapAV Dataset Card for VCapAV VCapAV is a large-scale audio-visual deepfake detection dataset focused on non-speech environmental sounds. It introduces new multimodal deepfake scenarios using both Text-to-Audio (TTA) and Video-to-Audio (V2A) pipelines, together with Text-to-Video (TTV) synthesis.The dataset contains 90,990 clips, totaling 252.75 hours, and supports audio-only, visual-only, and audio-visual detection tasks. Dataset Description VCapAV addresses the lack of… See the full description on the dataset page: https://huggingface.co/datasets/WailyWang/VCapAV.audiotext-to-audio100K<n<1M0 likes117 downloads10mo agoHugging Face24saeedzou /iemocap-vc-wavlm-layer-6-temporalaudio1K<n<10K0 likes100 downloads2mo agoHugging Face25DynamicSuperb /NoiseDetection_VCTK_MUSAN-Music Dataset Card for "NoiseDetectionmusic_VCTKMusan" More Information needed audion<1K1 likes71 downloads3y agoHugging Face26jimregan /merged-vctk-cmuarctic-gbiaudio10K<n<100K0 likes70 downloads7mo agoHugging Face27Cafet /v-corpus-19.0-2024-09-13-mnaudio10K<n<100K0 likes64 downloads2y agoHugging Face28DynamicSuperb /VoiceConversion_VCTK Dataset Card for "VoiceConversion_VCTK" More Information needed audio1K<n<10K0 likes62 downloads3y agoHugging Face29JHU-SmileLab /NaturalVoices_VC_0.1 NaturalVoices VC 10% A large voice conversion (VC) dataset curated from spontaneous, in-the-wild podcast speech as part of the NaturalVoices project in collaboration with 🤗MSP Lab at CMU LTI. This release provides the 10% subset uniformly sampled from 870-hour VC dataset and subsets mainly intended for training and evaluating emotion-aware voice conversion systems but not limited to VC tasks. 📄 Paper: NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice… See the full description on the dataset page: https://huggingface.co/datasets/JHU-SmileLab/NaturalVoices_VC_0.1.audioaudio-to-audio10K<n<100K2 likes61 downloads11mo agoHugging Face30jordand /echo-embeddings-vctk-tar VCTK Speaker Embeddings (tarred) Items: 109 This dataset ships as a single tar at the repo root. Members preserve paths like VCTK/<id>/audio.mp3 and VCTK/<id>/speaker_latent.safetensors. See loader.py for example loading. Attribution: Contains audio and embeddings derived from the CSTR VCTK Corpus. Distributed under CC BY 4.0; attribution required. textn<1K0 likes59 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.