CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sanchit-gandhi /vctk Dataset Card for "vctk" More Information needed audio10K<n<100K2 likes2.7k downloads3y agoHugging Face02vcr-org /VCR-wiki-en-easy The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-easy.imagevisual-question-answering1M<n<10M2 likes2.6k downloads2y agoHugging Face03vcr-org /VCR-wiki-en-hard The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-hard.imagevisual-question-answering1M<n<10M2 likes2k downloads2y agoHugging Face045551z /VC-Tooler-SFT VC-Tooler-SFT Supervised cold-start trajectories for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use. 🔗 Links 📄 Paper: arXiv 🌐 Project Page: w1zheng.github.io/VC-Tooler 🤗 Hugging Face: VC-Tooler-SFT (this dataset) · VC-Tooler-RL 🧩 ModelScope: VC-Tooler-SFT (this dataset) · VC-Tooler-RL This dataset is the Stage I (supervised fine-tuning) trajectory bank used to teach a vision–language model to use visual tools as a compositional and adaptive… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-SFT.imagevisual-question-answering10K<n<100K3 likes1.3k downloads2mo agoHugging Face05jspaulsen /vctk VCTK This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning. The original dataset is available at: https://datashare.ed.ac.uk/handle/10283/3443. Reproducing This repository notably lacks a requirements.txt file. There's likely a missing dependency or two, but roughly: pydub tqdm torch torchaudio… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/vctk.audiotext-to-speech10K<n<100K1 likes1k downloads1y agoHugging Face065551z /VC-Tooler-RL VC-Tooler-RL Reinforcement-learning data for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use. 🔗 Links 📄 Paper: arXiv 🌐 Project Page: w1zheng.github.io/VC-Tooler 🤗 Hugging Face: VC-Tooler-SFT · VC-Tooler-RL (this dataset) 🧩 ModelScope: VC-Tooler-SFT · VC-Tooler-RL (this dataset) This dataset is the Stage II (agentic RL) data used to refine the cold-started VC-Tooler policy through interaction with a tool environment. Unlike the SFT bank… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-RL.textvisual-question-answering10K<n<100K2 likes1k downloads2mo agoHugging Face07vcr-org /VCR-wiki-zh-easy The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-zh-easy.imagevisual-question-answering100K<n<1M1 likes692 downloads2y agoHugging Face08sionic-ai /vcrimage1M<n<10M0 likes636 downloads1y agoHugging Face09vcr-org /VCR-wiki-zh-hard The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-zh-hard.imagevisual-question-answering100K<n<1M1 likes611 downloads2y agoHugging Face10Milana /vctk_dataset_no_unknowntext10K<n<100K0 likes541 downloads2y agoHugging Face11Milana /VCTK_DATASET_RESAMPLEDtext10K<n<100K0 likes508 downloads2y agoHugging Face12AdoCleanCode /vc_5.3M-5.4Mtext10K<n<100K0 likes474 downloads8mo agoHugging Face13Milana /vctk_resampled_16k_balancedaudio10K<n<100K0 likes381 downloads2y agoHugging Face14Codec-SUPERB /noisy_vctk_16k_synth Dataset Card for "noisy_vctk_16k_synth" More Information needed audio100K<n<1M0 likes352 downloads3y agoHugging Face15TitouanCh /VCCtext100K<n<1M0 likes347 downloads1y agoHugging Face16Rowan /vcr VCR v1.0 VCR is the Visual Commonsense Reasoning dataset from "From Recognition to Cognition: Visual Commonsense Reasoning" (CVPR 2019). This Hugging Face version has two loadable configs: image_examples: the default viewer-friendly config, one row per unique image, with grouped annotations. questions: one row per original VCR question/answer/rationale example. The original annotation JSONL files are also included under original_annotations/ for legacy compatibility.… See the full description on the dataset page: https://huggingface.co/datasets/Rowan/vcr.imagevisual-question-answering100K<n<1M0 likes341 downloads3mo agoHugging Face17AdoCleanCode /vc_5.4M-5.8M_0.5stext10K<n<100K0 likes338 downloads8mo agoHugging Face18nicolas-lynn /vcc-perturb Virtual Cell Challenge 2025 H1 hESC training-set perturbation atlas CRISPRi (gene knockdown) in H1 human embryonic stem cells. Single-cell expression in log-normalized counts. Generated 2026-06-08 as one of three companion atlases (Norman, Replogle, VCC). File schema (each config / single-config repo) File Shape Description pseudobulks.h5ad (50, n_genes) 50 control pseudobulks (15 cells each, log-normalized means). Cell-type-specific baseline.… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-lynn/vcc-perturb.tabulartabular-regressionn<1K0 likes335 downloads3mo agoHugging Face19AdoCleanCode /vc_5.1M-5.2Mtext10K<n<100K0 likes332 downloads8mo agoHugging Face20saeedzou /vctk-48khzgated Dataset Card for VCTK (48kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.audioautomatic-speech-recognition10K<n<100K1 likes332 downloads2mo agoHugging Face21philgzl /vctk VCTK This is a mirror of the VCTK Corpus. The original files were converted from FLAC to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz Channels: 1 Format: Opus Splits: train_mic1: 90 speakers, 33.6 hours, 35987 utterances train_mic2: 90 speakers, 33.6 hours, 35987 utterances val_mic1: 10 random speakers unseen during training: p238, p244, p254, p263, p265, p272, p288, p294, p305, and p335. 4.0 hours, 4179 utterances. val_mic2: Same speakers as val_mic1.… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/vctk.audio10K<n<100K0 likes327 downloads5mo agoHugging Face22GEditBench-v2 /VCReward-BenchVCReward-Bench includes 3,506 expert-annotated preference pairs for evaluating assessment models of image editing in Visual Consistency. 🚀 Quick Start! Clone github repo git clone https://github.com/ZhangqiJiang07/GEditBench_v2.git cd GEditBench_v2 Use our autopipeline CIL for evaluation # (optional, or you can invoke the CLIs directly with `python -m src.cli.<tool>`) ./scripts/install_autopipeline.sh # you can use `python -m src.cli.autogen… See the full description on the dataset page: https://huggingface.co/datasets/GEditBench-v2/VCReward-Bench.image1K<n<10K4 likes282 downloads6mo agoHugging Face23AdoCleanCode /vc_0.1M-0.3Mtext100K<n<1M0 likes276 downloads8mo agoHugging Face24psusac /vctk-ttsaudio100K<n<1M0 likes264 downloads3mo agoHugging Face25saeedzou /vctk-16khz Dataset Card for VCTK (16kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.audioautomatic-speech-recognition10K<n<100K0 likes264 downloads2mo agoHugging Face26pranitchawla /VCDB-Core-AudioVideo VCDB Core Audio-Video Retrieval This repository packages synchronized video and extracted audio from the 528-video core set of VCDB as a symmetric video+audio-to-video+audio retrieval task for MTEB/MOEB. The separate 100,000-video background collection is not included. Terms and provenance The source dataset is provided by Fudan University for research purposes only. The source authors and Fudan University make no warranties about the dataset, including… See the full description on the dataset page: https://huggingface.co/datasets/pranitchawla/VCDB-Core-AudioVideo.audioother10K<n<100K0 likes263 downloads29d agoHugging Face27Scicom-intl /Evaluation-Multilingual-VC Evaluation-Multilingual-VC We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon, Filter languages that support by Whisper Large V3 to evaluate WER automatically, Only take test set, sort by up votes. Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows. Only build first 500 rows for each language Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.audio10K<n<100K0 likes252 downloads6mo agoHugging Face28tevien /vctk_p225_allcolsaudion<1K0 likes251 downloads2y agoHugging Face29AdoCleanCode /vc_5M-5.4M_1stext100K<n<1M0 likes244 downloads8mo agoHugging Face30AdoCleanCode /vc_3.8M-4.2M_1stext100K<n<1M0 likes235 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.