datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vctk
Dataset Card for "vctk"
More Information needed
VCR-wiki-en-easy
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-easy.VCR-wiki-en-hard
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-hard.VC-Tooler-SFT
VC-Tooler-SFT
Supervised cold-start trajectories for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use.
🔗 Links
📄 Paper: arXiv
🌐 Project Page: w1zheng.github.io/VC-Tooler
🤗 Hugging Face: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
🧩 ModelScope: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
This dataset is the Stage I (supervised fine-tuning) trajectory bank used to teach a
vision–language model to use visual tools as a compositional and adaptive… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-SFT.vctk
VCTK
This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning.
The original dataset is available at: https://datashare.ed.ac.uk/handle/10283/3443.
Reproducing
This repository notably lacks a requirements.txt file. There's likely a missing dependency or two, but roughly:
pydub
tqdm
torch
torchaudio… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/vctk.VC-Tooler-RL
VC-Tooler-RL
Reinforcement-learning data for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use.
🔗 Links
📄 Paper: arXiv
🌐 Project Page: w1zheng.github.io/VC-Tooler
🤗 Hugging Face: VC-Tooler-SFT · VC-Tooler-RL (this dataset)
🧩 ModelScope: VC-Tooler-SFT · VC-Tooler-RL (this dataset)
This dataset is the Stage II (agentic RL) data used to refine the cold-started VC-Tooler policy
through interaction with a tool environment. Unlike the SFT bank… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-RL.VCR-wiki-zh-easy
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-zh-easy.vcrVCR-wiki-zh-hard
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-zh-hard.vctk_dataset_no_unknownVCTK_DATASET_RESAMPLEDvc_5.3M-5.4Mvctk_resampled_16k_balancednoisy_vctk_16k_synth
Dataset Card for "noisy_vctk_16k_synth"
More Information needed
VCCvcr
VCR v1.0
VCR is the Visual Commonsense Reasoning dataset from "From Recognition to Cognition: Visual Commonsense Reasoning" (CVPR 2019).
This Hugging Face version has two loadable configs:
image_examples: the default viewer-friendly config, one row per unique image, with grouped annotations.
questions: one row per original VCR question/answer/rationale example.
The original annotation JSONL files are also included under original_annotations/ for legacy compatibility.… See the full description on the dataset page: https://huggingface.co/datasets/Rowan/vcr.vc_5.4M-5.8M_0.5svcc-perturb
Virtual Cell Challenge 2025 H1 hESC training-set perturbation atlas
CRISPRi (gene knockdown) in H1 human embryonic stem cells. Single-cell expression in log-normalized counts.
Generated 2026-06-08 as one of three companion atlases (Norman, Replogle, VCC).
File schema (each config / single-config repo)
File
Shape
Description
pseudobulks.h5ad
(50, n_genes)
50 control pseudobulks (15 cells each, log-normalized means). Cell-type-specific baseline.… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-lynn/vcc-perturb.vc_5.1M-5.2Mvctk-48khz
Dataset Card for VCTK (48kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.vctk
VCTK
This is a mirror of the VCTK Corpus.
The original files were converted from FLAC to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
train_mic1: 90 speakers, 33.6 hours, 35987 utterances
train_mic2: 90 speakers, 33.6 hours, 35987 utterances
val_mic1: 10 random speakers unseen during training: p238, p244, p254, p263, p265, p272, p288, p294, p305, and p335. 4.0 hours, 4179 utterances.
val_mic2: Same speakers as val_mic1.… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/vctk.VCReward-BenchVCReward-Bench includes 3,506 expert-annotated preference pairs for evaluating assessment models of image editing in Visual Consistency.
🚀 Quick Start!
Clone github repo
git clone https://github.com/ZhangqiJiang07/GEditBench_v2.git
cd GEditBench_v2
Use our autopipeline CIL for evaluation
# (optional, or you can invoke the CLIs directly with `python -m src.cli.<tool>`)
./scripts/install_autopipeline.sh
# you can use `python -m src.cli.autogen… See the full description on the dataset page: https://huggingface.co/datasets/GEditBench-v2/VCReward-Bench.vc_0.1M-0.3Mvctk-ttsvctk-16khz
Dataset Card for VCTK (16kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version at the original 48kHz sample rate is also available: saeedzou/vctk-48khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-16khz.VCDB-Core-AudioVideo
VCDB Core Audio-Video Retrieval
This repository packages synchronized video and extracted audio from the
528-video core set of VCDB as a symmetric video+audio-to-video+audio
retrieval task for MTEB/MOEB. The separate 100,000-video background collection
is not included.
Terms and provenance
The source dataset is provided by Fudan University for research purposes
only. The source authors and Fudan University make no warranties about the
dataset, including… See the full description on the dataset page: https://huggingface.co/datasets/pranitchawla/VCDB-Core-AudioVideo.Evaluation-Multilingual-VC
Evaluation-Multilingual-VC
We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon,
Filter languages that support by Whisper Large V3 to evaluate WER automatically,
Only take test set, sort by up votes.
Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows.
Only build first 500 rows for each language
Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.vctk_p225_allcolsvc_5M-5.4M_1svc_3.8M-4.2M_1s
