datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse
Vchitect Team1
1Shanghai Artificial Intelligence Laboratory
Paper |
Project Page |
Data Overview
The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.
It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.IC-VCO-Dataset
IC-VCO-Dataset
This dataset package contains the two IC-VCO training subsets:
sft: supervised fine-tuning examples.
preference: visual contrastive preference examples.
The two subsets intentionally use different schemas, so they are represented as separate Hugging Face dataset configurations instead of separate splits under a single configuration. Each configuration has a train split.
Loading From Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/IC-VCO-Dataset.vctk
Dataset Card for "vctk"
More Information needed
VCR-wiki-en-easy
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-easy.VCR-wiki-en-hard
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-hard.VC-Tooler-SFT
VC-Tooler-SFT
Supervised cold-start trajectories for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use.
🔗 Links
📄 Paper: arXiv
🌐 Project Page: w1zheng.github.io/VC-Tooler
🤗 Hugging Face: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
🧩 ModelScope: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
This dataset is the Stage I (supervised fine-tuning) trajectory bank used to teach a
vision–language model to use visual tools as a compositional and adaptive… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-SFT.vctk
VCTK
This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning.
The original dataset is available at: https://datashare.ed.ac.uk/handle/10283/3443.
Reproducing
This repository notably lacks a requirements.txt file. There's likely a missing dependency or two, but roughly:
pydub
tqdm
torch
torchaudio… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/vctk.VC-Tooler-RL
VC-Tooler-RL
Reinforcement-learning data for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use.
🔗 Links
📄 Paper: arXiv
🌐 Project Page: w1zheng.github.io/VC-Tooler
🤗 Hugging Face: VC-Tooler-SFT · VC-Tooler-RL (this dataset)
🧩 ModelScope: VC-Tooler-SFT · VC-Tooler-RL (this dataset)
This dataset is the Stage II (agentic RL) data used to refine the cold-started VC-Tooler policy
through interaction with a tool environment. Unlike the SFT bank… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-RL.VCTKVchitect_T2V_DataVerse_256p_8fps_wdshttps://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse resampled to 256p. Intended for training https://github.com/NilanEkanayake/TiTok-Video
VCRBench
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models
Authors: Pritam Sarkar and Ali Etemad
This repository provides the official implementation of VCRBench.
Usage
Please check our GitHub repo for the details of usage: VCRBench
from dataset import VCRBench
dataset=VCRBench(question_file="data.json",
video_root="./",
mode='default',
)
for sample in dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pritamqu/VCRBench.VCR-wiki-zh-easy
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-zh-easy.vcrVCR-wiki-zh-hard
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-zh-hard.vctk_dataset_no_unknownVCTK_DATASET_RESAMPLEDvcell-perturbation-source-data
ConvergeCELL Source Datasets — v1.0.0
Cached H5ADs of every single-cell RNA-seq dataset registered in the
ConvergeCELL data catalog. Mirrors the original sources (GEO, figshare,
Tabula Sapiens, CellxGene) so downstream code has a single, fast, versioned
endpoint to fetch from.
Each row of every h5ad is one cell; each column is one gene (HGNC symbol).
The exact obs/var schema follows whatever the original source provided —
this bundle does not re-annotate, harmonize, or QC. For a… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-lynn/vcell-perturbation-source-data.vc_5.3M-5.4Mvctk-fullvctk_resampled_16k_balancednoisy_vctk_16k_synth
Dataset Card for "noisy_vctk_16k_synth"
More Information needed
VCCvcr
VCR v1.0
VCR is the Visual Commonsense Reasoning dataset from "From Recognition to Cognition: Visual Commonsense Reasoning" (CVPR 2019).
This Hugging Face version has two loadable configs:
image_examples: the default viewer-friendly config, one row per unique image, with grouped annotations.
questions: one row per original VCR question/answer/rationale example.
The original annotation JSONL files are also included under original_annotations/ for legacy compatibility.… See the full description on the dataset page: https://huggingface.co/datasets/Rowan/vcr.vc_5.4M-5.8M_0.5svcc-perturb
Virtual Cell Challenge 2025 H1 hESC training-set perturbation atlas
CRISPRi (gene knockdown) in H1 human embryonic stem cells. Single-cell expression in log-normalized counts.
Generated 2026-06-08 as one of three companion atlases (Norman, Replogle, VCC).
File schema (each config / single-config repo)
File
Shape
Description
pseudobulks.h5ad
(50, n_genes)
50 control pseudobulks (15 cells each, log-normalized means). Cell-type-specific baseline.… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-lynn/vcc-perturb.vc_5.1M-5.2Mvctk-48khz
Dataset Card for VCTK (48kHz)
This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset.
A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz.
Dataset Summary
This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.vctk
VCTK
This is a mirror of the VCTK Corpus.
The original files were converted from FLAC to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
train_mic1: 90 speakers, 33.6 hours, 35987 utterances
train_mic2: 90 speakers, 33.6 hours, 35987 utterances
val_mic1: 10 random speakers unseen during training: p238, p244, p254, p263, p265, p272, p288, p294, p305, and p335. 4.0 hours, 4179 utterances.
val_mic2: Same speakers as val_mic1.… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/vctk.VCReward-BenchVCReward-Bench includes 3,506 expert-annotated preference pairs for evaluating assessment models of image editing in Visual Consistency.
🚀 Quick Start!
Clone github repo
git clone https://github.com/ZhangqiJiang07/GEditBench_v2.git
cd GEditBench_v2
Use our autopipeline CIL for evaluation
# (optional, or you can invoke the CLIs directly with `python -m src.cli.<tool>`)
./scripts/install_autopipeline.sh
# you can use `python -m src.cli.autogen… See the full description on the dataset page: https://huggingface.co/datasets/GEditBench-v2/VCReward-Bench.vc_0.1M-0.3M
