CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Vchitect /Vchitect_T2V_DataVerse Vchitect-T2V-Dataverse Vchitect Team1  1Shanghai Artificial Intelligence Laboratory  Paper | Project Page | Data Overview The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models. It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.texttext-to-video1M<n<10M11 likes56k downloads1y agoHugging Face02OPPOer /IC-VCO-Dataset IC-VCO-Dataset &nbsp; &nbsp; This dataset package contains the two IC-VCO training subsets: sft: supervised fine-tuning examples. preference: visual contrastive preference examples. The two subsets intentionally use different schemas, so they are represented as separate Hugging Face dataset configurations instead of separate splits under a single configuration. Each configuration has a train split. Loading From Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/IC-VCO-Dataset.imagevisual-question-answering100K<n<1M0 likes4k downloads1mo agoHugging Face03sanchit-gandhi /vctk Dataset Card for "vctk" More Information needed audio10K<n<100K2 likes2.7k downloads3y agoHugging Face04vcr-org /VCR-wiki-en-easy The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-easy.imagevisual-question-answering1M<n<10M2 likes2.6k downloads2y agoHugging Face05vcr-org /VCR-wiki-en-hard The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-hard.imagevisual-question-answering1M<n<10M2 likes2k downloads2y agoHugging Face065551z /VC-Tooler-SFT VC-Tooler-SFT Supervised cold-start trajectories for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use. 🔗 Links 📄 Paper: arXiv 🌐 Project Page: w1zheng.github.io/VC-Tooler 🤗 Hugging Face: VC-Tooler-SFT (this dataset) · VC-Tooler-RL 🧩 ModelScope: VC-Tooler-SFT (this dataset) · VC-Tooler-RL This dataset is the Stage I (supervised fine-tuning) trajectory bank used to teach a vision–language model to use visual tools as a compositional and adaptive… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-SFT.imagevisual-question-answering10K<n<100K3 likes1.3k downloads2mo agoHugging Face07jspaulsen /vctk VCTK This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning. The original dataset is available at: https://datashare.ed.ac.uk/handle/10283/3443. Reproducing This repository notably lacks a requirements.txt file. There's likely a missing dependency or two, but roughly: pydub tqdm torch torchaudio… See the full description on the dataset page: https://huggingface.co/datasets/jspaulsen/vctk.audiotext-to-speech10K<n<100K1 likes1k downloads1y agoHugging Face085551z /VC-Tooler-RL VC-Tooler-RL Reinforcement-learning data for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use. 🔗 Links 📄 Paper: arXiv 🌐 Project Page: w1zheng.github.io/VC-Tooler 🤗 Hugging Face: VC-Tooler-SFT · VC-Tooler-RL (this dataset) 🧩 ModelScope: VC-Tooler-SFT · VC-Tooler-RL (this dataset) This dataset is the Stage II (agentic RL) data used to refine the cold-started VC-Tooler policy through interaction with a tool environment. Unlike the SFT bank… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-RL.textvisual-question-answering10K<n<100K2 likes1k downloads2mo agoHugging Face09badayvedat /VCTKaudio10K<n<100K6 likes1k downloads2y agoHugging Face10NilanE /Vchitect_T2V_DataVerse_256p_8fps_wdshttps://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse resampled to 256p. Intended for training https://github.com/NilanEkanayake/TiTok-Video text100K<n<1M0 likes825 downloads1y agoHugging Face11pritamqu /VCRBench VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Authors: Pritam Sarkar and Ali Etemad This repository provides the official implementation of VCRBench. Usage Please check our GitHub repo for the details of usage: VCRBench from dataset import VCRBench dataset=VCRBench(question_file="data.json", video_root="./", mode='default', ) for sample in dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pritamqu/VCRBench.textvideo-text-to-textn<1K1 likes718 downloads1y agoHugging Face12vcr-org /VCR-wiki-zh-easy The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-zh-easy.imagevisual-question-answering100K<n<1M1 likes692 downloads2y agoHugging Face13sionic-ai /vcrimage1M<n<10M0 likes636 downloads1y agoHugging Face14vcr-org /VCR-wiki-zh-hard The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-zh-hard.imagevisual-question-answering100K<n<1M1 likes611 downloads2y agoHugging Face15Milana /vctk_dataset_no_unknowntext10K<n<100K0 likes541 downloads2y agoHugging Face16Milana /VCTK_DATASET_RESAMPLEDtext10K<n<100K0 likes508 downloads2y agoHugging Face17nicolas-lynn /vcell-perturbation-source-data ConvergeCELL Source Datasets — v1.0.0 Cached H5ADs of every single-cell RNA-seq dataset registered in the ConvergeCELL data catalog. Mirrors the original sources (GEO, figshare, Tabula Sapiens, CellxGene) so downstream code has a single, fast, versioned endpoint to fetch from. Each row of every h5ad is one cell; each column is one gene (HGNC symbol). The exact obs/var schema follows whatever the original source provided — this bundle does not re-annotate, harmonize, or QC. For a… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-lynn/vcell-perturbation-source-data.textn<1K0 likes480 downloads4mo agoHugging Face18AdoCleanCode /vc_5.3M-5.4Mtext10K<n<100K0 likes474 downloads8mo agoHugging Face19confit /vctk-fullaudio0 likes383 downloads2y agoHugging Face20Milana /vctk_resampled_16k_balancedaudio10K<n<100K0 likes381 downloads2y agoHugging Face21Codec-SUPERB /noisy_vctk_16k_synth Dataset Card for "noisy_vctk_16k_synth" More Information needed audio100K<n<1M0 likes352 downloads3y agoHugging Face22TitouanCh /VCCtext100K<n<1M0 likes347 downloads1y agoHugging Face23Rowan /vcr VCR v1.0 VCR is the Visual Commonsense Reasoning dataset from "From Recognition to Cognition: Visual Commonsense Reasoning" (CVPR 2019). This Hugging Face version has two loadable configs: image_examples: the default viewer-friendly config, one row per unique image, with grouped annotations. questions: one row per original VCR question/answer/rationale example. The original annotation JSONL files are also included under original_annotations/ for legacy compatibility.… See the full description on the dataset page: https://huggingface.co/datasets/Rowan/vcr.imagevisual-question-answering100K<n<1M0 likes341 downloads3mo agoHugging Face24AdoCleanCode /vc_5.4M-5.8M_0.5stext10K<n<100K0 likes338 downloads8mo agoHugging Face25nicolas-lynn /vcc-perturb Virtual Cell Challenge 2025 H1 hESC training-set perturbation atlas CRISPRi (gene knockdown) in H1 human embryonic stem cells. Single-cell expression in log-normalized counts. Generated 2026-06-08 as one of three companion atlases (Norman, Replogle, VCC). File schema (each config / single-config repo) File Shape Description pseudobulks.h5ad (50, n_genes) 50 control pseudobulks (15 cells each, log-normalized means). Cell-type-specific baseline.… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-lynn/vcc-perturb.tabulartabular-regressionn<1K0 likes335 downloads3mo agoHugging Face26AdoCleanCode /vc_5.1M-5.2Mtext10K<n<100K0 likes332 downloads8mo agoHugging Face27saeedzou /vctk-48khzgated Dataset Card for VCTK (48kHz) This is a re-packaged, HuggingFace-native version of the CSTR VCTK Corpus, provided as a ready-to-use datasets object (audio decoded via the Audio feature) rather than a loading-script-based dataset. A companion version resampled to 16kHz is also available: saeedzou/vctk-16khz. Dataset Summary This CSTR VCTK Corpus includes around 44 hours of speech data uttered by 110 English speakers with various accents. Each speaker reads out… See the full description on the dataset page: https://huggingface.co/datasets/saeedzou/vctk-48khz.audioautomatic-speech-recognition10K<n<100K1 likes332 downloads2mo agoHugging Face28philgzl /vctk VCTK This is a mirror of the VCTK Corpus. The original files were converted from FLAC to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz Channels: 1 Format: Opus Splits: train_mic1: 90 speakers, 33.6 hours, 35987 utterances train_mic2: 90 speakers, 33.6 hours, 35987 utterances val_mic1: 10 random speakers unseen during training: p238, p244, p254, p263, p265, p272, p288, p294, p305, and p335. 4.0 hours, 4179 utterances. val_mic2: Same speakers as val_mic1.… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/vctk.audio10K<n<100K0 likes327 downloads4mo agoHugging Face29GEditBench-v2 /VCReward-BenchVCReward-Bench includes 3,506 expert-annotated preference pairs for evaluating assessment models of image editing in Visual Consistency. 🚀 Quick Start! Clone github repo git clone https://github.com/ZhangqiJiang07/GEditBench_v2.git cd GEditBench_v2 Use our autopipeline CIL for evaluation # (optional, or you can invoke the CLIs directly with `python -m src.cli.<tool>`) ./scripts/install_autopipeline.sh # you can use `python -m src.cli.autogen… See the full description on the dataset page: https://huggingface.co/datasets/GEditBench-v2/VCReward-Bench.image1K<n<10K4 likes282 downloads6mo agoHugging Face30AdoCleanCode /vc_0.1M-0.3Mtext100K<n<1M0 likes276 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.