datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vitra-ego4d-videoLlama-VITS_data
Dataset Card for Llama-VITS_data
The dataset repository contains data related with our work "Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness", encapsulating:
Filtered dataset EmoV_DB_bea_sem
Filelists with semantic embeddings
Model checkpoints
Human evaluation templates
Dataset Details
Paper: Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
Curated by: Xincan Feng, Akifumi Yoshimoto
Funded by: CyberAgent Inc
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/xincan/Llama-VITS_data.ViT-FineTunevitaminc
Details
Fact Verification dataset created for Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence (Schuster et al., NAACL 21`) based on Wikipedia edits (revisions).
For more details see: https://github.com/TalSchuster/VitaminC
When using this dataset, please cite the paper:
BibTeX entry and citation info
@inproceedings{schuster-etal-2021-get,
title = "Get Your Vitamin {C}! Robust Fact Verification with Contrastive Evidence",
author =… See the full description on the dataset page: https://huggingface.co/datasets/tals/vitaminc.VITRA-1M
VITRA-1M: Human Hand V-L-A Dataset
Dataset Summary
VITRA-1M is a large-scale Human Hand Visual-Language-Action (V-L-A) dataset constructed as described in the paper Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos. It contains 1.2 million short episodes with segmented language annotations, camera parameters (corrected intrinsics/extrinsics), and 3D hand reconstructions (left and right… See the full description on the dataset page: https://huggingface.co/datasets/VITRA-VLA/VITRA-1M.ViTTiny1022
ViTTiny1022
The dataset for Scaling Up Parameter Generation: A Recurrent Diffusion Approach.
Requirement
Install torch and other dependencies
conda install pytorch==2.3.1 torchvision==0.18.1 torchaudio==2.3.1 pytorch-cuda=12.1 -c pytorch -c nvidia
pip install timm einops seaborn openpyxl
Usage
Test one checkpoint
cd ViTTiny1022
python test.py ./chechpoint_test/0000_acc0.9613_class0314_condition_cifar10_vittiny.pth
# python test.py… See the full description on the dataset page: https://huggingface.co/datasets/MTDoven/ViTTiny1022.vitra-dinotxt-featuresfeatures-dinov3-vith16plus-224-imagenet-22k-wdsbehavior-1k-2025-challenge-vjepa2-vitg-demo-embeddings
V-JEPA 2 ViT-G Embeddings — BEHAVIOR-1K 2025 Challenge Demos (62h)
Precomputed video embeddings for a 62-hour subsample of the
BEHAVIOR-1K 2025 challenge demonstrations,
extracted with the V-JEPA 2 ViT-g encoder.
The goal is to make downstream experimentation faster and more reproducible by eliminating
repeated video decoding and encoder forward passes — lowering the barrier for teams
without access to large GPU clusters.
Field
Value
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/quastAI/behavior-1k-2025-challenge-vjepa2-vitg-demo-embeddings.CLIP-ViT-H-14-laion2B-s32B-b79K-all-checkpointsThis repository contains the intermediate checkpoints for the model https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K.
Each "epoch" corresponds to an additional (32B / 256) samples seen, consituting total of 256 "epochs"
The purpose of releasing these checkpoints and optimizer states is to enable analysis.
For the first 121 "epochs", training was done with float16 mixed precision before switching to bfloat16 after a loss blow up.
vitac500stanford_mask_vit_rawVITON-HD-TESTvitco
ViTco
165,847,195 Vietnamese documents from 4 public corpora, 370.2 GB of Parquet, one schema
This dataset is the pinned public Vietnamese corpora as gao read them, every source put to one contract and one schema, before any cleaning.
Contents
What is it
What is in it
Where the text came from
How it is laid out
Reading it
What you can build with it
One row
The columns
What this repo is
What ships and what does not
Things to know before you use it
What this is… See the full description on the dataset page: https://huggingface.co/datasets/open-index/vitco.svi-benchmark
Stable Video Infinity (SVI) Benchmark Dataset
This benchmark dataset is introduced in the paper:
Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
by Wuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao, Alexandre Alahi (2025).
Project page: https://stable-video-infinity.github.io/homepage/
Code: https://github.com/vita-epfl/Stable-Video-Infinity
Abstract
We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vita/svi-benchmark.VITRA-TeleData
VITRA Teleoperation Dataset
Dataset Summary
This dataset contains real-world robot teleoperation demonstrations collected
using a 7-DoF robotic arm equipped with a dexterous hand and a head-mounted RGB
camera. Each episode provides synchronized numerical state/action data
and video recordings. The dataset is used for finetuning in the project VITRA: Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
Project… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/VITRA-TeleData.so-vits-svc-4.0-ru-The_Witcher_3_Wild_HuntЭто тренировочные данные моделей голосов персонажей из "Ведьмак 3: Дикая охота" для so-vits-svc-4.1.1
ViTextRender-500K
Vietnamese Text Render 500K Dataset
A large-scale dataset containing 500K Vietnamese text rendering image-text pairs for training generative models to improve text rendering performance.
Dataset Structure
image: Rendered text image in PNG format
text: Corresponding text content
filename: Original filename
Usage
This dataset is designed for fine-tuning generative models to improve text rendering capabilities on Vietnamese language.
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/pixxu/ViTextRender-500K.TrCaption-trclip-vitl14-e10VITON-HD-edit
VITON-HD-edit
This repository contains the VITON-HD-edit dataset presented in the paper CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation.
Github Repository | Paper (arXiv)
Dataset Overview
The VITON-HD-edit dataset is a public benchmark built to support three evaluations:
Image-editing Virtual Try-On (VTO)
Instance-level visual-prompt segmentation
Spatial controllability
Training the editing model requires triplets $(p… See the full description on the dataset page: https://huggingface.co/datasets/NXN-Labs/VITON-HD-edit.food101-vit-processedenergy-consumption-hourly-spainVitaSet
VitaSet: Vision-Tactile VQA Dataset
Overview
VitaSet is a vision-tactile Visual Question Answering dataset for physical property reasoning. The dataset combines RGB vision and tactile sensing for material property understanding, containing 5,145 human-verified QA pairs across three tasks: hardness classification, material property description, and surface roughness classification.
Hardware: Franka Emika Panda robot + GelSight Mini tactile sensor
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Bupt-Joy/VitaSet.sam2-vit-bVITATECS
Dataset Card for VITATECS
Dataset Description
Dataset Summary
VITATECS is a diagnostic VIdeo-Text dAtaset for the evaluation of TEmporal Concept underStanding.
[2023/11/27] We have updated a new version of VITATECS which is generated using ChatGPT. The previous version generated by OPT-175B can be found here.
Languages
English.
Dataset Structure
Usage
aspect = 'Type' #… See the full description on the dataset page: https://huggingface.co/datasets/lscpku/VITATECS.vite-selfbench
Vite Selfbench
Vite Selfbench is a 27-task software-engineering benchmark for coding agents, packaged for the Harbor evaluation framework. Each task asks an agent to implement a change in a frozen revision of vitejs/vite, then checks the resulting patch with task-specific tests. Harbor provides isolated task environments and runs verification separately from the agent.
This repository contains the raw evaluation only. It does not include model outputs, scores, costs, or… See the full description on the dataset page: https://huggingface.co/datasets/dari-ai/vite-selfbench.ViTTAVideo Test-Time Adaptation for Action Recognition (CVPR 2023)
Project Page
GitHub Repo
Arxiv Paper
Dataset Description
This dataset repo contains the following two datasets:
Kinetics400_val_corruptions: 12 corruption types for the 19877 validation videos on Kinetics400.
SSv2_val_corruptions: 12 corruption types for the 24777 validation videos on Something-Something v2.
VitaBench
🌱VitaBench: Benchmarking LLM Agents
with Versatile Interactive Tasks
📃 Paper • 🌐 Website • 🏆 Leaderboard • 🛠️ Code • 🤗 Dataset
🔔 News
[2026-01] Qwen3-Max-Thinking reported our Vita-Bench to evaluate and demonstrate its tool use capabilities (the averge score of 4 domains)!We invite the community to adopt Vita-Bench as the definitive touchstone for tool use performance assessment, and we appreciate diverse utilization & interpretation of our benchmark… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/VitaBench.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13wolt-food-clip-ViT-B-32-embeddings
wolt-food-clip-ViT-B-32-embeddings
Qdrant's Food Discovery demo relies on the dataset of food images from the Wolt
app. Each point in the collection represents a dish with a single image. The image is represented as a vector of 512
float numbers.
Generation process
The embeddings generated with clip-ViT-B-32 model have been generated using the following code snippet:
from PIL import Image
from sentence_transformers import SentenceTransformer
image_path =… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/wolt-food-clip-ViT-B-32-embeddings.
