datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vitaminc
Details
Fact Verification dataset created for Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence (Schuster et al., NAACL 21`) based on Wikipedia edits (revisions).
For more details see: https://github.com/TalSchuster/VitaminC
When using this dataset, please cite the paper:
BibTeX entry and citation info
@inproceedings{schuster-etal-2021-get,
title = "Get Your Vitamin {C}! Robust Fact Verification with Contrastive Evidence",
author =… See the full description on the dataset page: https://huggingface.co/datasets/tals/vitaminc.sam2-vit-bViTeX-Dataset
ViTeX-Dataset
🌐 Project page ·
📊 Dataset ·
🧪 Benchmark code ·
🤖 Model & Inference code ·
🏆 Leaderboard
Paired real-video dataset for video scene text editing: given a source video, a binary text-region mask, and a (source string → target string) pair, replace only the masked scene text across all frames while preserving the rest of the scene.
Accepted to NeurIPS 2026 E&D Track.
Authors: Xinghao Chen, Xiangbo Gao, Jiongze Yu… See the full description on the dataset page: https://huggingface.co/datasets/ViTeX-Bench/ViTeX-Dataset.sam2-vit-sVitaBench-2.0Long-VITA-DataBioMed-VITAL-instructions
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
Authors: Hejie Cui*, Lingjun Mao*, Xin Liang, Jieyu Zhang, Hui Ren, Quanzheng Li, Xiang Li, Carl Yang
BioMed-VITAL Instructions Files
Data file name
File Size
Sample Size
BioMed-VITAL-instructions-60K.json
127 MB
60K
BioMed-VITAL-instructions-80K.json
156 MB
80K
BioMed-VITAL-instructions-150K.json
309 MB
60K + 10K + 80K
BioMed-VITAL-instructions-210K.json
463 MB
80K + 10K + 60K +… See the full description on the dataset page: https://huggingface.co/datasets/mao1207/BioMed-VITAL-instructions.vitamins-supplements-reviews
Dataset Card for turkish-nlp-suite/vitamins-supplements-reviews
Dataset Summary
Turkish sentiment analysis dataset from customer reviews about supplement and vitamin products. The dataset is scraped from Vitaminler.com and contains
customer reviews and star rating about vitamin and supplement products.
Each customer review in the Vitamins and Supplements Reviews Dataset describes a customer’s experience with a supplement product in terms of the product’s effectiveness… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/vitamins-supplements-reviews.gigahands-vitra-mano
GigaHands → VITRA Stage-1, official-MANO annotations
VITRA Stage-1 hand annotations for GigaHands, with all joint positions taken from GigaHands'
official MANO fit instead of mixing in triangulated keypoints. Annotations only — no videos
(get those from GigaHands; the mapping is described in §5).
episodes
13,247 (train 11,904 / test 1,343)
frames
3,395,733
camera
brics-odroid-001_cam0 (static rig; one constant extrinsic per scene)
source
GigaHands params/ +… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/gigahands-vitra-mano.so-vits-svcr3al-vit-quantization-codex-trace
R3AL ViT Quantization — Codex Agent Trace
Codex session trace for installing the R3AL CLI and agent skill, exporting
google/vit-base-patch16-224 to ONNX, performing dynamic INT8 post-training
quantization on R3AL, and evaluating model size, Apple-arm64 CPU latency, and
prediction fidelity on a 100-image ImageNet validation sample.
The original Codex JSONL format is preserved for Hugging Face's native Agent
Trace viewer. Credential values, email addresses, unrelated Gmail/Slack… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/r3al-vit-quantization-codex-trace.Fight_Vitiligo_FAQ_Nepali_Health_QA_Dataset
Fight Vitiligo FAQ — Nepali Health Q&A Dataset
1. Overview
This dataset is a pure Nepali (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about vitiligo (भिटिलिगो) — a real, non-synthetic health-education FAQ set. Unlike template-generated MCQ datasets, this one contains naturally written, open-ended, explanatory Q&A pairs authored around a single health topic: vitiligo and related skin health.
Each… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Fight_Vitiligo_FAQ_Nepali_Health_QA_Dataset.azm-backup-20260909-depth-anything-v2-vitl
depth-anything-v2-vitl
Backup preserving the original files.
1 source files; 1,341,395,338 bytes.
Download all files to one directory, then verify with sha256sum -c SHA256SUMS.
azm-backup-20260909-depth-anything-v2-vitg
depth-anything-v2-vitg
Backup preserving the original files.
1 source files; 5,032,735,482 bytes.
Download all files to one directory, then verify with sha256sum -c SHA256SUMS.
VitsPerTrainModelazm-backup-20260909-depth-anything-v2-vits
depth-anything-v2-vits
Backup preserving the original files.
1 source files; 99,218,434 bytes.
Download all files to one directory, then verify with sha256sum -c SHA256SUMS.
movies_CLIP_ViT-L14
🎬 Movie Frame & Caption Dataset
📖 Introduction
This dataset was created from multiple movies across 10 genres, with approximately 3 movies per genre.From each movie, frames were extracted periodically, and AI-generated captions (BLIP) were assigned to each frame.A total of 93,813 frames were extracted.
This dataset can be used for tasks such as:
Video understanding
Multimodal learning (image + text)
Image captioning
Vision-language retrieval
📂 Data… See the full description on the dataset page: https://huggingface.co/datasets/thaotien/movies_CLIP_ViT-L14.Pedagogy-R1-benchmark-ViTutorHome-Assistant-Requests-V2
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/Vitinf/Home-Assistant-Requests-V2.vi-th-zh-parallel-corpus
VI/TH→ZH Parallel Corpus
Dataset Summary
A sentence-aligned parallel corpus for Vietnamese→Chinese (vi→zh) and Thai→Chinese (th→zh) machine translation. The dataset is released as two configs (vi-zh, th-zh) under a single repository, sharing the same schema, cleaning pipeline, and license.
Source languages: Vietnamese (vi), Thai (th)
Target language: Simplified Chinese (zh)
Primary task: Machine translation (vi→zh, th→zh); also usable for multilingual pretraining… See the full description on the dataset page: https://huggingface.co/datasets/maxwellziweiwei/vi-th-zh-parallel-corpus.Arabic_MSCOCO_1st_ViT-B-16-plus-240Arabic_3M_5M_ViT-B-16-plus-240
Loading the training split as follows:
from datasets import load_dataset
ds_train = load_dataset("Arabic-Clip/Arabic_3M_5M_ViT-B-16-plus-240", split="train")
ds_train
# Dataset({
# features: ['index', 'url', 'en_caption', 'embeddings_en', 'caption_ar'],
# num_rows: 200000
# })
Loading the validation split as follows:
from datasets import load_dataset
ds_validation = load_dataset("Arabic-Clip/Arabic_3M_5M_ViT-B-16-plus-240", split="validation")… See the full description on the dataset page: https://huggingface.co/datasets/Arabic-Clip/Arabic_3M_5M_ViT-B-16-plus-240.Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationArabic_3M_5M_ViT-B-16-SigLIP-512
Loading the training split as follows:
from datasets import load_dataset
ds_train = load_dataset("Arabic-Clip/Arabic_3M_5M_ViT-B-16-SigLIP-512", split="train")
ds_train
# Dataset({
# features: ['index', 'url', 'en_caption', 'embeddings_en', 'caption_ar'],
# num_rows: 2000000
# })
Loading the validation split as follows:
from datasets import load_dataset
ds_validation = load_dataset("Arabic-Clip/Arabic_3M_5M_ViT-B-16-SigLIP-512", split="validation")… See the full description on the dataset page: https://huggingface.co/datasets/Arabic-Clip/Arabic_3M_5M_ViT-B-16-SigLIP-512.vitl
Model Card for DINOv3
DINOv3 is a family of versatile vision foundation models that outperforms the specialized state of the art across a broad range of settings, without fine-tuning. DINOv3 produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models.
Model Details
These are Vision Transformer and ConvNeXt models trained following the method described in… See the full description on the dataset page: https://huggingface.co/datasets/simon123905/vitl.IMAD
Dataset Summary
This dataset contains data from the paper IMage Augmented multi-modal Dialogue: IMAD.
The main feature of this dataset is the novelty of the task. It has been generated specifically for the purpose of image interpretation in a dialogue context.
Some of the dialogue utterances have been replaced with images, allowing a generative model to be trained to restore the initial utterance.
The dialogues are sourced from multiple dialogue datasets (DailyDialog, Commonsense… See the full description on the dataset page: https://huggingface.co/datasets/VityaVitalich/IMAD.GraphCode-Bench-500-v0
GraphCode-Bench-500-v0
GraphCode-Bench is a benchmark for evaluating LLMs on call-graph reasoning — given a function in a real-world repository, can a model identify which functions call it (upstream) or which functions it calls (downstream), across 1 and 2 hops?
Models are evaluated agentically: they receive read-only filesystem tools (list_directory, read_file, search_in_file) and up to 10 turns to explore the codebase before producing an answer.
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/VittorioRossi/GraphCode-Bench-500-v0.vitamins-supplements-NER
Dataset Card for turkish-nlp-suite/vitamins-supplements-NER
Dataset Summary
The Vitamins and Supplements NER Dataset is a NER dataset containing customer reviews with entity and span annotations. User reviews were collected from a popular supplement products e-
commerce website Vitaminler.com.
Each customer review in the Vitamins and Supplements NER Dataset describes a customer’s experience with a supplement product in terms of that product’s effectiveness, side… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/vitamins-supplements-NER.mscoco_captions_ViT-B-16-SigLIP-512mscoco_captions_en_ar_all_splits_ViT_B_16_plus_240
