datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ViewSpatial-Bench
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Dataset Description
We introduce ViewSpatial-Bench, a comprehensive benchmark with over 5,700 question-answer pairs across 1,000+ 3D scenes from ScanNet and MS-COCO validation sets. This benchmark evaluates VLMs' spatial localization capabilities from multiple perspectives, specifically testing both egocentric (camera) and allocentric (human subject) viewpoints across… See the full description on the dataset page: https://huggingface.co/datasets/lidingm/ViewSpatial-Bench.viet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.view2space-v1
VIEW2SPACE v1
VIEW2SPACE v1 is a multi-view vision-language evaluation dataset for spatial reasoning.
Associated paper:
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations - ECCV 2026 🚀
arXiv: 2603.16506
Project Page: Project Page
Related VIEW2SPACE Releases
Training release: Pokerme/view2space-train
4B model checkpoint: Pokerme/view2space_4b
Collection: Pokerme/view2space
The public release is organized into three subsets:
count… See the full description on the dataset page: https://huggingface.co/datasets/Pokerme/view2space-v1.Vietnamese-yfcc15m-OpenAICLIPvietprofs
VietProfs
VietProfs is a community-maintained directory of Vietnamese and Vietnamese-diaspora academics at universities and eligible public or nonprofit scholarly research institutes worldwide.
This repository publishes the current complete, unmodified roster from VietProfs. It is a JSON array whose records describe a person's current academic appointment and may include research areas, education, honors, public profile links, and portrait source URLs. Portrait image files are… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhvuh/vietprofs.vqa-food-vietnamesevietnamese-medicinal-herb-vqa
Vietnamese Medicinal Herb VQA
This dataset contains Vietnamese visual question answering samples for medicinal herbs and traditional medicinal materials.
Each JSONL row has the following fields:
{
"id": "P000002676",
"name": "Cây Hẹ",
"image": "images/P000002266.jpg",
"question": "Hoa trong ảnh có màu gì?",
"answer": "Màu trắng"
}
Splits
Split
Images
QA pairs
Train
3032
23211
Validation
379
2998
Test
380
2952
The test split includes a… See the full description on the dataset page: https://huggingface.co/datasets/gdllelf/vietnamese-medicinal-herb-vqa.fact-checking-vietnamese-newsVietnamese-Food-VQA
Vietnamese Food VQA Dataset
Bộ dữ liệu Hỏi đáp Hình ảnh (VQA) chuyên biệt cho các món ăn Việt Nam. Được xây dựng phục vụ cho mục đích nghiên cứu và học tập trong lĩnh vực Deep Learning.
Thông tin bộ dữ liệu
Ngôn ngữ: Tiếng Việt
Lĩnh vực: Ẩm thực Việt Nam (Phở, Bún bò, Bánh mì, Cơm tấm, v.v.)
Số lượng mẫu: ~2500 - 10000 QA pairs (bao gồm dữ liệu đã augmented)
Định dạng:
Ảnh: .jpg
Nhãn: .json (train, val, test)
Cấu trúc thư mục trên Hugging Face
data/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Hoppaaa/Vietnamese-Food-VQA.
