datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vima
Dataset Card for "vima"
More Information needed
ViMD_Dataset
Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges (Main EMNLP 2024)
Introduction
This document presents the accompanying dataset for the paper titled "Multi-Dialect Vietnamese: Task, Dataset, Baseline Models, and Challenges". The dataset, referred to as the Vietnamese Multi-Dialect (ViMD) dataset, is a comprehensive resource designed to capture the linguistic diversity represented by 63 provincial dialects spoken across Vietnam. The paper is… See the full description on the dataset page: https://huggingface.co/datasets/nguyendv02/ViMD_Dataset.GeoMeld
🌍 GeoMeld Multi-Modal Earth Observation Dataset (WebDataset)
GeoMeld is a large-scale multi-modal remote sensing dataset introduced in our CVPRW 2026 paper on semantically grounded foundation modeling.
GeoMeld contains approximately 2.5 million spatially aligned samples spanning heterogeneous sensing modalities and spatial resolutions, paired with semantically grounded captions generated through an agentic pipeline.
The dataset is designed to support multimodal representation… See the full description on the dataset page: https://huggingface.co/datasets/vimageiitb/GeoMeld.ViMed-PET-CT
ViMed-PET-CT
📅 Update: April 23, 2026
🐛 Bug Fixes: Corrected field mismatches (blank/missing fields) and date/filename inconsistencies. Restored missing metadata for patient 1701 (Dec 2023).
✨ New Feature: Added English translations of reports (/reports_en) using Gemma-4-26B-A4B-it.
ℹ️ About the dataset
🍴 Forked and optimized compression of dacthai2807/ViMed-PET, converting .npy and chunked zip files into .npz files.
📝 Better annotation and guideline.
📂… See the full description on the dataset page: https://huggingface.co/datasets/thainamhoang/ViMed-PET-CT.ViMed-PET-part1
Dataset description for three years: 2017, 2018, 2019
This dataset contains data from three years (2017, 2018, 2019). Each year has several month folders, which are named as THANG {month}.
Each year folder is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract all data folders.
Folder structure after extraction
Each folder named THANG {month} of a year is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2807/ViMed-PET-part1.vima_rawViMedCSS
🩺 ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset (LREC 2026)
📖 Overview
ViMedCSS is a Vietnamese medical speech dataset for code-switching ASR, where each utterance contains at least one non-Vietnamese (mainly English) medical term embedded in Vietnamese speech.
📊 Dataset Statistics
Split Statistics (from ViMedCSS-Metadata)
Split
# Rows
Duration (hours)
Avg duration (s)
Total CS terms
train
11,832
24.30
7.39
12,314… See the full description on the dataset page: https://huggingface.co/datasets/tensorxt/ViMedCSS.ViMed-PET-part2
Dataset description for year 2023
This dataset contains data from 8 months: January to September, except August, stored in the following folders respectively:
THANG 1
THANG 2
...
THANG 7
THANG 9
The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract the data folders.
Folder structure after extraction
Each folder named THANG {month} is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/Duc2305/ViMed-PET-part2.ViMD_ReigonGroup
Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges (Main EMNLP 2024)
Introduction
This document presents the accompanying dataset for the paper titled "Multi-Dialect Vietnamese: Task, Dataset, Baseline Models, and Challenges". The dataset, referred to as the Vietnamese Multi-Dialect (ViMD) dataset, is a comprehensive resource designed to capture the linguistic diversity represented by 63 provincial dialects spoken across Vietnam. The paper is… See the full description on the dataset page: https://huggingface.co/datasets/notlee203/ViMD_ReigonGroup.En-ViMedNER
En-ViMedNER Dataset
Introduction
En-ViMedNER is the first English–Vietnamese parallel biomedical Named Entity Recognition (NER) corpus annotated with UMLS semantic types. It is derived from MedMentions ST21pv (MM-ST21pv) and provides a shared, language-neutral label space so that Vietnamese biomedical NER can be compared directly with existing UMLS-based English resources such as MedMentions.
The corpus contains 4,392 PubMed abstract pairs, 44,892… See the full description on the dataset page: https://huggingface.co/datasets/nhuvo/En-ViMedNER.ViMoGen-228K
ViMoGen-228K: The Quest for Generalizable Motion Generation
Project Page | Paper | Code | MBench Leaderboard
ViMoGen-228K is a large-scale dataset comprising 228,000 high-quality motion samples. It integrates high-fidelity optical MoCap data with semantically annotated motions from web videos and synthesized samples generated by state-of-the-art video generation (ViGen) models. The dataset includes both text-motion pairs and text-video-motion triplets, designed to expand… See the full description on the dataset page: https://huggingface.co/datasets/wruisi/ViMoGen-228K.vimeo90k_triplet
Dataset Card for "vimeo90k_triplet"
More Information needed
smart-repro-imagenet-resnet50-logitsViMed-PET-part3
Dataset description for year 2023
This dataset contains data from three months: October, November, and December, stored in the following folders respectively:
THANG 10
THANG 11
THANG 12
The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract the data folders.
Folder structure after extraction
Each folder named THANG {month} is divided into 3 subfolders, corresponding to 2… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2k/ViMed-PET-part3.ViMRHP
ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation
This dataset is presented in the paper ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation.
Dataset Usage Note
This dataset is originally constructed and provided for the Multimodal Review Helpfulness Prediction (MRHP) task. In addition to the main fields designed for MRHP, we… See the full description on the dataset page: https://huggingface.co/datasets/trucmtnguyen/ViMRHP.vimeo1080pvimeo-90k-mini
Vimeo-90k-Mini
A 25% random subset of the official Vimeo-90k Triplet dataset,
used for video frame interpolation tasks.
Splits
train: ~13,000 triplets
test: ~1000 triplets
Structure
Each example contains three consecutive video frames (im1, im2, im3).
The task is typically to predict im2 given im1 and im3.
Original Dataset
Paper: Video Enhancement with Task-Oriented Flow
Authors: Tianfan Xue et al.
vimeo-90k-medium
Vimeo-90k-Medium
A 50% random subset of the official Vimeo-90k Triplet dataset,
used for video frame interpolation tasks.
Splits
train: ~26,000 triplets
test: ~2000 triplets
Structure
Each example contains three consecutive video frames (im1, im2, im3).
The task is typically to predict im2 given im1 and im3.
Original Dataset
Paper: Video Enhancement with Task-Oriented Flow
Authors: Tianfan Xue et al.
icml26332-fullbatch-results
Reproduction results — "Full-Batch Gradient Descent Outperforms One-Pass SGD" (ICML 2026 #26332)
Raw and aggregated results for an independent reproduction of
arXiv:2602.02431 /
OpenReview QItZDBVCT0.
All numbers were produced locally on 2x NVIDIA RTX 4000 Ada with the scripts in the
reproduction logbook (scripts/sim.py, scripts/sweep_spherical.py,
scripts/sweep_online_sgd.py, scripts/sweep_squared_gd.py, scripts/spectral_audit.py,
scripts/thm32_bound.py).
file
contents… See the full description on the dataset page: https://huggingface.co/datasets/vimarsh/icml26332-fullbatch-results.Ranked-Vimeo-90K
Ranked-Vimeo-90K
A saliency-ranked video dataset for intelligent perceptual machine video coding.
Overview
Ranked-Vimeo-90K is a large-scale video dataset with pixel-level saliency ranking maps, derived from the Vimeo-90K septuplet dataset. It provides per-frame saliency ordering for 7-frame video clips, designed for research on saliency-aware video compression, intelligent perceptual coding, and machine vision oriented video transmission.
Authors… See the full description on the dataset page: https://huggingface.co/datasets/spzhu/Ranked-Vimeo-90K.ViMedCSS-Cop
🩺 ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset (LREC 2026)
📖 Overview
ViMedCSS is a Vietnamese medical speech dataset for code-switching ASR, where each utterance contains at least one non-Vietnamese (mainly English) medical term embedded in Vietnamese speech.
📊 Dataset Statistics
Split Statistics (from ViMedCSS-Metadata)
Split
# Rows
Duration (hours)
Avg duration (s)
Total CS terms
train
11,832
24.30
7.39
12,314… See the full description on the dataset page: https://huggingface.co/datasets/shannonnonshan/ViMedCSS-Cop.ViMUL-Bench
ViMUL-Bench: A Culturally-diverse Multilingual Multimodal Video Benchmark
Overview
The evaluation toolkit to be used is lmms-eval. This toolkit facilitates the evaluation of models across multiple tasks and languages.
Key Features
🌍 14 Languages: English, Chinese, Spanish, French, German, Hindi, Arabic, Russian, Bengali, Urdu, Sinhala, Tamil, Swedish, Japanese🎭 15 Categories: Including 8 culturally diverse categories (lifestyles, festivals, foods… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ViMUL-Bench.ViMathQAvimgolf_challenges_and_solutionsChallenges and descriptions from vimgolf.com
Scraper code at: https://github.com/James4Ever0/agi_computer_control/tree/master/scrape_vimgolf_challenges_and_solutions
Dataclasses used in: https://github.com/James4Ever0/vimgolf-gym/blob/main/vimgolf_gym/dataclasses.py
File structure:
challenges
|_ <challenge_hash>
|_ metadata.json
|_ challenge.json
|_ worst_solution.json
worst_solution.json is scraped from vimgolf.com.
Anonymous viewer can only… See the full description on the dataset page: https://huggingface.co/datasets/James4Ever0/vimgolf_challenges_and_solutions.ViMU
ViMU: Benchmarking Video Metaphorical Understanding
Qi Li, Xinchao Wang*
*Corresponding author
xML Lab, National University of Singapore
Our GitHub repository contains the evaluation scripts for ViMU, a benchmark for video metaphorical understanding. The code evaluates multimodal models on four tasks:
Open-ended interpretation (OE)
Evidence grounding (EG)
Rhetoric mechanism identification (RM)
Social value signal identification (SV)
Directory Structure
Expected… See the full description on the dataset page: https://huggingface.co/datasets/LIQIIIII/ViMU.vimgolf-public-challenges-inspect-evalvimedaqa-nli-g1vimqa
VIMQA
VIMQA is a Vietnamese dataset for advanced reasoning and explainable multi-hop
question answering. Each question requires combining facts from two different
Vietnamese Wikipedia articles, and every example ships with sentence-level
supporting facts so a model's reasoning chain can be evaluated, not just its
final answer.
The schema follows the HotpotQA
convention, so tooling written for HotpotQA transfers with minimal changes.
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/nguyenlab/vimqa.vimeo2kAR-VIM
AR-VIM: Dataset for Visual Information Manipulation in Augmented Reality
This is the official dataset of paper: Detecting Visual Information Manipulation Attacks in Augmented Reality: A Multimodal Semantic Reasoning Approach. The paper has been selected as the Best Paper Award by IEEE ISMAR 2025, and selected as a special issue of IEEE TVCG Journal.
This dataset contains paired Raw and Augmented videos collected to evaluate visual information manipulation… See the full description on the dataset page: https://huggingface.co/datasets/HarbingerKX/AR-VIM.
