datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.cardiac_cine_mnms
M&Ms (Cardiac Cine-MRI)
Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation
Views: SAX (ED/ES)
Splits: train / val / test
Data Structure (per example)
sax_ed, sax_ed_gt
sax_es, sax_es_gt
Optional: sax_t (if present)
Metadata columns listed below
Columns
Imaging
pid
sax_ed, sax_ed_gt, sax_es, sax_es_gt
sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.Traffic-sign-detection-VietNam
Vietnam Traffic Sign Detection Dataset
This repository contains the dataset for detecting road traffic signs in Vietnam using the state-of-the-art YOLO object detection model.
📂 Repository Structure
The dataset is structured in the standard YOLO format, containing images and corresponding annotations divided into training, validation, and testing sets.
├── classid.xlsx # Excel file mapping class IDs to names
├── dataset/
│ ├── train/ #… See the full description on the dataset page: https://huggingface.co/datasets/star092304/Traffic-sign-detection-VietNam.cardiac_cine_mnms2
M&Ms2 (Cardiac Cine-MRI, RV Focus)
Processed NIfTI cine-MRI data derived from the M&Ms2 challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation (RV focus)
Views: SAX + LAX 4C (LAX 2C if present)
Splits: train / val / test
Data Structure (per example)
SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt
LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt
LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt
Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.VietJobsTraffic-sign-detection-VietNam
Vietnam Traffic Sign Detection Dataset
This repository contains the dataset for detecting road traffic signs in Vietnam using the state-of-the-art YOLO object detection model.
📂 Repository Structure
The dataset is structured in the standard YOLO format, containing images and corresponding annotations divided into training, validation, and testing sets.
├── classid.xlsx # Excel file mapping class IDs to names
├── dataset/
│ ├── train/ #… See the full description on the dataset page: https://huggingface.co/datasets/Minh124689/Traffic-sign-detection-VietNam.cardiac_cine_emidec
EMIDEC (Delayed-Enhancement CMR)
Processed NIfTI EMIDEC dataset for myocardial infarction assessment.
Dataset Summary
Modality: DE-MRI
Task: Segmentation (LV, myocardium, infarction, no-reflow)
Splits: train / val / test
Data Structure (per example)
image
label
Metadata columns listed below
Columns
Imaging
pid
image, label
Metadata (all columns)
gender, age, tobacco, overweight, arterial_hypertension, diabetes
family_history, ecg, troponin… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_emidec.bridge_episode_viewpoints
Bridge Episode Viewpoints
This repository provides the selected sampled viewpoint for each Bridge episode in OXE-AugE.
Files
File
Description
episode_best_viewpoint_train.csv
Train split episode-to-viewpoint mapping.
episode_best_viewpoint_test.csv
Test split episode-to-viewpoint mapping.
Columns
Column
Description
episode_id
Episode index in the Bridge split.
best_viewpoint
Selected sampled viewpoint index for that episode.
cardiac_cine_myops2020
MyoPS 2020 (Multi-Sequence CMR)
Processed NIfTI multi-sequence CMR data (C0/DE/T2) for myocardial pathology segmentation.
Dataset Summary
Modality: CMR (C0, DE, T2)
Task: Segmentation of myocardium/pathology
Splits: train / test
Data Structure (per example)
c0, de, t2, label
Metadata columns listed below
Columns
Imaging
pid
c0, de, t2, label
Metadata (all columns)
orig_spacing_x, orig_spacing_y, orig_spacing_z
n_slices
crop_lower_x… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_myops2020.vietnamese-toxic-commentfull-dubai-pulsecardiac_cine_kaggle
Kaggle Cardiac Cine-MRI
Processed NIfTI cine-MRI sequences from the Kaggle cardiac dataset.
Dataset Summary
Modality: CMR cine MRI
Views: SAX, LAX 2C, LAX 4C (time series)
Splits: train / val (if present)
Data Structure (per example)
sax_t, lax_2c_t, lax_4c_t
Metadata columns listed below
Columns
Imaging
pid
sax_t, lax_2c_t, lax_4c_t
Metadata (all columns)
n_slices, n_frames
original_sax_spacing_x, original_sax_spacing_y… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_kaggle.vietnamese-healthcare-dataset
Vietnamese Healthcare Synthetic Patient Records
This dataset contains synthetic Vietnamese healthcare identity records from multiple source systems, plus a canonical synthetic patient table used by the generator.
All records are synthetic and are intended for entity resolution, record linkage, and Vietnamese identity-field preprocessing experiments.
Included Files
Only the following CSV files are included in this upload:
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/adachankawai/vietnamese-healthcare-dataset.adcumen-viewer-emotions
AdCumen Viewer Emotions Dataset
Dataset for the paper "Decoding Viewer Emotions in Video Ads" by Alexey Antonov, Shravan Sampath Kumar, Jiefei Wei, William Headley, Orlando Wood, and Giovanni Montana, published in Nature Scientific Reports.
Code: github.com/gmontana/DecodingViewerEmotions
Model weights: dnamodel/tsam-viewer-emotions
Dataset Description
The dataset consists of 26,637 five-second video clips extracted from video advertisements, annotated for seven… See the full description on the dataset page: https://huggingface.co/datasets/dnamodel/adcumen-viewer-emotions.hanzi-sino-vietnamese
HSK × Sino-Vietnamese (Hán-Việt) character dataset
768 HSK characters joined with their Sino-Vietnamese (Hán-Việt) readings, radical breakdowns and hand-written memory hooks in Vietnamese.
Open HSK wordlists are plentiful. The Sino-Vietnamese layer is what is missing from all of them — and it is the layer that matters most for the ~1 million Vietnamese speakers studying Chinese, because roughly 60% of Vietnamese vocabulary descends from Chinese. A learner meeting 学 (xué) already… See the full description on the dataset page: https://huggingface.co/datasets/kaihanzi/hanzi-sino-vietnamese.vietnamese-legal-corpus-20k-rawvietnamese-caucu-comments
Vietnamese Cau Cuu Facebook Comments
Dataset Summary
This dataset contains Vietnamese Facebook comments collected from a natural-disaster discussion thread and auto-labeled for binary emergency detection.
The target task is to detect whether a comment is a real-time rescue request (cau_cuu) versus a non-emergency comment (khong_phai_cau_cuu).
This release is intended as a bootstrap dataset for triage modeling and should be treated as a weakly supervised resource. Human… See the full description on the dataset page: https://huggingface.co/datasets/dat201204/vietnamese-caucu-comments.vietnamese_toxic_corecardiac_cine_msd
MSD Cardiac — Task02_Heart (Left Atrium Segmentation)
Processed NIfTI data from the Medical Segmentation Decathlon Task02 (Heart).
The goal is to segment the left atrium from mono-modal MR images.
Dataset Summary
Modality: MRI
Task: Left atrium segmentation
Patients: 30 total (20 train, 10 test)
Labels: 0 = background, 1 = left atrium
Splits: train (with labels), test (images only, no public labels)
Data Structure (per patient)
Each patient directory… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_msd.ott-viewer-dropoff-retention
🎬 OTT Viewer Drop-Off & Retention Risk Dataset (v1.0)
📌 Overview
This dataset provides episode-level viewer behavior data for OTT (streaming) TV series, focused on drop-off patterns, retention risk, and engagement dynamics across episodes and seasons.
Unlike traditional catalog datasets (genres, ratings, cast), this dataset is designed to support realistic retention analysis, similar to how streaming platforms study when and why viewers stop watching.
Each row… See the full description on the dataset page: https://huggingface.co/datasets/Eklavya16/ott-viewer-dropoff-retention.Vietnam-higher-education-lawViewSpatial-Bench-vlmevalviet-robot-laundry-reminder-log
viet-robot-laundry-reminder-log
A synthetic dataset for laundry-related reminders generated by a home robot.
Columns:
id: row id
timestamp_local: local time string
basket_location: where the laundry basket is
fill_level_pct: estimated fill level 0–100
days_since_last_wash: integer days
reminder_sent: yes/no
user_action: did_wash, snoozed, ignored
note_en: short English note
License
MIT
vietnam-deforestation-risk-sample
🌳 Vietnam Deforestation Risk Sample Dataset
A public derived demonstration dataset representing a subset of forest pixels in Gia Lai Province, Vietnam, configured for teaching, reproducibility audits, and modeling workflows.
Dataset Structure
system:index: Unique pixel identifier.
defo: Deforestation flag.
district: Ecological district (e.g. KBang or MangYang).
elevation / mean_elevation_m: Elevation in meters.
mean_slope_deg: Slope in degrees.… See the full description on the dataset page: https://huggingface.co/datasets/MahdiFattahi/vietnam-deforestation-risk-sample.synthetic_dropout_dataset_vietnam_100k_final_lhu
Student Dropout Prediction Dataset (Lac Hong University - Synthetic)
This dataset is a synthetically generated dataset representing student academic and behavioral data at Lac Hong University.
It is intended for machine learning tasks that predict student dropout risks.
📊 Features
StudentID: Unique student identifier (format: 1YYxxxxxx)
LUC: Lack of University Commitment (Likert 1–5)
DCC: Degree Commitment Conflict (Likert 1–5)
ITM: Ineffective Time Management (Likert… See the full description on the dataset page: https://huggingface.co/datasets/LHUThacSi/synthetic_dropout_dataset_vietnam_100k_final_lhu.viet-robot-security-events
viet-robot-security-events
A synthetic dataset of simple security-related events in a smart home.
Columns:
id: row id
event_type: door_open, window_open, motion_detected, noise_high, device_disconnected
location: where the event happened
time_hour: integer hour of day (0-23)
is_night: yes/no
human_confirmed: yes/no (whether a human was seen)
note: short English note
For demos only, not real security logs.
License
MIT
viet-robot-music-scenes
viet-robot-music-scenes
A small synthetic dataset of music scenes and preferences for a home robot.
Columns:
id: row id
scene_id: identifier of the scene
description_vi: description in Vietnamese
description_en: description in English
time_of_day: morning, afternoon, evening, night
mood: calm, focused, energetic, sleep
volume_level: 1–10 suggested volume
Manually authored for demo purposes.
License
MIT
viet-robot-energy-usage-log-v2
viet-robot-energy-usage-log
A small synthetic dataset with coarse energy usage estimates for different
robot activities.
Columns:
id: row id
activity: cleaning, mapping, idle_docked, patrolling, voice_only
duration_min: duration in minutes
energy_wh: estimated energy used in Wh
time_of_day: morning, afternoon, evening, night
note_en: short English context
Values are made up for demo purposes and not based on real hardware.
License
MIT
viet-robot-weather-scenes
viet-robot-weather-scenes
A tiny synthetic dataset of weather scenes and suggested indoor robot modes.
Columns:
id: row id
weather: sunny, cloudy, raining, storm, hot, cold
time_of_day: morning, afternoon, evening, night
outside_temp_c: approximate outdoor temperature
outside_humidity_pct: approximate outdoor humidity
suggested_mode: normal, quiet, safe_check_windows, extra_dry, stay_docked
note_en: short English note
For demo and educational purposes only.
License… See the full description on the dataset page: https://huggingface.co/datasets/hoangs/viet-robot-weather-scenes.
