datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.cardiac_cine_mnms
M&Ms (Cardiac Cine-MRI)
Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation
Views: SAX (ED/ES)
Splits: train / val / test
Data Structure (per example)
sax_ed, sax_ed_gt
sax_es, sax_es_gt
Optional: sax_t (if present)
Metadata columns listed below
Columns
Imaging
pid
sax_ed, sax_ed_gt, sax_es, sax_es_gt
sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.Vietnam-Stock-Symbols-and-Metadata
Vietnam Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Vietnam.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Vietnam-Stock-Symbols-and-Metadata.Traffic-sign-detection-VietNam
Vietnam Traffic Sign Detection Dataset
This repository contains the dataset for detecting road traffic signs in Vietnam using the state-of-the-art YOLO object detection model.
📂 Repository Structure
The dataset is structured in the standard YOLO format, containing images and corresponding annotations divided into training, validation, and testing sets.
├── classid.xlsx # Excel file mapping class IDs to names
├── dataset/
│ ├── train/ #… See the full description on the dataset page: https://huggingface.co/datasets/star092304/Traffic-sign-detection-VietNam.cardiac_cine_mnms2
M&Ms2 (Cardiac Cine-MRI, RV Focus)
Processed NIfTI cine-MRI data derived from the M&Ms2 challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation (RV focus)
Views: SAX + LAX 4C (LAX 2C if present)
Splits: train / val / test
Data Structure (per example)
SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt
LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt
LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt
Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.VietJobsvietnamese_sms_dataset
Bộ dữ liệu SMS lừa đảo tiếng Việt được đảm bảo chất lượng (Official Release)
(English Below)
Chào mừng bạn đến với kho lưu trữ chính thức của Bộ dữ liệu SMS lừa đảo tiếng Việt được đảm bảo chất lượng.
Đây là một bộ dữ liệu được xây dựng nhằm phục vụ nghiên cứu trong các lĩnh vực an ninh mạng, xử lý ngôn ngữ tự nhiên (NLP) và học máy, với trọng tâm là bài toán phát hiện tin nhắn SMS rác/lừa đảo.
Bộ dữ liệu này được tổng hợp từ các tin nhắn SMS thực tế trong cuộc sống. Không… See the full description on the dataset page: https://huggingface.co/datasets/trannguyenthaituan/vietnamese_sms_dataset.turn-detection-vietnameseNguồn dữ liệu: vi-wiki-conversational-search
Tỷ lệ Complete:Incomplete = 244304:366451
Đã lưu 610755 samples vào training_data.csv
Đã lưu 6189 samples vào test_data.csv
vietnamese_news_human_ai
Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT
This is the official dataset accompanying the paper Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT, which was accepted at ICCIES 2025 and published in Computational Intelligence in Engineering Science (Springer CCIS, vol. 2587).
You can read the paper here: Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT
Abstract
The emergence… See the full description on the dataset page: https://huggingface.co/datasets/ICCIES-2025-DetectAI/vietnamese_news_human_ai.View-Spatial-BenchVietOnlineNews
VietOnlineNews: Vietnamese Online News Topic Classification Dataset
Dataset Description
VietOnlineNews is a Vietnamese online news dataset constructed for the task of single-label multi-class topic classification. Each sample corresponds to one news article and is assigned exactly one main topic label through the category field.
The dataset was collected from multiple Vietnamese online news sources and processed through a data cleaning pipeline to remove… See the full description on the dataset page: https://huggingface.co/datasets/VLUS06/VietOnlineNews.Traffic-sign-detection-VietNam
Vietnam Traffic Sign Detection Dataset
This repository contains the dataset for detecting road traffic signs in Vietnam using the state-of-the-art YOLO object detection model.
📂 Repository Structure
The dataset is structured in the standard YOLO format, containing images and corresponding annotations divided into training, validation, and testing sets.
├── classid.xlsx # Excel file mapping class IDs to names
├── dataset/
│ ├── train/ #… See the full description on the dataset page: https://huggingface.co/datasets/Minh124689/Traffic-sign-detection-VietNam.vietnamese-poetry-corpuscardiac_cine_emidec
EMIDEC (Delayed-Enhancement CMR)
Processed NIfTI EMIDEC dataset for myocardial infarction assessment.
Dataset Summary
Modality: DE-MRI
Task: Segmentation (LV, myocardium, infarction, no-reflow)
Splits: train / val / test
Data Structure (per example)
image
label
Metadata columns listed below
Columns
Imaging
pid
image, label
Metadata (all columns)
gender, age, tobacco, overweight, arterial_hypertension, diabetes
family_history, ecg, troponin… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_emidec.vietnamese-speech-recognition
Vietnamese Speech Dataset
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.vietnamese_summarization_vr_vrp_resources
Vietnamese Summarization VR/VRP Resources
This repository consolidates the experimental resources associated with the paper:
Reinforcement Learning With Verifier Guidance and Penalty Shaping for Vietnamese Summarization Using Small Language Models
It contains:
CSV exports for Hugging Face Data Viewer,
Links to the released best checkpoints,
The link to the frozen evaluator MultiEvalSumViet2.
Representative Source Code
Dataset files used in the paper
Split… See the full description on the dataset page: https://huggingface.co/datasets/phuongntc/vietnamese_summarization_vr_vrp_resources.bridge_episode_viewpoints
Bridge Episode Viewpoints
This repository provides the selected sampled viewpoint for each Bridge episode in OXE-AugE.
Files
File
Description
episode_best_viewpoint_train.csv
Train split episode-to-viewpoint mapping.
episode_best_viewpoint_test.csv
Test split episode-to-viewpoint mapping.
Columns
Column
Description
episode_id
Episode index in the Bridge split.
best_viewpoint
Selected sampled viewpoint index for that episode.
cardiac_cine_myops2020
MyoPS 2020 (Multi-Sequence CMR)
Processed NIfTI multi-sequence CMR data (C0/DE/T2) for myocardial pathology segmentation.
Dataset Summary
Modality: CMR (C0, DE, T2)
Task: Segmentation of myocardium/pathology
Splits: train / test
Data Structure (per example)
c0, de, t2, label
Metadata columns listed below
Columns
Imaging
pid
c0, de, t2, label
Metadata (all columns)
orig_spacing_x, orig_spacing_y, orig_spacing_z
n_slices
crop_lower_x… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_myops2020.vietnamese-comment-sentiment
Vietnamese Comment Sentiment Dataset
Overview
This dataset contains Vietnamese comments collected from various social networks to facilitate sentiment analysis. Each comment is labeled to indicate its sentiment, making it useful for natural language processing tasks.
Data Source
The comments were crawled from several social network platforms, ensuring a diverse range of expressions and contexts within Vietnamese language usage.
Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/minhtoan/vietnamese-comment-sentiment.full-dubai-pulsele-hoi-viet-nam
Lễ hội, lễ tết và ngày nghỉ Việt Nam
Vietnamese festivals, seasonal feasts and public holidays
1. Mô tả · Description
Lễ hội dân gian, lễ tết và ngày nghỉ lễ ở Việt Nam, kèm ngày âm lịch hoặc dương lịch cố định của từng dịp.
Vietnamese folk festivals, seasonal feasts and public holidays, each with its fixed lunar or solar date.
Số dòng · Rows: 48
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/le-hoi-viet-nam.VietNews-Abs-Sum
VietNews-Abs-Sum
A dataset for Vietnamese Abstractive Summarization task.It includes all articles from Vietnews (VNDS) dataset which was released by Van-Hau Nguyen et al.The articles were collected from tuoitre.vn, vnexpress.net, and nguoiduatin.vn online newspaper by the authors.
Introduction
This dataset was extracted from Train/Val/Test split of Vietnews dataset. All files from test_tokenized, train_tokenized and val_tokenized directories are fetched and preprocessed… See the full description on the dataset page: https://huggingface.co/datasets/ithieund/VietNews-Abs-Sum.cardiac_cine_kaggle
Kaggle Cardiac Cine-MRI
Processed NIfTI cine-MRI sequences from the Kaggle cardiac dataset.
Dataset Summary
Modality: CMR cine MRI
Views: SAX, LAX 2C, LAX 4C (time series)
Splits: train / val (if present)
Data Structure (per example)
sax_t, lax_2c_t, lax_4c_t
Metadata columns listed below
Columns
Imaging
pid
sax_t, lax_2c_t, lax_4c_t
Metadata (all columns)
n_slices, n_frames
original_sax_spacing_x, original_sax_spacing_y… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_kaggle.vietnamese-social-comments
🇻🇳 Bộ dữ liệu phân loại bình luận tiếng Việt
Bộ dữ liệu này bao gồm 4.896 bình luận tiếng Việt được thu thập từ nhiều nền tảng mạng xã hội phổ biến như TikTok, Facebook, YouTube,...Mỗi bình luận được gán nhãn theo 2 cấp độ:
label: thể hiện cảm xúc hoặc thái độ tổng thể.
category: phân loại chi tiết theo ngữ nghĩa hoặc mục đích cụ thể của câu.
🔖 Cấu trúc dữ liệu
Trường
Kiểu dữ liệu
Mô tả
comment
string
Văn bản bình luận (có thể viết tắt, không dấu… See the full description on the dataset page: https://huggingface.co/datasets/vanhai123/vietnamese-social-comments.vietnamese-healthcare-dataset
Vietnamese Healthcare Synthetic Patient Records
This dataset contains synthetic Vietnamese healthcare identity records from multiple source systems, plus a canonical synthetic patient table used by the generator.
All records are synthetic and are intended for entity resolution, record linkage, and Vietnamese identity-field preprocessing experiments.
Included Files
Only the following CSV files are included in this upload:
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/adachankawai/vietnamese-healthcare-dataset.adcumen-viewer-emotions
AdCumen Viewer Emotions Dataset
Dataset for the paper "Decoding Viewer Emotions in Video Ads" by Alexey Antonov, Shravan Sampath Kumar, Jiefei Wei, William Headley, Orlando Wood, and Giovanni Montana, published in Nature Scientific Reports.
Code: github.com/gmontana/DecodingViewerEmotions
Model weights: dnamodel/tsam-viewer-emotions
Dataset Description
The dataset consists of 26,637 five-second video clips extracted from video advertisements, annotated for seven… See the full description on the dataset page: https://huggingface.co/datasets/dnamodel/adcumen-viewer-emotions.vietnam-normalize-24kvietnamese-error-correction-corpus
Data Summary
The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets.
• Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text.
• Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs.
• Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.hanzi-sino-vietnamese
HSK × Sino-Vietnamese (Hán-Việt) character dataset
768 HSK characters joined with their Sino-Vietnamese (Hán-Việt) readings, radical breakdowns and hand-written memory hooks in Vietnamese.
Open HSK wordlists are plentiful. The Sino-Vietnamese layer is what is missing from all of them — and it is the layer that matters most for the ~1 million Vietnamese speakers studying Chinese, because roughly 60% of Vietnamese vocabulary descends from Chinese. A learner meeting 学 (xué) already… See the full description on the dataset page: https://huggingface.co/datasets/kaihanzi/hanzi-sino-vietnamese.
