datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.vietnamese_sms_dataset
Bộ dữ liệu SMS lừa đảo tiếng Việt được đảm bảo chất lượng (Official Release)
(English Below)
Chào mừng bạn đến với kho lưu trữ chính thức của Bộ dữ liệu SMS lừa đảo tiếng Việt được đảm bảo chất lượng.
Đây là một bộ dữ liệu được xây dựng nhằm phục vụ nghiên cứu trong các lĩnh vực an ninh mạng, xử lý ngôn ngữ tự nhiên (NLP) và học máy, với trọng tâm là bài toán phát hiện tin nhắn SMS rác/lừa đảo.
Bộ dữ liệu này được tổng hợp từ các tin nhắn SMS thực tế trong cuộc sống. Không… See the full description on the dataset page: https://huggingface.co/datasets/trannguyenthaituan/vietnamese_sms_dataset.turn-detection-vietnameseNguồn dữ liệu: vi-wiki-conversational-search
Tỷ lệ Complete:Incomplete = 244304:366451
Đã lưu 610755 samples vào training_data.csv
Đã lưu 6189 samples vào test_data.csv
vietnamese_news_human_ai
Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT
This is the official dataset accompanying the paper Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT, which was accepted at ICCIES 2025 and published in Computational Intelligence in Engineering Science (Springer CCIS, vol. 2587).
You can read the paper here: Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT
Abstract
The emergence… See the full description on the dataset page: https://huggingface.co/datasets/ICCIES-2025-DetectAI/vietnamese_news_human_ai.vietnamese-poetry-corpusvietnamese-speech-recognition
Vietnamese Speech Dataset
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.vietnamese_summarization_vr_vrp_resources
Vietnamese Summarization VR/VRP Resources
This repository consolidates the experimental resources associated with the paper:
Reinforcement Learning With Verifier Guidance and Penalty Shaping for Vietnamese Summarization Using Small Language Models
It contains:
CSV exports for Hugging Face Data Viewer,
Links to the released best checkpoints,
The link to the frozen evaluator MultiEvalSumViet2.
Representative Source Code
Dataset files used in the paper
Split… See the full description on the dataset page: https://huggingface.co/datasets/phuongntc/vietnamese_summarization_vr_vrp_resources.vietnamese-comment-sentiment
Vietnamese Comment Sentiment Dataset
Overview
This dataset contains Vietnamese comments collected from various social networks to facilitate sentiment analysis. Each comment is labeled to indicate its sentiment, making it useful for natural language processing tasks.
Data Source
The comments were crawled from several social network platforms, ensuring a diverse range of expressions and contexts within Vietnamese language usage.
Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/minhtoan/vietnamese-comment-sentiment.vietnamese-social-comments
🇻🇳 Bộ dữ liệu phân loại bình luận tiếng Việt
Bộ dữ liệu này bao gồm 4.896 bình luận tiếng Việt được thu thập từ nhiều nền tảng mạng xã hội phổ biến như TikTok, Facebook, YouTube,...Mỗi bình luận được gán nhãn theo 2 cấp độ:
label: thể hiện cảm xúc hoặc thái độ tổng thể.
category: phân loại chi tiết theo ngữ nghĩa hoặc mục đích cụ thể của câu.
🔖 Cấu trúc dữ liệu
Trường
Kiểu dữ liệu
Mô tả
comment
string
Văn bản bình luận (có thể viết tắt, không dấu… See the full description on the dataset page: https://huggingface.co/datasets/vanhai123/vietnamese-social-comments.vietnamese-healthcare-dataset
Vietnamese Healthcare Synthetic Patient Records
This dataset contains synthetic Vietnamese healthcare identity records from multiple source systems, plus a canonical synthetic patient table used by the generator.
All records are synthetic and are intended for entity resolution, record linkage, and Vietnamese identity-field preprocessing experiments.
Included Files
Only the following CSV files are included in this upload:
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/adachankawai/vietnamese-healthcare-dataset.vietnamese-error-correction-corpus
Data Summary
The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets.
• Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text.
• Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs.
• Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.hanzi-sino-vietnamese
HSK × Sino-Vietnamese (Hán-Việt) character dataset
768 HSK characters joined with their Sino-Vietnamese (Hán-Việt) readings, radical breakdowns and hand-written memory hooks in Vietnamese.
Open HSK wordlists are plentiful. The Sino-Vietnamese layer is what is missing from all of them — and it is the layer that matters most for the ~1 million Vietnamese speakers studying Chinese, because roughly 60% of Vietnamese vocabulary descends from Chinese. A learner meeting 学 (xué) already… See the full description on the dataset page: https://huggingface.co/datasets/kaihanzi/hanzi-sino-vietnamese.vietnamese-nom-poetry-translationvietnamese-toxic-commentvietnamese-sentiment-analysisvietnamese-SA-dataset
Vietnamese Sentiment Dataset
Dataset Description
This dataset contains Vietnamese user reviews labeled with sentiment classes for sentiment analysis tasks.
The dataset is designed for:
Sentiment classification
Vietnamese NLP research
Text classification benchmarking
Fine-tuning language models
Dataset Structure
Column
Type
Description
text
string
Vietnamese review/comment text
label
string
Sentiment label (pos, neg, neu)… See the full description on the dataset page: https://huggingface.co/datasets/h-i-e-u/vietnamese-SA-dataset.vietnamese-nom-latin-translationvietnamese-legal-corpus-20k-rawVietnamese_sentimentVietnamese-normThis dataset use Vinorm and Llama to normalize Vietnamese text
For example:
33/4 -> ba mươi ba tháng tư
43 tỷ USD -> bốn mươi ba tỉ đô la
Covid-19 -> covid mười chín
lần thứ VI -> lần thứ sáu
33% -> ba mươi ba phần trăm
U23 -> u hai mươi ba
iPhone 14 -> iphone mười bốn
năm 2023 -> năm hai không hai mươi ba
vietnamese-caucu-comments
Vietnamese Cau Cuu Facebook Comments
Dataset Summary
This dataset contains Vietnamese Facebook comments collected from a natural-disaster discussion thread and auto-labeled for binary emergency detection.
The target task is to detect whether a comment is a real-time rescue request (cau_cuu) versus a non-emergency comment (khong_phai_cau_cuu).
This release is intended as a bootstrap dataset for triage modeling and should be treated as a weakly supervised resource. Human… See the full description on the dataset page: https://huggingface.co/datasets/dat201204/vietnamese-caucu-comments.vietnamese_sms_phishing_sample
Quality-Assured Vietnamese SMS Phishing Dataset (Sample Evaluation Corpus)
This repository provides an open-access 300-message evaluation sample of the Quality-Assured Vietnamese SMS Phishing Dataset, a leakage-controlled benchmark constructed using a closed-loop multi-annotator agreement framework and near-duplicate template deduplication (J ≥ 0.85).
Authors & Research Team
Trần Nguyễn Thái Tuấn (Lead Researcher & Project Lead)
Lê Hoàng Khang (Researcher)
Nguyễn… See the full description on the dataset page: https://huggingface.co/datasets/trannguyenthaituan/vietnamese_sms_phishing_sample.Accent-and-accentless-Vietnamese-datasetXNLI_Vietnamese_tripletsCitation:
@InProceedings{conneau2018xnli,
author = {Conneau, Alexis
and Rinott, Ruty
and Lample, Guillaume
and Williams, Adina
and Bowman, Samuel R.
and Schwenk, Holger
and Stoyanov, Veselin},
title = {XNLI: Evaluating Cross-lingual Sentence Representations},
booktitle = {Proceedings of the 2018 Conference on Empirical Methods
in Natural Language Processing},
year =… See the full description on the dataset page: https://huggingface.co/datasets/haiFrHust/XNLI_Vietnamese_triplets.vietnamese_news_human_ai
Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT
This is the official dataset accompanying the paper Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT, which was accepted at ICCIES 2025 and published in Computational Intelligence in Engineering Science (Springer CCIS, vol. 2587).
You can read the paper here: Detecting AI-Generated Vietnamese News Articles with Multilingual-E5 and BERT
Citation Information… See the full description on the dataset page: https://huggingface.co/datasets/AnhNguyen2299/vietnamese_news_human_ai.vietnamese-diacritic-restoration-corpusI have downloaded it from Kaggle. I sincerely thank the author for making it available.
vietnamese-sport-newspapers-summarizationvietnamese_legal_datasetvietnamese-poemvietnamese-news-dataset
📰 Bộ dữ liệu Phân loại Chủ đề Tin tức Tiếng Việt
Đây là bộ dữ liệu nhỏ gồm các đoạn tin tức hoặc tiêu đề ngắn bằng tiếng Việt, mỗi dòng được gắn nhãn chủ đề tương ứng. Bộ dữ liệu phù hợp cho các bài toán phân loại văn bản nhiều lớp trong xử lý ngôn ngữ tự nhiên (NLP).
📂 Cấu trúc bộ dữ liệu
Mỗi dòng trong file CSV gồm hai cột:
content: Nội dung tin tức (dưới dạng tiêu đề hoặc đoạn ngắn)
label: Nhãn chủ đề thuộc một trong các nhóm sau:
giáo dục
thể thao
giải trí
công… See the full description on the dataset page: https://huggingface.co/datasets/vanhai123/vietnamese-news-dataset.
