datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-News
Dataset Card for "VietnameseNewsparquet"
More Information needed
vietnamese-tokenizedTranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.vietnam-listed-companies-financial-statements
Vietnamese Listed Companies Financial Data
Overview
This dataset provides standardized financial statement data for Vietnamese listed companies.
The original financial information was collected from publicly available financial statements and annual reports published through official stock exchange portals and company websites.
The source data has been transformed into a standardized long-format Parquet structure for research, educational, analytical, and… See the full description on the dataset page: https://huggingface.co/datasets/thanhnp-uel/vietnam-listed-companies-financial-statements.Agentic-Multi-SWE-RLvietnamese_curated_dataset
Dataset Description
Vietnamese Curated Text Dataset. This dataset is collected from multiple open Vietnamese datasets, and curated with NeMo Curator
Developed by: Viettel Solutions
Language: Vietnamese
Details
Please visit our Tech Blog post on NVIDIA's plog page for details. Link
Data Collection
We utilize a combination of datasets that contain samples in Vietnamese language, ensuring a robust and representative text corpus. These datasets include:
The… See the full description on the dataset page: https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset.Vietnamese-Traditional-Musicvietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/dolly-vn/dolly-audio-1000h-vietnamese.vietnamese-corpus
Dataset Card for "vietnamese-corpus"
More Information needed
Nemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.Vietnam-Celeb
unofficial mirror of Vietnam-Celeb dataset
official announcement:
https://www.isca-archive.org/interspeech_2023/pham23b_interspeech.html
https://github.com/Vietnam-Celeb/Vietnam-Celeb
https://huggingface.co/datasets/hustep-lab/Vietnam-Celeb
official download:
Part 0: https://drive.google.com/file/d/1pMuT3DFzSwib7SVcRS8VkDwPuLTsemSG/view?usp=share_link
Part 1: https://drive.google.com/file/d/1xayHt2HRqE1aJ4HvtUT40_9XlgvfDfRY/view?usp=share_linkPart 2:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/Vietnam-Celeb.vietnamese-corpus
Vietnamese Combined Corpus
Dataset Statistics
Total documents: {<15M:,}
Wikipedia articles: {>1.3M:,}
News articles: {>13M:,}
Text documents: {>200K:,}
Processing Details
Processed using Apache Spark
Minimum document length: {10} characters
Text cleaning applied:
HTML/special character removal
Whitespace normalization
URL removal
Empty document filtering
Data Format
Each document has:
'text': The document content
'source': Origin of the… See the full description on the dataset page: https://huggingface.co/datasets/BlossomsAI/vietnamese-corpus.youtube-center-vietnamese-asrlaion2B-multi-Vietnamese-subset
Dataset Card for LAION-2B-multi Vietnamese subset
Dataset Summary
Filter the Vietnamese subset from Laion2B-multi
To get the subset of your language, check out this notebook
vietnam-real-estates
🏠 Tinix Vietnam Real Estate Listings (2025-2026)
Tinix Vietnam Real Estate Listings 2025-2026 là bộ dữ liệu bất động sản Việt Nam quy mô lớn được thu thập và xử lý bởi TiniX AI, bao gồm đúng 3.500.744 tin đăng bán/cho thuê bất động sản từ tháng 6/2025 đến tháng 3/2026 sau khi đã qua bước lọc loại hình nghiêm ngặt (loại bỏ Nhà mặt phố, Nhà trong ngõ). Đây là tài nguyên phục vụ nghiên cứu về thị trường bất động sản, xây dựng mô hình định giá nhà, phân tích xu hướng thị trường tại… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/vietnam-real-estates.mmarco-vietnamese-split
Dataset Summary
This dataset contains Vietnamese split of the mMARCO dataset.
Subset
Split
# Rows
triples
train
39,780,811
queries
train
808,731
queries
dev.full
101,093
queries
dev
6980
collection
collection
8,841,823
Note: triples contains (query, positive, negative) triples and can be used for training embeddings models.
Citing
@article{DBLP:journals/corr/abs-2108-13897,
author = {Luiz Bonifacio and
Israel Campiotti and… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/mmarco-vietnamese-split.Vietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.vietnamese-music-dataset
Vietnamese Music Dataset
A collection of 4,820 Vietnamese music tracks with matching cover thumbnails and per-track metadata collected from YouTube, packaged as an audiofolder dataset.
Repository structure
Path
Contents
Count
audio/
MP3 audio files, named by YouTube video ID
4,820
images/
PNG cover thumbnails, same IDs as audio/
4,820
data/
Parquet metadata files, one per collection session
31
Metadata schema
Each Parquet file in… See the full description on the dataset page: https://huggingface.co/datasets/Toan-Minh-Duong-Son/vietnamese-music-dataset.vietnamese-text-pretokenizedolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/fnooub/dolly-audio-1000h-vietnamese.vietnamese_curated_textsvietnamese-packed-2048vietnamese_students_feedbackStudents’ feedback is a vital resource for the interdisciplinary research involving the combining of two different
research fields between sentiment analysis and education.
Vietnamese Students’ Feedback Corpus (UIT-VSFC) is the resource consists of over 16,000 sentences which are
human-annotated with two different tasks: sentiment-based and topic-based classifications.
To assess the quality of our corpus, we measure the annotator agreements and classification evaluation on the
UIT-VSFC corpus. As a result, we obtained the inter-annotator agreement of sentiments and topics with more than over
91% and 71% respectively. In addition, we built the baseline model with the Maximum Entropy classifier and achieved
approximately 88% of the sentiment F1-score and over 84% of the topic F1-score.Vietnamese-msMARCO-ggtranslatedmerged_vietnamese_instruction_datasetdataset_vietnamesedolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/quanghd96/dolly-audio-1000h-vietnamese.Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned
This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 4 hours for 500 examples.
data-voice-vietnamese-restaurant-quan-oc
Vietnamese Restaurant Order Speech
This dataset contains Vietnamese spoken restaurant orders paired with text transcripts. Each utterance typically includes a table number, item quantities, dishes, drinks, and add-ons.
Dataset Structure
Files are split into subdirectories by filename-derived speaker_code to satisfy Hugging Face repository file-count limits:
metadata.csv: one row per audio sample.
audio/{speaker_code}/*.wav: mono WAV audio files.… See the full description on the dataset page: https://huggingface.co/datasets/EmilyNguyen235/data-voice-vietnamese-restaurant-quan-oc.
