datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.vietnam-real-estates
🏠 Tinix Vietnam Real Estate Listings (2025-2026)
Tinix Vietnam Real Estate Listings 2025-2026 là bộ dữ liệu bất động sản Việt Nam quy mô lớn được thu thập và xử lý bởi TiniX AI, bao gồm đúng 3.500.744 tin đăng bán/cho thuê bất động sản từ tháng 6/2025 đến tháng 3/2026 sau khi đã qua bước lọc loại hình nghiêm ngặt (loại bỏ Nhà mặt phố, Nhà trong ngõ). Đây là tài nguyên phục vụ nghiên cứu về thị trường bất động sản, xây dựng mô hình định giá nhà, phân tích xu hướng thị trường tại… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/vietnam-real-estates.Vietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.Vietnamese-msMARCO-ggtranslatedTranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5 over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm
Languages: Vietnamese (vi) answers · English (en) reasoning
Modality: image + text → text
Records: 544,795… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-ImageText-Reasoning.laion-2b-vietnamese-subset
Dataset Card for "laion-2b-vietnamese-subset"
More Information needed
ipfs_vietnam_laws_ir
Vietnam legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_vietnam_laws (revision 5088d10cf67c8db7ebb77bfc0e46615422512fd2) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Vietnam prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_vietnam_laws_ir.TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5-397B-A17B
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm.. Mình rất welcome cho các hợp tác liên quan tới building… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/TranNhiem-Vietnamese-DocumentImage-Reasoning.Traffic-sign-detection-VietNam
Vietnam Traffic Sign Detection Dataset
This repository contains the dataset for detecting road traffic signs in Vietnam using the state-of-the-art YOLO object detection model.
📂 Repository Structure
The dataset is structured in the standard YOLO format, containing images and corresponding annotations divided into training, validation, and testing sets.
├── classid.xlsx # Excel file mapping class IDs to names
├── dataset/
│ ├── train/ #… See the full description on the dataset page: https://huggingface.co/datasets/star092304/Traffic-sign-detection-VietNam.vietnam-real-estates-2
🏠 Tinix Vietnam Real Estate Listings (2025-2026)
Tinix Vietnam Real Estate Listings 2025-2026 là bộ dữ liệu bất động sản Việt Nam quy mô lớn được thu thập và xử lý bởi TiniX AI, bao gồm đúng 3.500.744 tin đăng bán/cho thuê bất động sản từ tháng 6/2025 đến tháng 3/2026 sau khi đã qua bước lọc loại hình nghiêm ngặt (loại bỏ Nhà mặt phố, Nhà trong ngõ). Đây là tài nguyên phục vụ nghiên cứu về thị trường bất động sản, xây dựng mô hình định giá nhà, phân tích xu hướng thị trường tại… See the full description on the dataset page: https://huggingface.co/datasets/vduydong/vietnam-real-estates-2.TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm..
Languages: Vietnamese (vi) answers · English (en) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-DocumentImage-Reasoning.vietnamese-job-descriptions
💼 Tinix Vietnam Job Description
1. 📌 Giới Thiệu Tinix Vietnam Job Description
Tinix Vietnam Job Description là bộ dữ liệu tuyển dụng tiếng Việt ở định dạng CSV, gồm các tin tuyển dụng có cấu trúc về chức danh, công ty, mức lương, địa điểm, loại hợp đồng, ngành nghề, yêu cầu kinh nghiệm, trình độ học vấn, mô tả công việc, phúc lợi, yêu cầu ứng viên và năm đăng tin.
Bộ dữ liệu được thiết kế cho các bài toán NLP và phân tích thị trường lao động tại Việt Nam, đặc biệt trong… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/vietnamese-job-descriptions.vietnam-real-estates
🏠 Tinix Vietnam Real Estate Listings (2025)
Tinix Vietnam Real Estate Listings 2025 là bộ dữ liệu bất động sản Việt Nam quy mô lớn được thu thập và xử lý bởi TiniX AI, bao gồm đúng 1.000.000 tin đăng bán/cho thuê bất động sản trong năm 2025 sau khi đã qua bước lọc loại hình. Đây là tài nguyên phục vụ nghiên cứu về thị trường bất động sản, xây dựng mô hình định giá nhà, phân tích xu hướng thị trường, và các ứng dụng địa lý không gian (GIS) tại Việt Nam.
A refined Vietnam real… See the full description on the dataset page: https://huggingface.co/datasets/vduydong/vietnam-real-estates.Traffic-sign-detection-VietNam
Vietnam Traffic Sign Detection Dataset
This repository contains the dataset for detecting road traffic signs in Vietnam using the state-of-the-art YOLO object detection model.
📂 Repository Structure
The dataset is structured in the standard YOLO format, containing images and corresponding annotations divided into training, validation, and testing sets.
├── classid.xlsx # Excel file mapping class IDs to names
├── dataset/
│ ├── train/ #… See the full description on the dataset page: https://huggingface.co/datasets/Minh124689/Traffic-sign-detection-VietNam.Miriad-Tooluse-Prompts-StratifiedKFold-View-Patch-1vietnamese-toxic-commentvietnamese-healthcare-dataset
Vietnamese Healthcare Synthetic Patient Records
This dataset contains synthetic Vietnamese healthcare identity records from multiple source systems, plus a canonical synthetic patient table used by the generator.
All records are synthetic and are intended for entity resolution, record linkage, and Vietnamese identity-field preprocessing experiments.
Included Files
Only the following CSV files are included in this upload:
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/adachankawai/vietnamese-healthcare-dataset.vietnamese-evidence-retrieval-indexes
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large at revision e928944361ca7d4c80f80d19bec52ebad55a4f7f.
Rows: 52,605
Source embedding shards: 11
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Vietnamese tokenization: Unicode word tokens, no stemming and no stopword removal
row_id in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes.Vietnamese_literature_VuTrongPhung
Vu Trong Phung Literature Chunks
This dataset consists of Vietnamese literary texts written by author Vũ Trọng Phụng, one of the most influential figures of 20th-century Vietnamese literature.
The dataset includes both short stories and novels, and has been split into smaller chunks based on the number of tokens.
Chunking Strategy
We used the tokenizer from vinai/PhoGPT-4B to split the original text into chunks of less than 512 tokens.
Each row in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/trieunh/Vietnamese_literature_VuTrongPhung.hanzi-sino-vietnamese
HSK × Sino-Vietnamese (Hán-Việt) character dataset
768 HSK characters joined with their Sino-Vietnamese (Hán-Việt) readings, radical breakdowns and hand-written memory hooks in Vietnamese.
Open HSK wordlists are plentiful. The Sino-Vietnamese layer is what is missing from all of them — and it is the layer that matters most for the ~1 million Vietnamese speakers studying Chinese, because roughly 60% of Vietnamese vocabulary descends from Chinese. A learner meeting 学 (xué) already… See the full description on the dataset page: https://huggingface.co/datasets/kaihanzi/hanzi-sino-vietnamese.vietnamese-evidence-retrieval-indexes-v2-r1
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v2-r1 at revision 2a18d35b6ea2e078db95c1aacdc2a28947268b4e.
Rows: 63,699
Source embedding shards: 13
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Dense input: title + text
Dense rows: deduplicated by content hash
Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes-v2-r1.banking_sentiment_vietnameseaio2025-vietnamese-exam
AIO2025 Vietnamese AI Exam Dataset (v3 - Normalized)
Bộ dữ liệu câu hỏi trắc nghiệm AI tiếng Việt từ Kỳ thi AI Việt Nam 2025 (AIO2025).
Dataset Summary
Split
Số câu
Train
109
Test
34
Tổng
143
Ngôn ngữ: Tiếng Việt
Chủ đề: AI, Machine Learning, Deep Learning, Computer Vision, NLP
Nguồn: AIO2025 Vietnam AI Exam
Version: v3 (normalized, verified by eval_dataset.py)
Schema (13 trường)
Trường
Mô tả
id
Mã định danh duy nhất… See the full description on the dataset page: https://huggingface.co/datasets/vudang449/aio2025-vietnamese-exam.vietnam-real-estates
🏠 Tinix Vietnam Real Estate Listings (2025-2026)
Tinix Vietnam Real Estate Listings 2025-2026 là bộ dữ liệu bất động sản Việt Nam quy mô lớn được thu thập và xử lý bởi TiniX AI, bao gồm đúng 3.500.744 tin đăng bán/cho thuê bất động sản từ tháng 6/2025 đến tháng 3/2026 sau khi đã qua bước lọc loại hình nghiêm ngặt (loại bỏ Nhà mặt phố, Nhà trong ngõ). Đây là tài nguyên phục vụ nghiên cứu về thị trường bất động sản, xây dựng mô hình định giá nhà, phân tích xu hướng thị trường tại… See the full description on the dataset page: https://huggingface.co/datasets/oceanNG/vietnam-real-estates.Vietnamese-Openorca-Multiplechoice-gg-translatedEval-RAG-Vietnamesevietnamese-evidence-retrieval-indexes-v3-1
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
aiMy144/vietnamese-evidence-corpus-embeddings-e5-large-v3-1 at revision ea0826c0a9d4273eec267eb66b6dbbfe64d5fb11.
Rows: 53,747
Source embedding shards: 11
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Dense input: title + text
Dense rows: deduplicated by content hash
Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/vietnamese-evidence-retrieval-indexes-v3-1.vietnamese-legal-corpus-20k-rawvietnamese-caucu-comments
Vietnamese Cau Cuu Facebook Comments
Dataset Summary
This dataset contains Vietnamese Facebook comments collected from a natural-disaster discussion thread and auto-labeled for binary emergency detection.
The target task is to detect whether a comment is a real-time rescue request (cau_cuu) versus a non-emergency comment (khong_phai_cau_cuu).
This release is intended as a bootstrap dataset for triage modeling and should be treated as a weakly supervised resource. Human… See the full description on the dataset page: https://huggingface.co/datasets/dat201204/vietnamese-caucu-comments.medical_vietnamese_datasets
