datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.Vietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.Vietnamese-msMARCO-ggtranslatedTranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5 over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm
Languages: Vietnamese (vi) answers · English (en) reasoning
Modality: image + text → text
Records: 544,795… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-ImageText-Reasoning.laion-2b-vietnamese-subset
Dataset Card for "laion-2b-vietnamese-subset"
More Information needed
TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5-397B-A17B
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm.. Mình rất welcome cho các hợp tác liên quan tới building… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/TranNhiem-Vietnamese-DocumentImage-Reasoning.TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm..
Languages: Vietnamese (vi) answers · English (en) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-DocumentImage-Reasoning.vietnamese-job-descriptions
💼 Tinix Vietnam Job Description
1. 📌 Giới Thiệu Tinix Vietnam Job Description
Tinix Vietnam Job Description là bộ dữ liệu tuyển dụng tiếng Việt ở định dạng CSV, gồm các tin tuyển dụng có cấu trúc về chức danh, công ty, mức lương, địa điểm, loại hợp đồng, ngành nghề, yêu cầu kinh nghiệm, trình độ học vấn, mô tả công việc, phúc lợi, yêu cầu ứng viên và năm đăng tin.
Bộ dữ liệu được thiết kế cho các bài toán NLP và phân tích thị trường lao động tại Việt Nam, đặc biệt trong… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/vietnamese-job-descriptions.vietnamese-evidence-retrieval-indexes-v2-r1
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v2-r1 at revision 2a18d35b6ea2e078db95c1aacdc2a28947268b4e.
Rows: 63,699
Source embedding shards: 13
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Dense input: title + text
Dense rows: deduplicated by content hash
Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes-v2-r1.vietnamese-healthcare-dataset
Vietnamese Healthcare Synthetic Patient Records
This dataset contains synthetic Vietnamese healthcare identity records from multiple source systems, plus a canonical synthetic patient table used by the generator.
All records are synthetic and are intended for entity resolution, record linkage, and Vietnamese identity-field preprocessing experiments.
Included Files
Only the following CSV files are included in this upload:
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/adachankawai/vietnamese-healthcare-dataset.vietnamese-evidence-retrieval-indexes
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large at revision e928944361ca7d4c80f80d19bec52ebad55a4f7f.
Rows: 52,605
Source embedding shards: 11
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Vietnamese tokenization: Unicode word tokens, no stemming and no stopword removal
row_id in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes.Vietnamese_literature_VuTrongPhung
Vu Trong Phung Literature Chunks
This dataset consists of Vietnamese literary texts written by author Vũ Trọng Phụng, one of the most influential figures of 20th-century Vietnamese literature.
The dataset includes both short stories and novels, and has been split into smaller chunks based on the number of tokens.
Chunking Strategy
We used the tokenizer from vinai/PhoGPT-4B to split the original text into chunks of less than 512 tokens.
Each row in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/trieunh/Vietnamese_literature_VuTrongPhung.hanzi-sino-vietnamese
HSK × Sino-Vietnamese (Hán-Việt) character dataset
768 HSK characters joined with their Sino-Vietnamese (Hán-Việt) readings, radical breakdowns and hand-written memory hooks in Vietnamese.
Open HSK wordlists are plentiful. The Sino-Vietnamese layer is what is missing from all of them — and it is the layer that matters most for the ~1 million Vietnamese speakers studying Chinese, because roughly 60% of Vietnamese vocabulary descends from Chinese. A learner meeting 学 (xué) already… See the full description on the dataset page: https://huggingface.co/datasets/kaihanzi/hanzi-sino-vietnamese.vietnamese-toxic-commentaio2025-vietnamese-exam
AIO2025 Vietnamese AI Exam Dataset (v3 - Normalized)
Bộ dữ liệu câu hỏi trắc nghiệm AI tiếng Việt từ Kỳ thi AI Việt Nam 2025 (AIO2025).
Dataset Summary
Split
Số câu
Train
109
Test
34
Tổng
143
Ngôn ngữ: Tiếng Việt
Chủ đề: AI, Machine Learning, Deep Learning, Computer Vision, NLP
Nguồn: AIO2025 Vietnam AI Exam
Version: v3 (normalized, verified by eval_dataset.py)
Schema (13 trường)
Trường
Mô tả
id
Mã định danh duy nhất… See the full description on the dataset page: https://huggingface.co/datasets/vudang449/aio2025-vietnamese-exam.banking_sentiment_vietnameseVietnamese-Openorca-Multiplechoice-gg-translatedvietnamese-evidence-corpus-chunked-e5-v3
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
Chunked with multilingual-E5 token budget
Prefix-aware chunking using `passage: {title}
`
Sentence-aware overlap to preserve local context
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.Eval-RAG-Vietnamesevietnamese-legal-corpus-20k-rawmedical_vietnamese_datasetsvietnamese-caucu-comments
Vietnamese Cau Cuu Facebook Comments
Dataset Summary
This dataset contains Vietnamese Facebook comments collected from a natural-disaster discussion thread and auto-labeled for binary emergency detection.
The target task is to detect whether a comment is a real-time rescue request (cau_cuu) versus a non-emergency comment (khong_phai_cau_cuu).
This release is intended as a bootstrap dataset for triage modeling and should be treated as a weakly supervised resource. Human… See the full description on the dataset page: https://huggingface.co/datasets/dat201204/vietnamese-caucu-comments.wit_vietnamese_subset
Dataset Card for "wit_vietnamese_subset"
More Information needed
vietnamese-restaurant-review-sentiment-datasetvietnamese_10classes_train_test_splitvietnamese-evidence-corpus-chunked
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.vietnamese-evidence-corpus-embeddings-e5-large-v2vietnamese_toxic_corevietnamese-evidence-corpus-chunked-e5-v2
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.slam-en-es-vietnamese-prompts
SLAM English-Spanish Knowledge Tracing with Vietnamese Prompts
This dataset is a derived version of bihungba1101/slam-en-es-knowledge-tracing. It preserves every source row and field and adds prompt_vi, a machine-generated Vietnamese translation of the Spanish prompt.
The source track contains English learners who already speak Spanish. Consequently:
prompt is the original Spanish text shown to the learner when text is available.
prompt_vi is a Vietnamese translation of prompt.… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/slam-en-es-vietnamese-prompts.
