datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.Nemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned
This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 4 hours for 500 examples.
vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
2,393 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/vietnamese-legal-documents.vietnamese-legal-instruct
Vietnamese Legal Instruction Dataset
Dataset: huggingface.co/datasets/duyet/vietnamese-legal-instruct | Source code: github.com/duyet/vietnamese-legal-documents-dataset
Instruction-following dataset built from th1nhng0/vietnamese-legal-documents — 127K Vietnamese legal documents from vbpl.vn (Government Legal Document Portal, Ministry of Justice).
467,732 training pairs across 14 QA types with deep Vietnamese legal hierarchy knowledge. Every document has a full_text pair for content… See the full description on the dataset page: https://huggingface.co/datasets/duyet/vietnamese-legal-instruct.Vietnam-Law-Raw-Datavietnamese-news-copus-segmented
Dataset Card for vietnamese-news-copus-segmented
Dataset Summary
This dataset is a refined collection of Vietnamese news articles, originally sourced from ademax/binhvq-news-corpus. It has been processed through a specialized pipeline for cleaning, normalization, and word segmentation. It is ideal for training Vietnamese Language Models (LLMs), word embeddings, or text classification tasks.
Original Source: ademax/binhvq-news-corpus
Language: Vietnamese (vi)… See the full description on the dataset page: https://huggingface.co/datasets/trungbb8/vietnamese-news-copus-segmented.Vietnamese-Legal-Documents
Vietnamese Legal Documents Dataset
1. Dataset Summary
Raw data: tmnam20/BKAI-Legal-Retrieval
The Vietnamese Legal Documents Dataset is a benchmark dataset designed for legal information retrieval in the Vietnamese language. It consists of:
A corpus of legal documents.
Train/test splits containing natural language queries and their corresponding relevant documents.
This dataset is intended to support research and development in:
Information Retrieval (IR)… See the full description on the dataset page: https://huggingface.co/datasets/YuITC/Vietnamese-Legal-Documents.KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k
This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 9 hours for 2k examples.
Usage
from datasets import load_dataset
kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1
### Dataset Summary
`magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`.
The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.vsec-vietnamese-spell-correction
VSEC: Vietnamese Spell Correction Dataset
Dataset Description
VSEC (Vietnamese Spell Correction) is a comprehensive dataset for Vietnamese spelling error detection and correction, containing 9,341 sentences with 11,202 human-made misspellings across 5,211 unique error types. This dataset represents the largest publicly available collection of Vietnamese spelling errors with syllable-level annotations, making it an invaluable resource for developing and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/vsec-vietnamese-spell-correction.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedVietnamese-nampdn-ai-tiny-webtext-gg-translatedvietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
2,393 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/vietnamese-legal-documents.vietnamese_health_dataset
Team and Homepage
Official Website: https://aienthusiasm.vn
Hugging Face Organization: https://huggingface.co/ai-enthusiasm-community
Contact
If you encounter any issues with the dataset or have any inquiries, please feel free to reach out to us via email at: aienthusiasm.team@gmail.com
Dataset Structure
The dataset is provided in a flattened tabular format, optimized for the Hugging Face Dataset Viewer and high-speed Parquet processing.… See the full description on the dataset page: https://huggingface.co/datasets/ai-enthusiasm-community/vietnamese_health_dataset.Vietnam-History-1M-ViLearn more on GitHub: https://github.com/MinhxThanh/Vietnam-History-Chat-Datasets
Thông số chính
Tổng số mẫu: 1,000,000
Tỷ lệ có reasoning (analysis): ~77.99%
Tỷ lệ chỉ trả lời (final-only): ~22.01%
Định dạng: messages theo ShareGPT/ChatML
Mẫu có reasoning: system → user → assistant (analysis) → assistant (final)
Mẫu final-only: system → user → assistant (final)
vietnamese-poetry-corpusvietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn
Language:… See the full description on the dataset page: https://huggingface.co/datasets/anhquan12/vietnamese-legal-documents.Vietnamese-SFT-Corpus-V2
🇻🇳 Vietnamese SFT Corpus V2.1 (Balanced Safety & Anti-Over-Refusal)
Vietnamese SFT Corpus V2.1 là tập dữ liệu Tinh chỉnh có Giám sát (Supervised Fine-Tuning - SFT) chuẩn công nghiệp dành cho mô hình ngôn ngữ lớn (LLM) tiếng Việt. Tập dữ liệu được thiết kế nhằm phục vụ huấn luyện trợ lý ảo thông minh, hội thoại tự nhiên, suy luận logic, đồng thời đặc trị triệt để hiện tượng "từ chối lười biếng / từ chối nhầm" (Lazy Refusal / Over-Refusal) vốn xuất hiện phổ biến ở các mô… See the full description on the dataset page: https://huggingface.co/datasets/Phuc-HugigFace/Vietnamese-SFT-Corpus-V2.vietnamese-legal-qa
thangvip/vietnamese-legal-qa
Dataset Description
This dataset contains Vietnamese legal documents with automatically generated question-answer pairs. Each document includes comprehension questions of varying difficulty levels (easy, medium, hard) and types (factual, interpretation, analytical, application).
Dataset Structure
Data Fields
doc_name: Name of the legal document
doc_type_name: Type of document (e.g., "Luật" for Law)
article_content:… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/vietnamese-legal-qa.vietnamese-sft-10k
Vietnamese Instruction-Following Dataset (10K)
This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts.
Format: JSONL (one object per line)
Fields: "prompt" (instruction or user message), "response" (assistant reply)
Language:… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-sft-10k.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn
Language:… See the full description on the dataset page: https://huggingface.co/datasets/vanhthefirst/vietnamese-legal-documents.vietnamese-roleplay-realm
🇻🇳 Vietnamese Role-play Realm Dataset
This is a dataset of GPT-generated Vietnamese characters made to increase the ability of open-source language models to role-play.
It contains 446 characters generated by GPT-3.5
Each character will have 20 topics generated by ChatGPT. And each topic will have a conversation corresponding with it
In 446 characters, there are 400 general characters and 46 Vietnamese characters.
To construct this dataset, we follow a four-step process:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vietnamese-roleplay-realm.Vietnamese-Legal-QA-RAG
Vietnamese Legal QA Dataset for RAG Evaluation (420 rows)
Dataset Summary
Vietnamese-Legal-QA-RAG is a specialized dataset designed specifically for evaluating Retrieval-Augmented Generation (RAG) systems in the Vietnamese language.
Containing 420 meticulously curated rows, this dataset focuses on the Vietnamese legal domain. It is built not just to test simple fact retrieval, but to rigorously evaluate an LLM's ability to perform multi-hop reasoning and, crucially, to… See the full description on the dataset page: https://huggingface.co/datasets/NamSyntax/Vietnamese-Legal-QA-RAG.Vietnam-History-200K-ViBộ dữ liệu 200,000 mẫu (tiếng Việt) về lịch sử Việt Nam 905–2025, đúng định dạng messages cho mô hình tư duy (analysis) và mô hình thường (final-only).
Thông tin chính
Ngôn ngữ: Tiếng Việt
Phạm vi: 905–2025 (các mốc, nhân vật, trận đánh, văn bản, cải cách, thời kỳ…)
Định dạng: messages theo ShareGPT/ChatML
Mẫu có reasoning (~78%): system → user → assistant (analysis) → assistant (final)
Mẫu final-only (~22%): system → user → assistant (final)
Tính đa dạng câu hỏi: theo năm, theo sự… See the full description on the dataset page: https://huggingface.co/datasets/minhxthanh/Vietnam-History-200K-Vi.vietnam-normalize-24kVietnamese_literature_VuTrongPhung
Vu Trong Phung Literature Chunks
This dataset consists of Vietnamese literary texts written by author Vũ Trọng Phụng, one of the most influential figures of 20th-century Vietnamese literature.
The dataset includes both short stories and novels, and has been split into smaller chunks based on the number of tokens.
Chunking Strategy
We used the tokenizer from vinai/PhoGPT-4B to split the original text into chunks of less than 512 tokens.
Each row in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/trieunh/Vietnamese_literature_VuTrongPhung.combined-vietnamese-legal-text
thangvip/combined-vietnamese-legal-text
Dataset Description
This is a combined Vietnamese legal dataset with question-answer pairs formatted in a single text column. It combines two datasets:
thangvip/vietnamese-legal-qa (9,715 examples)
thangvip/law-reading-comprehension-qa-filtered (205,369 examples)
Dataset Structure
Data Fields
text: Combined text containing legal content followed by question-answer pairs in XML-like format
Format… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/combined-vietnamese-legal-text.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
1,335 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/pdt590/vietnamese-legal-documents.vietnamese-dpo-10k
Vietnamese DPO Dataset (10K)
This dataset contains 10,000 Vietnamese prompt-response pairs in the Direct Preference Optimization (DPO) format, including a "prompt", a "chosen" response (preferred), and a "rejected" response (less preferred or misaligned). It is intended for training language models to better align with human-preferred responses, particularly in edge cases involving social sensitivity, rudeness, or toxicity.
Format: JSONL (one object per line)
Fields:
"prompt":… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-dpo-10k.
