datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned
This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 4 hours for 500 examples.
vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
2,393 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/vietnamese-legal-documents.vietnamese-legal-instruct
Vietnamese Legal Instruction Dataset
Dataset: huggingface.co/datasets/duyet/vietnamese-legal-instruct | Source code: github.com/duyet/vietnamese-legal-documents-dataset
Instruction-following dataset built from th1nhng0/vietnamese-legal-documents — 127K Vietnamese legal documents from vbpl.vn (Government Legal Document Portal, Ministry of Justice).
467,732 training pairs across 14 QA types with deep Vietnamese legal hierarchy knowledge. Every document has a full_text pair for content… See the full description on the dataset page: https://huggingface.co/datasets/duyet/vietnamese-legal-instruct.Vietnamese-Multi-turn-Chat-AlpacaVietnamese-Legal-Documents
Vietnamese Legal Documents Dataset
1. Dataset Summary
Raw data: tmnam20/BKAI-Legal-Retrieval
The Vietnamese Legal Documents Dataset is a benchmark dataset designed for legal information retrieval in the Vietnamese language. It consists of:
A corpus of legal documents.
Train/test splits containing natural language queries and their corresponding relevant documents.
This dataset is intended to support research and development in:
Information Retrieval (IR)… See the full description on the dataset page: https://huggingface.co/datasets/YuITC/Vietnamese-Legal-Documents.vietnamese-medical-qa
Dataset Summary
Vietnamese-Medical-QA is a question-answering dataset in the healthcare domain, collected from edoctor and vinmec.
Size: After merging data from these two sources, obtained 9335 QA pairs.
Language: Vietnamese
Load with Datasets
from datasets import load_dataset
# Load dataset from huggingface
qa_dataset = load_dataset("hungnm/vietnamese-medical-qa")
# print a QA example
print(qa_dataset['train'][0])
{
"question": "Chào bác sĩ,\nRăng cháu hiện tại… See the full description on the dataset page: https://huggingface.co/datasets/hungnm/vietnamese-medical-qa.KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k
This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 9 hours for 2k examples.
Usage
from datasets import load_dataset
kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.Vietnamese-Locutusque-function-calling-chatml-gg-translatedVietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1
### Dataset Summary
`magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`.
The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.Vietnamese-alpaca-gpt4-gg-translatedVietnamese-Salesforce-xlam-function-calling-60k-gg-translatedVietnamese-nampdn-ai-tiny-webtext-gg-translatedvietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
2,393 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/vietnamese-legal-documents.vietnamese_health_dataset
Team and Homepage
Official Website: https://aienthusiasm.vn
Hugging Face Organization: https://huggingface.co/ai-enthusiasm-community
Contact
If you encounter any issues with the dataset or have any inquiries, please feel free to reach out to us via email at: aienthusiasm.team@gmail.com
Dataset Structure
The dataset is provided in a flattened tabular format, optimized for the Hugging Face Dataset Viewer and high-speed Parquet processing.… See the full description on the dataset page: https://huggingface.co/datasets/ai-enthusiasm-community/vietnamese_health_dataset.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn
Language:… See the full description on the dataset page: https://huggingface.co/datasets/anhquan12/vietnamese-legal-documents.Vietnamese-395k-meta-math-MetaMathQA-gg-translatedVietnamese-ShareGPT4Video-ShareGPT4Video-gg-translatedVietnamese-ShareGPT4Vision-gg-translatedVietnamese-TriviaQA-RC-gg-translatedVietnamese-liuhaotian-llava_v1_5_mix665k-gg-translatedVietnamese-ComplexWebQuestions-gg-translatedvietnamese-legal-qa
thangvip/vietnamese-legal-qa
Dataset Description
This dataset contains Vietnamese legal documents with automatically generated question-answer pairs. Each document includes comprehension questions of varying difficulty levels (easy, medium, hard) and types (factual, interpretation, analytical, application).
Dataset Structure
Data Fields
doc_name: Name of the legal document
doc_type_name: Type of document (e.g., "Luật" for Law)
article_content:… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/vietnamese-legal-qa.vietnamese-sft-10k
Vietnamese Instruction-Following Dataset (10K)
This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts.
Format: JSONL (one object per line)
Fields: "prompt" (instruction or user message), "response" (assistant reply)
Language:… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-sft-10k.vietnamese-vlm
Vietnamese Industries Insights
About Me
I'm Matteo Khan, a computer science apprentice at TW3 Partners, specializing in Generative AI and NLP. My focus is on creating datasets that improve AI's ability to process complex technical documents.
You can connect with me on LinkedIn: Matteo Khan
Dataset Details
Purpose / Mục Đích
Tiếng Việt:
Bộ dữ liệu này được tạo ra nhằm cung cấp cái nhìn tổng quan về các ngành công nghiệp chủ chốt… See the full description on the dataset page: https://huggingface.co/datasets/MatteoKhan/vietnamese-vlm.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn
Language:… See the full description on the dataset page: https://huggingface.co/datasets/vanhthefirst/vietnamese-legal-documents.Vietnamese-Legal-QA-RAG
Vietnamese Legal QA Dataset for RAG Evaluation (420 rows)
Dataset Summary
Vietnamese-Legal-QA-RAG is a specialized dataset designed specifically for evaluating Retrieval-Augmented Generation (RAG) systems in the Vietnamese language.
Containing 420 meticulously curated rows, this dataset focuses on the Vietnamese legal domain. It is built not just to test simple fact retrieval, but to rigorously evaluate an LLM's ability to perform multi-hop reasoning and, crucially, to… See the full description on the dataset page: https://huggingface.co/datasets/NamSyntax/Vietnamese-Legal-QA-RAG.combined-vietnamese-legal-text
thangvip/combined-vietnamese-legal-text
Dataset Description
This is a combined Vietnamese legal dataset with question-answer pairs formatted in a single text column. It combines two datasets:
thangvip/vietnamese-legal-qa (9,715 examples)
thangvip/law-reading-comprehension-qa-filtered (205,369 examples)
Dataset Structure
Data Fields
text: Combined text containing legal content followed by question-answer pairs in XML-like format
Format… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/combined-vietnamese-legal-text.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
1,335 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/pdt590/vietnamese-legal-documents.Vietnamese-OpenGVLab-ShareGPT-4o-gg-translated
