datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-Legal-QA
Vietnamese Legal QA — Question Specificity
Phân loại độ cụ thể của câu hỏi pháp luật dân sự Việt Nam: broad (hỏi khái
quát, phải tổng hợp nhiều chế định) hay narrow (hỏi vào một tình huống / một
điều luật xác định). Dùng để định tuyến truy vấn trong hệ RAG pháp luật.
Cấu trúc
Mỗi dòng là một câu hỏi kèm vết gán nhãn. Hai dòng cùng pair_id là một cặp
đối chứng sinh từ cùng một điều luật — một broad, một narrow.
Trường
Ý nghĩa
item_id, pair_id… See the full description on the dataset page: https://huggingface.co/datasets/ThanhVu101/Vietnamese-Legal-QA.vietnamese-legal-faissVietnamese-Legal-Chat-Dataset
VLSP 2025 Vietnamese Legal Dataset
This dataset is part of the VLSP 2025 Legal SLM Challenge, designed to evaluate and train large language models on Vietnamese legal reasoning tasks.
It follows the ShareGPT conversation format, enabling supervised fine-tuning (SFT) of chat-based LLMs such as Qwen3-4B-Vietnamese-Legal-Chat.
📘 Dataset Description
Tasks Included: Multiple Choice, Natural Language Inference (NLI), and Syllogistic Legal Reasoning.
Format: Follows the… See the full description on the dataset page: https://huggingface.co/datasets/luanngo/Vietnamese-Legal-Chat-Dataset.Vietnamese_Legal_Traffic_Judge_Prediction_QA
Public Dataset — Nghị định 168/2024/NĐ-CP
Tập dữ liệu hỏi đáp pháp luật giao thông đường bộ được xây dựng từ Nghị định 168/2024/NĐ-CP về xử phạt vi phạm hành chính trong lĩnh vực giao thông đường bộ.
Tổng quan
Train
Test
Tổng
Số mẫu
1.000
200
1.200
Tỉ lệ
~83%
~17%
100%
Dữ liệu được shuffle ngẫu nhiên (seed = 42) trước khi chia để đảm bảo phân phối đồng đều giữa hai tập.
Cấu trúc mỗi mẫu
{
"id": "official_00001",
"question_type":… See the full description on the dataset page: https://huggingface.co/datasets/hdv2709/Vietnamese_Legal_Traffic_Judge_Prediction_QA.vietnamese-legal-chatbot-fine-tuning-splitsvietnamese-legal-qa-mini-300
Vietnamese Legal Q&A — SFT Dataset
A domain-specific supervised fine-tuning dataset for
Vietnamese legal question answering, built for LLM
fine-tuning and instruction tuning.
Dataset Summary
This dataset contains 300 Vietnamese legal Q&A samples
covering common areas of Vietnamese civil, criminal,
labor, and administrative law. All samples are in
Vietnamese and follow the Alpaca format.
Split
Samples
Train
250
Validation
50
Total
300… See the full description on the dataset page: https://huggingface.co/datasets/Dang-DN-VN/vietnamese-legal-qa-mini-300.vietnamese-legal-claim-verificationvietnamese-legal-dataset
