datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ViBidLQA
📘 ViBidLQA — Vietnamese Bidding Law QA Dataset
ViBidLQA is a high-quality Vietnamese legal QA dataset specifically curated from the Vietnamese Bidding Law (No. 22/2023/QH15) and its Implementation Decree (No. 24/2024/NĐ-CP). This dataset was developed to support both extractive and abstractive QA tasks in the legal domain, and serves as a robust benchmark for evaluating Vietnamese Legal QA systems.
🔍 Motivation
Although existing datasets like ALQAC have been… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViBidLQA.ViLegalMCQ
ViLegalMCQ
ViLegalMCQ is a Vietnamese legal Multiple Choice Question Answering (MCQ) dataset released alongside the ViLegalLM suite. It is synthetically generated from the ALQAC legal corpus using Qwen3-8B with human filtering, providing training data for context-based legal MCQ tasks.
Paper: ViLegalLM: Language Models for Vietnamese Legal Text — Read paper
Resources: GitHub | ViLegalBERT | ViLegalQwen2.5-1.5B-Base | ViLegalQwen3-1.7B-Base
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViLegalMCQ.ViLegalTF
ViLegalTF
ViLegalTF is a Vietnamese legal True/False Question Answering (TF) dataset released alongside the ViLegalLM suite. It is synthetically generated from the ALQAC legal corpus using Qwen3-8B with human filtering, providing training data for context-based legal true/false judgment tasks.
Paper: ViLegalLM: Language Models for Vietnamese Legal Text — Read paper
Resources: GitHub | ViLegalBERT | ViLegalQwen2.5-1.5B-Base | ViLegalQwen3-1.7B-Base
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViLegalTF.
