datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ViBidLQA_v1
ViBidLQA: Vietnamese Bidding Legal Question Answering Dataset
Summary
ViBidLQA is a synthesized Vietnamese legal question-answering dataset built from the Vietnamese Bidding Law. It was created to address the scarcity of large-scale annotated datasets for legal AI in Vietnamese — a low-resource language setting. The dataset contains 3,013 QA pairs generated automatically using a large language model (Gemini) and verified by two domain experts. You may also like the new… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViBidLQA_v1.ViBidLQA
📘 ViBidLQA — Vietnamese Bidding Law QA Dataset
ViBidLQA is a high-quality Vietnamese legal QA dataset specifically curated from the Vietnamese Bidding Law (No. 22/2023/QH15) and its Implementation Decree (No. 24/2024/NĐ-CP). This dataset was developed to support both extractive and abstractive QA tasks in the legal domain, and serves as a robust benchmark for evaluating Vietnamese Legal QA systems.
🔍 Motivation
Although existing datasets like ALQAC have been… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViBidLQA.ViLegalMCQ
ViLegalMCQ
ViLegalMCQ is a Vietnamese legal Multiple Choice Question Answering (MCQ) dataset released alongside the ViLegalLM suite. It is synthetically generated from the ALQAC legal corpus using Qwen3-8B with human filtering, providing training data for context-based legal MCQ tasks.
Paper: ViLegalLM: Language Models for Vietnamese Legal Text — Read paper
Resources: GitHub | ViLegalBERT | ViLegalQwen2.5-1.5B-Base | ViLegalQwen3-1.7B-Base
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViLegalMCQ.ViLegalTF
ViLegalTF
ViLegalTF is a Vietnamese legal True/False Question Answering (TF) dataset released alongside the ViLegalLM suite. It is synthetically generated from the ALQAC legal corpus using Qwen3-8B with human filtering, providing training data for context-based legal true/false judgment tasks.
Paper: ViLegalLM: Language Models for Vietnamese Legal Text — Read paper
Resources: GitHub | ViLegalBERT | ViLegalQwen2.5-1.5B-Base | ViLegalQwen3-1.7B-Base
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViLegalTF.
