datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ViBidLQA
📘 ViBidLQA — Vietnamese Bidding Law QA Dataset
ViBidLQA is a high-quality Vietnamese legal QA dataset specifically curated from the Vietnamese Bidding Law (No. 22/2023/QH15) and its Implementation Decree (No. 24/2024/NĐ-CP). This dataset was developed to support both extractive and abstractive QA tasks in the legal domain, and serves as a robust benchmark for evaluating Vietnamese Legal QA systems.
🔍 Motivation
Although existing datasets like ALQAC have been… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViBidLQA.ViLegalMCQ
ViLegalMCQ
ViLegalMCQ is a Vietnamese legal Multiple Choice Question Answering (MCQ) dataset released alongside the ViLegalLM suite. It is synthetically generated from the ALQAC legal corpus using Qwen3-8B with human filtering, providing training data for context-based legal MCQ tasks.
Paper: ViLegalLM: Language Models for Vietnamese Legal Text — Read paper
Resources: GitHub | ViLegalBERT | ViLegalQwen2.5-1.5B-Base | ViLegalQwen3-1.7B-Base
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViLegalMCQ.ViSpanExtractQA
📘 ViSpanExtractQA — Vietnamese Span-based QA Benchmark
ViSpanExtractQA is a consolidated Vietnamese QA dataset designed for span-based extractive question answering tasks. It aggregates and harmonizes multiple high-quality resources to create a diverse, multilingual, and robust benchmark for Vietnamese QA systems.
🔍 Motivation
While several QA datasets exist in Vietnamese, most are limited in size, scope, or consistency. ViSpanExtractQA aims to bridge this gap by… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViSpanExtractQA.bimnext_math_vi_1ViLegalTF
ViLegalTF
ViLegalTF is a Vietnamese legal True/False Question Answering (TF) dataset released alongside the ViLegalLM suite. It is synthetically generated from the ALQAC legal corpus using Qwen3-8B with human filtering, providing training data for context-based legal true/false judgment tasks.
Paper: ViLegalLM: Language Models for Vietnamese Legal Text — Read paper
Resources: GitHub | ViLegalBERT | ViLegalQwen2.5-1.5B-Base | ViLegalQwen3-1.7B-Base
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViLegalTF.
