CoolFace
Datasetpublic

th1nhng0/vietnamese-legal-documents

Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
45likes1.1kdownloads
Dataset Card

Vietnamese Legal Documents

A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).

2026 Portal Refresh

This release migrates the active configs to the current VBPL Next.js catalog and public JSON gateway. It adds UUID and portal-prefixed records, removes duplicate content rows, refreshes document text and relationships, and retains records that disappeared from the live catalog as historical rows.

Migration note: metadata.id, content.id, relationships.doc_id, and relationships.other_doc_id are now strings. Consumers that previously joined on integer IDs must cast their keys to string before upgrading. The former mixed-schema legacy config is now exposed as the loadable legacy_metadata and legacy_content configs.

See CHANGELOG.md for the complete release summary.

Quick Start

python
from datasets import load_dataset

# Metadata for 171k documents
meta = load_dataset("th1nhng0/vietnamese-legal-documents", "metadata", split="data")
print(meta.to_pandas().head())

# Cross-document relationships (amendments, citations, repeals, …)
rels = load_dataset("th1nhng0/vietnamese-legal-documents", "relationships", split="data")
print(rels.to_pandas().head())

# Full-text HTML content for 170k documents
content = load_dataset("th1nhng0/vietnamese-legal-documents", "content", split="data")
print(content.to_pandas().head())

Join the two on id (metadata) ↔ doc_id (relationships):

python
import pandas as pd

df = meta.to_pandas()
rel = rels.to_pandas()

# Find all documents that cite document 10420
citing = rel[rel["other_doc_id"] == "10420"].merge(df, left_on="doc_id", right_on="id")
print(citing[["id", "title", "relationship"]])

Dataset Structure

The dataset has five configs:

ConfigSplitRowsDescription
metadatadata171,556One row per document — 17 metadata fields
contentdata170,824One unique raw HTML body per document
relationshipsdata1,033,255Unique directed edges between documents
legacy_metadatametadata518,601Older crawl — English field names, more docs
legacy_contentcontent518,235Plain-text content for older crawl

metadata

ColumnDescription
idUnique document ID (string; numeric, UUID, or portal-prefixed)
titleFull Vietnamese title
so_ky_hieuOfficial number, e.g. 115/NQ-HĐBCQG
ngay_ban_hanhIssuance date (DD/MM/YYYY)
loai_van_banType — Quyết định, Nghị quyết, Thông tư, …
ngay_co_hieu_lucEffective date
ngay_het_hieu_lucExpiry date (empty if still in effect)
nguon_thu_thapCollection source (e.g. Công báo)
ngay_dang_cong_baoOfficial Gazette publication date
nganhSector — Tài chính, Y tế, …
linh_vucLegal field / sub-domain
co_quan_ban_hanhIssuing authority (551 unique bodies)
chuc_danhSignatory title — Chủ tịch, Bộ trưởng, …
nguoi_kySignatory name
pham_viGeographical scope
thong_tin_ap_dungImplementation note
tinh_trang_hieu_lucEffect status — Còn hiệu lực, Hết hiệu lực toàn bộ, …

content

ColumnDescription
idDocument ID (join key → metadata.id)
content_htmlRaw HTML body of the document
Note: 732 metadata rows do not have a corresponding content row. The portal does not expose an HTML body for those records (some are PDF-only).

relationships

ColumnDescription
doc_idSource document ID (join key → metadata.id)
other_doc_idTarget document ID
relationshipEdge label from the current portal; archived old-only pairs retain their historical label

Every doc_id is present in metadata. Relationship targets may reference documents outside the live catalog; 17,706 distinct other_doc_id values do not have a metadata row in this snapshot.

Legacy configs

An older, larger crawl snapshot with ~518 k documents. Field names and enumerated values are in English (unlike the current configs which use Vietnamese originals). Dates are YYYY-MM-DD. Use this config when you need broader coverage at the cost of reduced metadata richness.

`legacy` / `metadata` split (518,601 rows):

ColumnDescription
idUnique document ID (int)
document_numberOfficial number, e.g. 115/NQ-HĐBCQG
titleFull Vietnamese title
legal_typeDocument type in English — Resolution, Decision, Circular, …
legal_sectorsSector in English — Finance, Education, …
issuing_authorityIssuing authority (Vietnamese name)
issuance_dateIssuance date (YYYY-MM-DD)
effect_dateEffective date (YYYY-MM-DD)
effectless_dateExpiry date (empty if still in effect)
effect_statusIn effect or Not in effect
signersSignatory name and ID, e.g. Trần Thanh Mẫn:2140

`legacy` / `content` split (518,235 rows):

ColumnDescription
idDocument ID (join key → legacy/metadata.id)
contentPlain-text body of the document
Note: The content column in legacy contains plain text, not HTML. The current content config stores raw HTML (content_html).
python
from datasets import load_dataset

legacy_meta = load_dataset("th1nhng0/vietnamese-legal-documents", "legacy_metadata", split="metadata")
legacy_content = load_dataset("th1nhng0/vietnamese-legal-documents", "legacy_content", split="content")

Data Collection

All data was scraped from vbpl.vn using the Scrapy crawler under `crawler/`. The single vbpl spider targets the current Next.js catalog and public JSON detail gateway, collecting metadata, HTML content, signers, fields, and relationships in one pass. IDs are strings because the migrated portal uses numeric IDs, UUIDs, and prefixed identifiers.

The current files were refreshed on 2026-07-23. Refresh rows are authoritative for overlapping IDs; missing refresh fields are filled from the prior dataset, and records no longer exposed by the live catalog are retained as historical rows.

For local rollback, the immediately preceding files are preserved under the Git-ignored data/archive_pre_refresh_2026-07-23/ directory.

bash
cd crawler
uv sync --extra validation

uv run scrapy crawl vbpl -a seed_file=data/ids.txt
uv run scrapy crawl vbpl -a seed_file=data/ids.txt -a proxy_file=proxies.txt
uv run scrapy crawl vbpl -a seed_file=data/ids.txt -a resume=1 -a resume_from=../data/metadata_raw.jsonl

# Crawl the complete catalog. This one pass includes metadata, relationships,
# and full-text HTML.
uv run scrapy crawl vbpl -a full=1 -a output=../data/metadata_raw.jsonl

# Safely continue an interrupted full crawl.
uv run scrapy crawl vbpl -a full=1 -a resume=1 -a resume_from=../data/metadata_raw.jsonl -a output=../data/metadata_raw.jsonl

# Stream the combined crawl into separate, versioned Parquet files.
uv run python build_full_dataset.py ../data/metadata_raw.jsonl ../data/refresh_YYYY-MM-DD

# Upsert a verified refresh into a staging directory. Refresh values win,
# missing values are recovered from the current data, and old-only rows remain.
uv run python upsert_dataset.py ../data/refresh_YYYY-MM-DD ../data/upsert_YYYY-MM-DD

# Validate card metadata, Parquet schemas/counts, uniqueness, and join integrity.
uv run --extra validation python validate_release.py ..

Both processing commands refuse to overwrite their output directory. Validate the staged Parquet files before promoting them into data/.

Limitations

  • —Coverage depends on what vbpl.vn has indexed; older or undigitized documents may be missing.
  • —The live catalog can change during a long crawl. Use -a resume=1 for a final catalog-only verification pass; zero newly scraped items means every currently listed ID is already present.
  • —Effect status reflects the portal at crawl time and may lag behind real-world changes.
  • —This is a snapshot, not a live mirror. Always cross-check with the portal for authoritative status.

Privacy

The dataset contains names of document signatories (public officials acting in their official capacity). No private citizen data is included.

Citation

bibtex
@dataset{ngo_thinh_2026_vietnamese_legal,
  title     = {Vietnamese Legal Documents},
  author    = {Thịnh Ngô},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents},
}

License

Vietnamese legal documents are public domain under the Law on Access to Information (No. 104/2016/QH13) and the Law on Promulgation of Legal Documents (No. 64/2025/QH15).

The compiled dataset (schema, processing, curation) is released under CC BY 4.0. Not a substitute for legal advice.