anhnon/vietnamese-corporate-legal-articles-fsm
Lexora Knowledge - Vietnamese Legal Documents Dataset Summary A structured Vietnamese legal knowledge base crawled from vbpl.vn (CSDL Quốc gia về Pháp luật - Vietnam's National Legal Database), published as 4 linked subsets: full documents, individual articles (Điều), the citation graph between documents/articles, and domain-concept tags. Load a specific subset with datasets.load_dataset("anhnon/vietnamese-corporate-legal-articles-fsm", "articles") etc. Intended… See the full description on the dataset page: https://huggingface.co/datasets/anhnon/vietnamese-corporate-legal-articles-fsm.
Lexora Knowledge - Vietnamese Legal Documents
Dataset Summary
A structured Vietnamese legal knowledge base crawled from vbpl.vn (CSDL Quốc gia về Pháp luật - Vietnam's National Legal Database), published as 4 linked subsets: full documents, individual articles (Điều), the citation graph between documents/articles, and domain-concept tags. Load a specific subset with datasets.load_dataset("anhnon/vietnamese-corporate-legal-articles-fsm", "articles") etc.
Intended for legal text corpora, retrieval-augmented generation (RAG) over Vietnamese law, citation-graph analysis, and fine-tuning language models on legal Vietnamese.
Supported Tasks
- Text retrieval / RAG:
documents.full_textfor whole-document context, orarticles.full_textfor precise, article-level passages (avoids retrieving an entire 100+ page law for one relevant Điều). - Text classification:
doc_type,effective_status,agency(documents subset) are usable as classification targets. - Text generation:
full_textfields are legal-domain Vietnamese suitable for domain-adaptive language model pretraining or fine-tuning. - Citation graph / knowledge graph tasks:
relationshipsencodes typed edges (CITES, AMENDS, REPLACES, BASEDON, GUIDEDBY, CONSOLIDATES, REPEALS...) between documents and articles - usable for relation classification, link prediction, or building a legal knowledge graph (this is the same graph the Lexora Knowledge pipeline loads into Neo4j). - Concept tagging / filtering:
entity_mentionstags each document with the domain concepts it discusses (Luật, Doanh nghiệp, Cơ quan, Thuế, Người lao động), usable to filter or weakly-label documents by topic.
Languages
Vietnamese (vi).
Dataset Structure
This dataset has 4 configs (subsets), sharing ids that let you join across them (articles.document_id → documents.id, relationships.source_id/target_id → documents.id or articles.id).
documents (240 rows)
One row per legal document (Luật, Nghị định, Thông tư...). full_text concatenates the document's chapters/sections/articles in order.
{
"id": "142847",
"title": "Luật Doanh nghiệp số 59/2020/QH14",
"doc_number": "59/2020/QH14",
"doc_type": "Luật",
"issue_date": "2020-06-17T00:00:00",
"effective_date": "2021-01-01T00:00:00",
"expiry_date": null,
"effective_status": "Còn hiệu lực",
"agency": "Quốc hội",
"signer": "Nguyễn Thị Kim Ngân",
"public_url": "https://vbpl.vn/van-ban/chi-tiet/luat-doanh-nghiep-so-59-2020-qh14--142847",
"legal_bases": ["Căn cứ Hiến pháp nước Cộng hòa xã hội chủ nghĩa Việt Nam"],
"full_text": "Điều 1. Phạm vi điều chỉnh\n..."
}articles (19449 rows)
One row per Điều (article), with its owning document/chapter/section context attached - the fine-grained retrieval unit.
{
"id": "142847_art_1",
"document_id": "142847",
"document_title": "Luật Doanh nghiệp số 59/2020/QH14",
"document_number": "59/2020/QH14",
"chapter_number": "I",
"chapter_title": "NHỮNG QUY ĐỊNH CHUNG",
"section_number": null,
"section_title": null,
"article_number": "1",
"article_title": "Phạm vi điều chỉnh",
"full_text": "Điều 1. Phạm vi điều chỉnh\n1. Luật này quy định...",
"public_url": "https://vbpl.vn/van-ban/chi-tiet/luat-doanh-nghiep-so-59-2020-qh14--142847"
}relationships (120259 rows)
The legal citation graph.
{
"source_id": "152951",
"source_label": "Document",
"source_title": "Luật sửa đổi, bổ sung một số điều của...",
"target_id": "142847",
"target_label": "Document",
"target_title": "Luật Doanh nghiệp số 59/2020/QH14",
"relation_type": "AMENDS",
"label_vn": "Văn bản sửa đổi bổ sung",
"context": "Điều 21 Luật Doanh nghiệp"
}entity_mentions (893 rows)
Document-level domain concept tags (Luật, Doanh nghiệp, Cơ quan, Thuế, Người lao động), keyword-extracted (no LLM/NLP).
{
"document_id": "142847",
"document_title": "Luật Doanh nghiệp số 59/2020/QH14",
"concept": "Doanh nghiệp",
"category": "DOANH_NGHIEP",
"mention_count": 214
}Data Splits
Each config is a single split, no train/test division - this is a corpus and knowledge graph, not a supervised benchmark.
Dataset Creation
Source Data
Crawled via BFS from a seed set of business/enterprise-law documents, following each document's official cross-references (references field) up to a configurable depth. Raw HTML content is cleaned and parsed into a physical structure (Chapter → Section → Article → Clause → Point) using rule-based Vietnamese legal-text patterns (no LLM/NLP).
Annotations
legal_bases, structural boundaries, citation relationships, and entity_mentions are all extracted with regex/rule-based parsing over the document's own text; no manual annotation or LLM was used.
Considerations for Using the Data
- Coverage: this release covers 240 documents from a depth-limited crawl seeded around Vietnamese enterprise/business law - not a complete mirror of vbpl.vn.
- Parsing limitations: quoted amendment text and appendices are excluded from structural parsing on a best-effort basis; rare edge cases in complex omnibus amendment laws may still be imperfect.
- Not legal advice:
effective_statusreflects vbpl.vn's status at crawl time and may be outdated - always verify against the official source (public_url) before relying on this data for legal decisions.
Licensing Information
The dataset compilation (structuring, entity tagging, metadata, and the citation graph) is licensed under CC-BY-4.0. The underlying legal texts themselves are not subject to copyright under Vietnamese law (Điều 15, Luật Sở hữu trí tuệ - văn bản quy phạm pháp luật is excluded from copyright protection).
Additional Information
- Homepage: https://lexora-knowledge.trunganh.tech/
- Repository: https://github.com/LexoraAI/Lexora-Knowledge
Generated from release v1.0 (240 documents, 19449 articles, 120259 relationships, 893 entity mentions) by the Lexora Knowledge pipeline (pipelines/publish_pipeline.py).
