CoolFace
Datasetpublic

anhnon/vietnamese-corporate-legal-articles-fsm

Lexora Knowledge - Vietnamese Legal Documents Dataset Summary A structured Vietnamese legal knowledge base crawled from vbpl.vn (CSDL Quốc gia về Pháp luật - Vietnam's National Legal Database), published as 4 linked subsets: full documents, individual articles (Điều), the citation graph between documents/articles, and domain-concept tags. Load a specific subset with datasets.load_dataset("anhnon/vietnamese-corporate-legal-articles-fsm", "articles") etc. Intended… See the full description on the dataset page: https://huggingface.co/datasets/anhnon/vietnamese-corporate-legal-articles-fsm.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes28downloads
Dataset Card

Lexora Knowledge - Vietnamese Legal Documents

Dataset Summary

A structured Vietnamese legal knowledge base crawled from vbpl.vn (CSDL Quốc gia về Pháp luật - Vietnam's National Legal Database), published as 4 linked subsets: full documents, individual articles (Điều), the citation graph between documents/articles, and domain-concept tags. Load a specific subset with datasets.load_dataset("anhnon/vietnamese-corporate-legal-articles-fsm", "articles") etc.

Intended for legal text corpora, retrieval-augmented generation (RAG) over Vietnamese law, citation-graph analysis, and fine-tuning language models on legal Vietnamese.

Supported Tasks

  • —Text retrieval / RAG: documents.full_text for whole-document context, or articles.full_text for precise, article-level passages (avoids retrieving an entire 100+ page law for one relevant Điều).
  • —Text classification: doc_type, effective_status, agency (documents subset) are usable as classification targets.
  • —Text generation: full_text fields are legal-domain Vietnamese suitable for domain-adaptive language model pretraining or fine-tuning.
  • —Citation graph / knowledge graph tasks: relationships encodes typed edges (CITES, AMENDS, REPLACES, BASEDON, GUIDEDBY, CONSOLIDATES, REPEALS...) between documents and articles - usable for relation classification, link prediction, or building a legal knowledge graph (this is the same graph the Lexora Knowledge pipeline loads into Neo4j).
  • —Concept tagging / filtering: entity_mentions tags each document with the domain concepts it discusses (Luật, Doanh nghiệp, Cơ quan, Thuế, Người lao động), usable to filter or weakly-label documents by topic.

Languages

Vietnamese (vi).

Dataset Structure

This dataset has 4 configs (subsets), sharing ids that let you join across them (articles.document_id → documents.id, relationships.source_id/target_id → documents.id or articles.id).

documents (240 rows)

One row per legal document (Luật, Nghị định, Thông tư...). full_text concatenates the document's chapters/sections/articles in order.

json
{
  "id": "142847",
  "title": "Luật Doanh nghiệp số 59/2020/QH14",
  "doc_number": "59/2020/QH14",
  "doc_type": "Luật",
  "issue_date": "2020-06-17T00:00:00",
  "effective_date": "2021-01-01T00:00:00",
  "expiry_date": null,
  "effective_status": "Còn hiệu lực",
  "agency": "Quốc hội",
  "signer": "Nguyễn Thị Kim Ngân",
  "public_url": "https://vbpl.vn/van-ban/chi-tiet/luat-doanh-nghiep-so-59-2020-qh14--142847",
  "legal_bases": ["Căn cứ Hiến pháp nước Cộng hòa xã hội chủ nghĩa Việt Nam"],
  "full_text": "Điều 1. Phạm vi điều chỉnh\n..."
}
FieldTypeDescription
idstringSource document id on vbpl.vn
titlestringDocument title
doc_numberstringOfficial document number (e.g. 59/2020/QH14)
doc_typestringDocument category (Luật, Nghị định, Thông tư...)
issue_datestring \nullISO date the document was issued
effective_datestring \nullISO date the document takes effect
expiry_datestring \nullISO date the document expires, if any
effective_statusstringCurrent legal status (Còn hiệu lực / Hết hiệu lực toàn bộ / Hết hiệu lực một phần)
agencystring \nullIssuing authority
signerstringSignatory name(s)
public_urlstringLink to the document on vbpl.vn
legal_baseslist[string]Legal grounds cited in the document's preamble ("Căn cứ...")
full_textstringFull structured text: chapters, sections, articles, clauses, points, in document order

articles (19449 rows)

One row per Điều (article), with its owning document/chapter/section context attached - the fine-grained retrieval unit.

json
{
  "id": "142847_art_1",
  "document_id": "142847",
  "document_title": "Luật Doanh nghiệp số 59/2020/QH14",
  "document_number": "59/2020/QH14",
  "chapter_number": "I",
  "chapter_title": "NHỮNG QUY ĐỊNH CHUNG",
  "section_number": null,
  "section_title": null,
  "article_number": "1",
  "article_title": "Phạm vi điều chỉnh",
  "full_text": "Điều 1. Phạm vi điều chỉnh\n1. Luật này quy định...",
  "public_url": "https://vbpl.vn/van-ban/chi-tiet/luat-doanh-nghiep-so-59-2020-qh14--142847"
}
FieldTypeDescription
idstringArticle id ({document_id}_art_{number})
document_idstringOwning document id
document_title, document_numberstringOwning document's title / official number
chapter_number, chapter_titlestring \nullOwning chapter, if any
section_number, section_titlestring \nullOwning section, if any
article_numberstringĐiều number
article_titlestringĐiều title
full_textstringFull text of this article, including its clauses/points
public_urlstringLink to the owning document on vbpl.vn

relationships (120259 rows)

The legal citation graph.

json
{
  "source_id": "152951",
  "source_label": "Document",
  "source_title": "Luật sửa đổi, bổ sung một số điều của...",
  "target_id": "142847",
  "target_label": "Document",
  "target_title": "Luật Doanh nghiệp số 59/2020/QH14",
  "relation_type": "AMENDS",
  "label_vn": "Văn bản sửa đổi bổ sung",
  "context": "Điều 21 Luật Doanh nghiệp"
}
FieldTypeDescription
source_id, target_idstringEndpoint ids (Document or Article)
source_label, target_labelstringDocument or Article
source_title, target_titlestring \nullResolved titles for readability
relation_typestringOne of CITES, AMENDS, REPLACES, BASED_ON, GUIDED_BY, CONSOLIDATES, CONSOLIDATED_DOC, REPEALS, REPEALED_BY, CORRECTS, SUSPENDED_BY, TRANSLATES, EXTENDS
label_vnstringVietnamese label for relation_type
contextstring \nullSurrounding text the relation was extracted from

entity_mentions (893 rows)

Document-level domain concept tags (Luật, Doanh nghiệp, Cơ quan, Thuế, Người lao động), keyword-extracted (no LLM/NLP).

json
{
  "document_id": "142847",
  "document_title": "Luật Doanh nghiệp số 59/2020/QH14",
  "concept": "Doanh nghiệp",
  "category": "DOANH_NGHIEP",
  "mention_count": 214
}
FieldTypeDescription
document_idstringDocument that mentions the concept
document_titlestringDocument title
conceptstringHuman-readable concept name
categorystringOne of LUAT, DOANH_NGHIEP, CO_QUAN, THUE, NGUOI_LAO_DONG
mention_countintNumber of times the concept appears in the document

Data Splits

Each config is a single split, no train/test division - this is a corpus and knowledge graph, not a supervised benchmark.

Dataset Creation

Source Data

Crawled via BFS from a seed set of business/enterprise-law documents, following each document's official cross-references (references field) up to a configurable depth. Raw HTML content is cleaned and parsed into a physical structure (Chapter → Section → Article → Clause → Point) using rule-based Vietnamese legal-text patterns (no LLM/NLP).

Annotations

legal_bases, structural boundaries, citation relationships, and entity_mentions are all extracted with regex/rule-based parsing over the document's own text; no manual annotation or LLM was used.

Considerations for Using the Data

  • —Coverage: this release covers 240 documents from a depth-limited crawl seeded around Vietnamese enterprise/business law - not a complete mirror of vbpl.vn.
  • —Parsing limitations: quoted amendment text and appendices are excluded from structural parsing on a best-effort basis; rare edge cases in complex omnibus amendment laws may still be imperfect.
  • —Not legal advice: effective_status reflects vbpl.vn's status at crawl time and may be outdated - always verify against the official source (public_url) before relying on this data for legal decisions.

Licensing Information

The dataset compilation (structuring, entity tagging, metadata, and the citation graph) is licensed under CC-BY-4.0. The underlying legal texts themselves are not subject to copyright under Vietnamese law (Điều 15, Luật Sở hữu trí tuệ - văn bản quy phạm pháp luật is excluded from copyright protection).

Additional Information

  • —Homepage: https://lexora-knowledge.trunganh.tech/
  • —Repository: https://github.com/LexoraAI/Lexora-Knowledge

Generated from release v1.0 (240 documents, 19449 articles, 120259 relationships, 893 entity mentions) by the Lexora Knowledge pipeline (pipelines/publish_pipeline.py).