CoolFace
Datasetpublic

tmquan/thuvienphapluat-vn-hdpl

Hỏi đáp pháp luật Việt Nam — Vietnamese Legal Q&A VI — Bộ dữ liệu 109.543 bài hỏi đáp pháp luật Việt Nam từ thuvienphapluat.vn/hoi-dap-phap-luat: câu hỏi ↔ câu trả lời dài có trích dẫn văn bản, kèm embedding Nemotron-3-Embed-8B và toạ độ giảm chiều trong không gian chung của 6 bộ dữ liệu pháp luật. EN — 109,543 Vietnamese legal question-and-answer articles from thuvienphapluat.vn, each pairing a legal question with a citation-grounded long-form answer, embedded with… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/thuvienphapluat-vn-hdpl.

sourceHugging Faceotherupdated 9d agoView on Hugging Face
0likes544downloads
Dataset Card

Hỏi đáp pháp luật Việt Nam — Vietnamese Legal Q&A

VI — Bộ dữ liệu 109.543 bài hỏi đáp pháp luật Việt Nam từ thuvienphapluat.vn/hoi-dap-phap-luat: câu hỏi ↔ câu trả lời dài có trích dẫn văn bản, kèm embedding Nemotron-3-Embed-8B và toạ độ giảm chiều trong không gian chung của 6 bộ dữ liệu pháp luật.

EN — 109,543 Vietnamese legal question-and-answer articles from thuvienphapluat.vn, each pairing a legal question with a citation-grounded long-form answer, embedded with Nemotron-3-Embed-8B and placed into a shared common-corpus t-SNE/UMAP space jointly fit over 6 Vietnamese legal datasets.

3. Overview

This is one of 6 Vietnamese legal datasets published under a single unified standard (documents / embeddings / reduces, plus dataset-specific extras). Each dataset carries its own cleaned content, a single 8B embedding table, and its rows' coordinates in one joint dimensionality-reduction space fit over all 6 corpora together (≈2.33M points, nvidia/Nemotron-3-Embed-8B-BF16, no PCA).

  • —Domain: Vietnamese legal Q&A (Hỏi đáp pháp luật) — a legal-question knowledge base with long-form, citation-grounded answers.
  • —Provenance: crawled from the public Hỏi đáp pháp luật section of THƯ VIỆN PHÁP LUẬT (thuvienphapluat.vn).
  • —Size: 109,543 Q&A · 27 legal areas · 19,219 extracted citations (73.6% of Q&A cite ≥1 law) · avg. answer 4,766 chars · published 2016-08-24 → 2026-08-18.
  • —Representative vector in the shared corpus: the answer embedding of each Q&A.

4. Configs

ConfigRowsWhat it is
documents (default)109,543one row per Q&A — question, answer, citations, area, metadata
embeddings109,543Nemotron-3-Embed-8B (4096-d) question & answer vectors, keyed by id
reduces438,172shared common-corpus 2-D/3-D t-SNE & UMAP coordinates (long)
pages (extra)109,543raw gzipped crawl HTML, one row per Q&A, keyed by id

5. Schema

`documents` (default) — one row per Q&A:

ColumnTypeNotes
idstringnumeric article id (from the URL); join key across all configs
url · sourcestringsource URL · site
questionstringthe legal question
answer · answer_htmlstringlong-form answer (plain text · original HTML)
category · areastringlegal domain — English slug · Vietnamese label
published_date · modified_datestringISO-8601
author · keywords · summarystring / listarticle metadata
citationslist<struct>parsed references (law_type, law_name, article, clause, point, ref, year, …)
num_citationsintnumber of citations
content_flags · content_flag_summarylist / structquality/content flags
answer_charsintanswer length

`embeddings` — Nemotron-3-Embed-8B vectors, keyed by id:

ColumnTypeNotes
idstringjoin key to documents
question_embeddinglist<float32>the question (model's query prompt)
answer_embeddinglist<float32>the answer (as a passage/document)
embedding_dimint4096
embedding_model_idstringnvidia/Nemotron-3-Embed-8B-BF16

`reduces` — long table of shared common-corpus coordinates:

ColumnTypeValues
idstringjoin key to documents
datasetstringhdpl (this dataset's tag in the joint corpus)
languagestringvi
methodstringtsne · umap
dimint2 · 3
coordslist<float>length = dim

`pages` (extra) — the raw crawl artifact, kept separate so the crawl itself can be studied:

ColumnTypeNotes
idstringjoin key to documents
page_html_gzbinaryraw crawled source page, gzip bytes (gzip.decompress(...).decode("utf-8") → the original HTML)

6. Embeddings

Every question and every answer is embedded with `nvidia/Nemotron-3-Embed-8B-BF16` (4096-d, one vector per field), living in the embeddings config as question_embedding and answer_embedding. The question uses the model's query prompt; the answer is encoded as a passage/document. Content is Vietnamese (`vi`). The answer embedding is this dataset's representative vector in the shared corpus below.

7. Dimensionality reduction — the common corpus

The reduces config is the centrepiece. It holds this dataset's rows' coordinates in one joint t-SNE / UMAP space fit over all 6 Vietnamese legal datasets together (≈2.33M points, nvidia/Nemotron-3-Embed-8B-BF16, no PCA). Because every dataset is embedded and reduced in the same space, coordinates are directly comparable across corpora.

  • —Long format: {id, dataset, language, method, dim, coords} — one row per (id × method × dim).
  • —Methods: tsne, umap · Dims: 2, 3 · All 109,543 documents present in every (method, dim) combination → 438,172 rows.
  • —Shared fit: a single reduction was fit over the concatenation of all 6 datasets' 8B vectors; dataset / language tag each point's origin.
python
red = load_dataset("tmquan/thuvienphapluat-vn-hdpl", "reduces")["train"].to_pandas()
umap2d = red[(red.method == "umap") & (red.dim == 2)]   # this dataset's points in the joint UMAP

8. Visualizations

The whole common corpus (all 6 datasets) in the shared space:

[image] [image] [image]

This dataset on its own (coloured by legal area) — coordinates taken from the shared 6-dataset reduce above (not a per-dataset re-fit), so this map is a zoom onto just this dataset's points within the same joint t-SNE/UMAP space:

[image] [image]

Citation Sankey — legal area → most-cited laws:

[image]

Luật được trích dẫn nhiều nhất / Most-cited lawQ&A
Bộ luật Lao động1,720
Luật Quản lý thuế1,041
Bộ luật hình sự885
Luật Đất đai764
Luật Thuế thu nhập cá nhân646
Bộ luật Dân sự474
Luật Bảo hiểm xã hội320
Luật Đấu thầu279
Luật Nhà ở222
Luật Đầu tư212

Areas (Lĩnh vực) — all 27

Lĩnh vực / AreaQ&A
Bộ máy hành chính21,293
Giáo dục14,712
Thuế - Phí - Lệ phí9,360
Lao động - Tiền lương7,334
Doanh nghiệp4,687
Bất động sản4,424
Tài nguyên - Môi trường4,300
Bảo hiểm3,507
Thương mại3,298
Giao thông - Vận tải3,272
Quyền dân sự3,092
Trách nhiệm hình sự3,056
Văn hóa - Xã hội2,951
Đầu tư2,940
Vi phạm hành chính2,439
Thể thao - Y tế2,414
Kế toán - Kiểm toán2,226
Công nghệ thông tin2,218
Tiền tệ - Ngân hàng2,052
Xây dựng - Đô thị1,925
Lĩnh vực khác1,809
Thủ tục tố tụng1,744
Dịch vụ pháp lý1,319
Tài chính nhà nước1,227
Xuất nhập khẩu961
Chứng khoán499
Sở hữu trí tuệ423

9. Analysis pointers

The shared coordinates support cross-dataset geometric analysis directly:

  • —kNN graph over the joint reduces (or the raw 8B vectors) to find nearest Q&A across all 6 corpora — e.g. link a legal question to statutes, judgments, or terminology entries living in sibling datasets.
  • —Clustering (HDBSCAN / k-means) on the 2-D/3-D coords to surface legal-topic structure, then compare cluster composition by dataset / area.
  • —Cross-dataset geometry: measure how the Q&A cloud overlaps the legislation / case-law clouds in the shared space (density, boundary regions, bridges) to study coverage and retrieval reach.

10. Usage

python
import gzip
from datasets import load_dataset

docs  = load_dataset("tmquan/thuvienphapluat-vn-hdpl", "documents")   # Q&A + metadata (no vectors, no raw page)
emb   = load_dataset("tmquan/thuvienphapluat-vn-hdpl", "embeddings")  # Nemotron-8B question/answer vectors
red   = load_dataset("tmquan/thuvienphapluat-vn-hdpl", "reduces")     # shared common-corpus 2D/3D t-SNE/UMAP coords
pages = load_dataset("tmquan/thuvienphapluat-vn-hdpl", "pages")       # raw gzipped crawl HTML, keyed by id

html  = gzip.decompress(pages["train"][0]["page_html_gz"]).decode("utf-8")   # raw source HTML (gzip bytes)

Join all configs on id.

11. Provenance & crawl

Content was crawled from the public Hỏi đáp pháp luật section of THƯ VIỆN PHÁP LUẬT (thuvienphapluat.vn) for research. The clean documents table (question, answer, parsed citations, area, metadata) is derived from the raw source pages preserved byte-for-byte in the `pages` extra config (gzipped HTML, keyed by id) so the full extraction pipeline is reproducible. Embeddings and shared-corpus coordinates were computed downstream with nvidia/Nemotron-3-Embed-8B-BF16.

12. License & citation

License: other — content © thuvienphapluat.vn, redistributed for research. This is informational legal material and is not legal advice; answers reflect the law as published at crawl time.

bibtex
@misc{tvpl_hdpl_2026,
  title        = {Hỏi đáp pháp luật Việt Nam — Vietnamese Legal Q&A (thuvienphapluat.vn)},
  author       = {TMQuan},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/tmquan/thuvienphapluat-vn-hdpl}},
  note         = {109,543 Vietnamese legal Q&A with parsed citations, Nemotron-3-Embed-8B question/answer embeddings, and shared common-corpus 2D/3D t-SNE/UMAP coordinates jointly fit over 6 Vietnamese legal datasets.}
}

@misc{tvpl_hdpl_source_2026,
  title        = {Hỏi đáp pháp luật — THƯ VIỆN PHÁP LUẬT},
  author       = {{Hỏi đáp pháp luật — THƯ VIỆN PHÁP LUẬT}},
  year         = {2026},
  howpublished = {\url{https://thuvienphapluat.vn/hoi-dap-phap-luat}},
  note         = {Official Vietnamese legal question-and-answer knowledge base published by THƯ VIỆN PHÁP LUẬT (thuvienphapluat.vn).}
}