CoolFace
Datasetpublic

tmquan/cbba-toaan-gov-vn

Vietnamese Bản án Corpus — congbobanan.toaan.gov.vn 🇻🇳 Tóm tắt. Bản án sơ thẩm/phúc thẩm/giám đốc thẩm/tái thẩm của Việt Nam, thu thập từ cổng công bố bản án congbobanan.toaan.gov.vn của Tòa án nhân dân tối cao. Ba cấu hình HF khoá theo doc_name/id: documents (nội dung siêu dữ liệu + trích dẫn), embeddings (vector 4096-D Nemotron-3-8B), reduces (toạ độ t-SNE/UMAP trong không gian chung 6 bộ dữ liệu). Tên cột và giá trị phân loại bằng tiếng Anh; chỉ nội dung pháp lý giữ tiếng… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/cbba-toaan-gov-vn.

sourceHugging Facecc-by-4.0updated 10d agoView on Hugging Face
0likes684downloads
Dataset Card

Vietnamese Bản án Corpus — congbobanan.toaan.gov.vn

🇻🇳 Tóm tắt. Bản án sơ thẩm/phúc thẩm/giám đốc thẩm/tái thẩm của Việt Nam, thu thập từ cổng công bố bản án `congbobanan.toaan.gov.vn` của Tòa án nhân dân tối cao. Ba cấu hình HF khoá theo doc_name/id: `documents` (nội dung + siêu dữ liệu + trích dẫn), `embeddings` (vector 4096-D Nemotron-3-8B), `reduces` (toạ độ t-SNE/UMAP trong không gian chung 6 bộ dữ liệu). Tên cột và giá trị phân loại bằng tiếng Anh; chỉ nội dung pháp lý giữ tiếng Việt. 🇬🇧 Summary. Vietnamese first-instance / appellate / cassation / retrial court judgments harvested from the Supreme People's Court (Tòa án nhân dân tối cao) judgment-portal. Three HF configs keyed by doc_name/id: `documents` (content + metadata + citations), `embeddings` (4096-D Nemotron-3-8B vectors), `reduces` (t-SNE/UMAP coordinates in the shared 6-dataset ViLA common-corpus space). Column names and categorical values are English; only the legal content stays Vietnamese.

This dataset is one of six that make up the ViLA legal common corpus. All six share a single embedding model and one joint t-SNE/UMAP fit, so their coordinates are directly comparable (see Dimensionality reduction).

1 · Tổng quan · Overview

🇻🇳 Bản phát hành này chứa 1.542.247 văn bản — tập con của lần thu thập có lớp text gốc trích xuất được từ PDF. · 🇬🇧 This release contains 1,542,247 documents — the subset of the crawl whose PDFs carry an extractable native text layer.

Chỉ số · MetricGiá trị · Value
Số văn bản · Documents1,542,247 (native-text subset of 2,101,504 crawled)
Ngôn ngữ · LanguageVietnamese content (vi); English column names + categorical values
Nguồn · Source<https://congbobanan.toaan.gov.vn/> — Tòa án nhân dân tối cao (Supreme People's Court)
Thu thập · Crawled2026-08
Phân loại vụ án · Case categories8 English enums: Marriage & Family, Civil, Criminal, Commercial, Administrative, Labor, Bankruptcy, Economic
Cấp xét xử · Instance levelFirst-instance · Appellate · Cassation · Retrial
Cấp toà · Court levelDistrict · Provincial · High · Supreme
Mô hình embedding · Embedding modelnvidia/Nemotron-3-Embed-8B-BF16 (4096-d), một vector/văn bản · one vector/document
Giảm chiều · Reductionshared common-corpus t-SNE + UMAP, 2-D & 3-D (no PCA)
Giấy phép · LicenseCC BY 4.0
Corpus liên quan · Sibling`tmquan/anle-toaan-gov-vn` — shares this document schema

🇻🇳 558.373 văn bản còn lại (26,6 %) không bị mất: 381.104 là ảnh scan không có lớp text và 177.015 có lớp text nhưng font nhúng thiếu bảng ToUnicode; chúng chờ OCR ở phiên bản sau. · 🇬🇧 The remaining 558,373 judgments (26.6%) are not lost: 381,104 are scanned images with no text layer and 177,015 have a text layer whose embedded font lacks a ToUnicode map; they await OCR in a later release.

2 · Cấu hình · Configs

🇻🇳 Ba cấu hình đều khoá theo cùng một id văn bản (doc_name trong documents /embeddings, id trong reduces) nên nối chéo được. · 🇬🇧 All three configs key on the same document id (doc_name in documents/embeddings, id in reduces), so they join.

ConfigSplitSố dòng · RowsFile(s)Nội dung · Contents
documents (default)train1,542,247documents-*.parquet (16)Bản án + siêu dữ liệu + trích dẫn · Judgment body + metadata + citations
embeddingstrain1,542,247embeddings-*.parquet (39)Vector 4096-D Nemotron-3-8B · 4096-D doc vectors
reducestrain6,168,988reduces-*.parquet (2)Toạ độ t-SNE/UMAP 2-D+3-D trong không gian chung · shared-space t-SNE/UMAP 2-D+3-D coords

🇻🇳 reduces có 6.168.988 dòng = 1.542.247 văn bản × {t-SNE,UMAP} × {2-D,3-D}. · 🇬🇧 reduces has 6,168,988 rows = 1,542,247 documents × {t-SNE,UMAP} × {2-D,3-D}.

3 · Lược đồ · Schema

documents — một dòng mỗi bản án · one row per judgment

Cột · ColumnKiểu · TypeGhi chú · Notes
doc_namestringKhoá chính · Primary key (portal document id)
sourcestringLuôn congbobanan.toaan.gov.vn
web_urlstringURL trang chi tiết · Judgment detail-page URL
pdf_urlstringURL PDF đã ký · Signed judgment-PDF URL
official_document_idstringSố hiệu như in trên văn bản, vd 2/2017/LĐ-PT · Official judgment number as printed
official_document_id_normalizedstringSố hiệu đã chuẩn hoá · Normalised form
numberdoubleSố thứ tự bản án · Sequence number (null when no_official_id)
yeardoubleNăm ban hành · Year from the judgment number
codestringMã loại văn bản, vd HS-ST, DS-PT, QĐST-HNGĐ · Document-type code (Vietnamese)
id_sourcestringCách trích số hiệu (regex) · How the id was extracted
categorystringLĩnh vực vụ án (English enum, 8 giá trị) · Case category (English enum)
instance_levelstringFirst-instance · Appellate · Cassation · Retrial
courtstringTên toà (nội dung tiếng Việt) · Court name (Vietnamese content)
court_levelstringDistrict · Provincial · High · Supreme
issued_datestringNgày ban hành (ISO YYYY-MM-DD) · Issue date
date_sourcestringCách trích ngày (regex) · How the date was extracted
precedent_numbernullToàn bộ null trong bản này · All-null in this release (typed column reserved)
is_precedentboolCó phải án lệ không · Whether flagged as a precedent (án lệ)
citations_lawlist&lt;struct&gt;Trích dẫn văn bản pháp luật (regex, có span) · Regex primary-law citations (with spans)
citations_caselist&lt;struct&gt;Trích dẫn bản án/quyết định khác (regex) · Regex case citations
citations_sourcestringCách trích dẫn (regex)
num_law_citationsint64Số citations_law
num_case_citationsint64Số citations_case
markdownstringThân bản án dạng markdown, chuẩn hoá NFC · Vietnamese-normalised (NFC) markdown body
markdown_charsint64Số ký tự của markdown
num_pagesint64Số trang PDF gốc · Source-PDF page count
confidencenullToàn bộ null trong bản này · All-null in this release (typed column reserved)
flagslist&lt;string&gt;Cờ chất lượng, vd no_official_id · Quality flags
citationslist&lt;struct&gt;Trích dẫn từ bộ phân tích cấu trúc (neo Điều/khoản/điểm + văn bản đích) · Structure-parser citations (Điều/khoản/điểm anchors + target act) — độc lập với citations_law/_case · independent of the regex layers
parent_actslist&lt;string&gt;Văn bản luật gốc phát hiện trong thân · Parent acts detected in the body
num_citations, num_articles, num_khoan, num_diemint64Số phần tử citations / số Điều / khoản / điểm theo cấu trúc · Structural counts

🇻🇳 `citations_law` (mỗi phần tử): ref, law_type, law_name, article, chapter, clause, point, section, kind, id, year, sentence_id[], span[[start,end]]. `citations_case`: id, code, domain, level, number, role, year, sentence_id[], span[]. `citations` (cấu trúc): source, target, dieu, khoan, diem, kind, raw. · 🇬🇧 See the same field lists for the nested citation structs; span values are character offsets into markdown, sentence_id are {doc_name}#{idx} anchors.

embeddings — một vector mỗi bản án · one vector per judgment

Cột · ColumnKiểu · TypeGhi chú · Notes
doc_namestringKhoá nối tới documents/reduces · Join key
embeddinglist&lt;double&gt;Vector 4096-D, mean-pooled trên các đoạn passage: , L2-normalised · 4096-D L2-normalised vector
embedding_dimint644096
embedding_model_idstringnvidia/Nemotron-3-Embed-8B-BF16

reduces — bảng dài, toạ độ trong không gian chung · long table, shared-space coordinates

Cột · ColumnKiểu · TypeGhi chú · Notes
idstringKhoá nối tới documents.doc_name · Join key to documents.doc_name
datasetstringLuôn cbba (nhãn bộ dữ liệu trong corpus chung) · Dataset tag within the common corpus
languagestringLuôn vi
methodstringtsne · umap
dimint642 · 3
coordslist&lt;double&gt;Toạ độ, độ dài = dim · Coordinates, length = dim

4 · Embeddings

🇻🇳 Mỗi văn bản được nhúng bằng `nvidia/Nemotron-3-Embed-8B-BF16` (4096 chiều): thân markdown được cắt đoạn với template passage: , encode, rồi mean-pool và L2-normalise thành một vector cho mỗi bản án. Chạy trên một NVIDIA GB10 qua vLLM. · 🇬🇧 Each document is embedded with `nvidia/Nemotron-3-Embed-8B-BF16` (4096-d): the markdown body is chunked with the passage: template, encoded, then mean-pooled and L2-normalised into one vector per judgment. Served on a single NVIDIA GB10 via vLLM. The config carries 1,542,247 vectors, one per row in documents.

5 · Dimensionality reduction — the shared common corpus

🇻🇳 Đây là điểm nhấn của bản phát hành. Toàn bộ 2.326.918 văn bản của cả 6 bộ dữ liệu ViLA được embed bằng Nemotron-3-Embed-8B (4096-D) rồi giảm chiều trong MỘT lần fit chung — t-SNE và UMAP, cả 2-D lẫn 3-D, không dùng PCA. Vì vậy toạ độ của cbba so sánh trực tiếp được với 5 bộ còn lại: cùng một không gian. Config reduces ở đây chỉ chứa các dòng của cbba trong không gian chung đó. · 🇬🇧 This is the centerpiece. All 2,326,918 documents across the six ViLA datasets are embedded with Nemotron-3-Embed-8B (4096-D) and reduced in ONE joint fit — t-SNE and UMAP, at both 2-D and 3-D, with no PCA. cbba's coordinates are therefore directly comparable to the other five datasets: it is one shared space. This reduces config holds only cbba's rows of that space.

Long schema {id, dataset, language, method, dim, coords} — methods {tsne, umap}, dims {2, 3}, all 1,542,247 documents present at each (method × dim) → 6,168,988 rows. There are no `pca` rows.

Bộ corpus chung · The shared corpus (all six datasets, one joint fit):

DatasetRowsLanguageRows
cbba (this)1,542,247vi2,309,754
tvpl-vbpl574,314en17,164
hdpl109,543
phapdien65,568
tnpl33,402
anle1,844
Total2,326,918Total2,326,918

6 · Trực quan hoá · Visualizations

Bản đồ toàn corpus chung · Global common-corpus maps

🇻🇳 Ba bản đồ dưới đây là một lần fit chung trên toàn bộ 2.326.918 văn bản của cả 6 bộ; cùng những hình này xuất hiện trên mọi card ViLA. · 🇬🇧 The three maps below are the single joint fit over all 2,326,918 documents of all six datasets; the same images appear on every ViLA card.

[image] [image] [image]

Riêng bộ cbba · This dataset, coloured by case category

🇻🇳 Toạ độ lấy trực tiếp từ config reduces (dim = 2), tô màu theo category. · 🇬🇧 Coordinates taken directly from the reduces config (dim = 2), coloured by category.

[image] [image]

7 · Gợi ý phân tích · Analysis pointers

🇻🇳 Vì toạ độ nằm trong không gian chung, có thể: (1) dựng đồ thị kNN trên coords (hoặc trên vector 4096-D) để truy hồi bản án tương tự; (2) phân cụm (HDBSCAN/k-means) trên toạ độ 2-D/3-D và đối chiếu cụm với category/court_level; (3) đo hình học chéo bộ dữ liệu — cbba (bản án) so với luật (tvpl-vbpl, phapdien), hỏi đáp (hdpl), án lệ (anle) trong cùng một không gian. · 🇬🇧 Because the coordinates live in the shared space you can: (1) build a kNN graph over coords (or over the 4096-D vectors) for similar-judgment retrieval; (2) cluster (HDBSCAN/k-means) on the 2-D/3-D coordinates and compare clusters against category/court_level; (3) measure cross-dataset geometry — cbba judgments vs statutes (tvpl-vbpl, phapdien), Q&A (hdpl) and precedents (anle) in one common space. Join reduces.id == documents.doc_name to attach the label facets.

8 · Cách dùng · How to use

python
from datasets import load_dataset

# Văn bản (mặc định) · Documents (default)
docs = load_dataset("tmquan/cbba-toaan-gov-vn", "documents", split="train")
print(docs[0]["official_document_id"], docs[0]["category"], docs[0]["court_level"])
print(docs[0]["markdown"][:200])

# Vector 4096-D · 4096-D vectors
emb = load_dataset("tmquan/cbba-toaan-gov-vn", "embeddings", split="train")
print(emb[0]["embedding_model_id"], emb[0]["embedding_dim"])  # nvidia/Nemotron-3-Embed-8B-BF16 4096

# Toạ độ không gian chung · Shared-space coordinates (long)
red = load_dataset("tmquan/cbba-toaan-gov-vn", "reduces", split="train")
umap2d = red.filter(lambda r: r["method"] == "umap" and r["dim"] == 2)
print(umap2d[0]["id"], umap2d[0]["coords"])  # [x, y]

# Nối chéo theo id văn bản · Join on the document id
red_by_doc = {(r["id"], r["method"], r["dim"]): r["coords"] for r in umap2d}
row = docs[0]
xy = red_by_doc[(row["doc_name"], "umap", 2)]

9 · Cách dữ liệu được tạo · Provenance & how the corpus was built

🇻🇳 Thu thập từ cổng công khai congbobanan.toaan.gov.vn; PDF đã ký được tải và phân giải bằng pypdf + một bộ "cmap healer" (sửa font); chỉ giữ văn bản có lớp text gốc trích xuất được (73,4 %), phần còn lại hoãn cho OCR. Thân bản án chuẩn hoá NFC và render sang markdown theo trang; siêu dữ liệu (category, instance_level, court_level, số hiệu, ngày) và trích dẫn được trích bằng regex kèm id câu + span ký tự; một lớp trích dẫn thứ hai đến từ bộ phân tích cấu trúc pháp luật. Sau đó embedding trên GB10 (Nemotron-3-Embed-8B, template passage: , L2-normalise) và giảm chiều chung với 5 bộ ViLA còn lại (t-SNE + UMAP, 2-D & 3-D, không PCA). · 🇬🇧 Harvested from the public congbobanan.toaan.gov.vn portal; signed PDFs are downloaded and parsed with pypdf plus a cmap healer; only documents with an extractable native text layer are kept (73.4%), the rest deferred to OCR. Bodies are NFC-normalised and rendered to per-page markdown; metadata (category, instance_level, court_level, number, date) and citations are regex-extracted with sentence ids + character spans, plus a second citation layer from the Vietnamese legal structure parser. Documents are then embedded on a GB10 (Nemotron-3-Embed-8B, passage: template, L2-normalised) and reduced jointly with the five other ViLA datasets (t-SNE + UMAP, 2-D & 3-D, no PCA).

Hạn chế đã biết · Known limitations

🇻🇳 (i) Độ phủ 73,4 % (còn 26,6 % chờ OCR). (ii) Lệch lĩnh vực: Marriage & Family chiếm ưu thế (≈47 %), rồi Civil, Criminal; Bankruptcy/Economic rất thưa; ~193k văn bản chưa gán category. (iii) Siêu dữ liệu & trích dẫn bằng regex có thể sai ở văn bản định dạng bất thường. (iv) precedent_number và confidence là cột đã khai kiểu nhưng toàn bộ null. (v) Dữ liệu vụ án thật — xem mục Riêng tư. · 🇬🇧 (i) 73.4% coverage (26.6% await OCR). (ii) Category skew: Marriage & Family dominates (≈47%), then Civil, Criminal; Bankruptcy/Economic are very sparse; ~193k documents have no category. (iii) Regex-extracted metadata & citations may be wrong on unusually formatted judgments. (iv) precedent_number and confidence are typed but all-null. (v) Real case data — see Personal / sensitive information.

Riêng tư & đạo đức · Personal / sensitive information

🇻🇳 Bản án được Tòa án nhân dân công bố công khai theo Nghị quyết 03/2017/NQ-HĐTP; toà đã ẩn danh một phần tên đương sự trước khi công bố. Văn bản vẫn có thể chứa thông tin cá nhân nhạy cảm. Hãy dùng ở mức tổng hợp/nghiên cứu và tuân thủ pháp luật Việt Nam về bảo vệ dữ liệu cá nhân. · 🇬🇧 Judgments are published by the People's Courts under Resolution 03/2017/NQ-HĐTP; the court partially anonymises party names before publication. The text can still contain sensitive personal information. Use at an aggregate/research level and comply with Vietnamese personal-data law.

10 · Giấy phép & trích dẫn · License & citation

🇻🇳 Bản án là văn bản công khai do Tòa án nhân dân tối cao công bố (Nghị quyết 03/2017/NQ-HĐTP). Bản phân phối lại này theo CC BY 4.0 kèm kỳ vọng: (1) ghi nhận nguồn congbobanan.toaan.gov.vn; (2) không trình bày bản trích xuất như văn bản pháp lý chính thức; (3) tôn trọng tính nhạy cảm của dữ liệu vụ án; (4) tuân thủ pháp luật Việt Nam. · 🇬🇧 Judgments are public documents published by the Supreme People's Court (Resolution 03/2017/NQ-HĐTP). This redistribution is CC BY 4.0 with the expectations that you (1) attribute congbobanan.toaan.gov.vn; (2) do not present the extraction as official legal text; (3) respect the sensitivity of case data; (4) comply with applicable Vietnamese law.

bibtex
@misc{cbba_2026,
  title        = {Vietnamese Bản án Corpus (congbobanan.toaan.gov.vn)},
  author       = {TMQuan},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/tmquan/cbba-toaan-gov-vn}},
  note         = {Document-level mirror of published Vietnamese court judgments with 4096-D Nemotron-3-Embed-8B embeddings and shared-space (6-dataset joint) t-SNE / UMAP 2-D & 3-D projections over 1,542,247 natively-extractable documents.}
}

@misc{congbobanan_toaan_2026,
  title        = {Cổng công bố bản án, quyết định của Tòa án},
  author       = {{Công bố bản án — Tòa án nhân dân tối cao}},
  year         = {2026},
  howpublished = {\url{https://congbobanan.toaan.gov.vn/}},
  note         = {Official judgment-publication portal of the Supreme People's Court of Vietnam. Judgments are public documents published under Resolution 03/2017/NQ-HDTP.}
}