tmquan/cbba-toaan-gov-vn
Vietnamese Bản án Corpus — congbobanan.toaan.gov.vn 🇻🇳 Tóm tắt. Bản án sơ thẩm/phúc thẩm/giám đốc thẩm/tái thẩm của Việt Nam, thu thập từ cổng công bố bản án congbobanan.toaan.gov.vn của Tòa án nhân dân tối cao. Ba cấu hình HF khoá theo doc_name/id: documents (nội dung siêu dữ liệu + trích dẫn), embeddings (vector 4096-D Nemotron-3-8B), reduces (toạ độ t-SNE/UMAP trong không gian chung 6 bộ dữ liệu). Tên cột và giá trị phân loại bằng tiếng Anh; chỉ nội dung pháp lý giữ tiếng… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/cbba-toaan-gov-vn.
Vietnamese Bản án Corpus — congbobanan.toaan.gov.vn
🇻🇳 Tóm tắt. Bản án sơ thẩm/phúc thẩm/giám đốc thẩm/tái thẩm của Việt Nam, thu thập từ cổng công bố bản án `congbobanan.toaan.gov.vn` của Tòa án nhân dân tối cao. Ba cấu hình HF khoá theodoc_name/id: `documents` (nội dung + siêu dữ liệu + trích dẫn), `embeddings` (vector 4096-D Nemotron-3-8B), `reduces` (toạ độ t-SNE/UMAP trong không gian chung 6 bộ dữ liệu). Tên cột và giá trị phân loại bằng tiếng Anh; chỉ nội dung pháp lý giữ tiếng Việt. 🇬🇧 Summary. Vietnamese first-instance / appellate / cassation / retrial court judgments harvested from the Supreme People's Court (Tòa án nhân dân tối cao) judgment-portal. Three HF configs keyed bydoc_name/id: `documents` (content + metadata + citations), `embeddings` (4096-D Nemotron-3-8B vectors), `reduces` (t-SNE/UMAP coordinates in the shared 6-dataset ViLA common-corpus space). Column names and categorical values are English; only the legal content stays Vietnamese.
This dataset is one of six that make up the ViLA legal common corpus. All six share a single embedding model and one joint t-SNE/UMAP fit, so their coordinates are directly comparable (see Dimensionality reduction).
1 · Tổng quan · Overview
🇻🇳 Bản phát hành này chứa 1.542.247 văn bản — tập con của lần thu thập có lớp text gốc trích xuất được từ PDF. · 🇬🇧 This release contains 1,542,247 documents — the subset of the crawl whose PDFs carry an extractable native text layer.
🇻🇳 558.373 văn bản còn lại (26,6 %) không bị mất: 381.104 là ảnh scan không có lớp text và 177.015 có lớp text nhưng font nhúng thiếu bảng ToUnicode; chúng chờ OCR ở phiên bản sau. · 🇬🇧 The remaining 558,373 judgments (26.6%) are not lost: 381,104 are scanned images with no text layer and 177,015 have a text layer whose embedded font lacks a ToUnicode map; they await OCR in a later release.
2 · Cấu hình · Configs
🇻🇳 Ba cấu hình đều khoá theo cùng một id văn bản (doc_name trong documents /embeddings, id trong reduces) nên nối chéo được. · 🇬🇧 All three configs key on the same document id (doc_name in documents/embeddings, id in reduces), so they join.
🇻🇳 reduces có 6.168.988 dòng = 1.542.247 văn bản × {t-SNE,UMAP} × {2-D,3-D}. · 🇬🇧 reduces has 6,168,988 rows = 1,542,247 documents × {t-SNE,UMAP} × {2-D,3-D}.
3 · Lược đồ · Schema
documents — một dòng mỗi bản án · one row per judgment
🇻🇳 `citations_law` (mỗi phần tử): ref, law_type, law_name, article, chapter, clause, point, section, kind, id, year, sentence_id[], span[[start,end]]. `citations_case`: id, code, domain, level, number, role, year, sentence_id[], span[]. `citations` (cấu trúc): source, target, dieu, khoan, diem, kind, raw. · 🇬🇧 See the same field lists for the nested citation structs; span values are character offsets into markdown, sentence_id are {doc_name}#{idx} anchors.
embeddings — một vector mỗi bản án · one vector per judgment
reduces — bảng dài, toạ độ trong không gian chung · long table, shared-space coordinates
4 · Embeddings
🇻🇳 Mỗi văn bản được nhúng bằng `nvidia/Nemotron-3-Embed-8B-BF16` (4096 chiều): thân markdown được cắt đoạn với template passage: , encode, rồi mean-pool và L2-normalise thành một vector cho mỗi bản án. Chạy trên một NVIDIA GB10 qua vLLM. · 🇬🇧 Each document is embedded with `nvidia/Nemotron-3-Embed-8B-BF16` (4096-d): the markdown body is chunked with the passage: template, encoded, then mean-pooled and L2-normalised into one vector per judgment. Served on a single NVIDIA GB10 via vLLM. The config carries 1,542,247 vectors, one per row in documents.
5 · Dimensionality reduction — the shared common corpus
🇻🇳 Đây là điểm nhấn của bản phát hành. Toàn bộ 2.326.918 văn bản của cả 6 bộ dữ liệu ViLA được embed bằng Nemotron-3-Embed-8B (4096-D) rồi giảm chiều trong MỘT lần fit chung — t-SNE và UMAP, cả 2-D lẫn 3-D, không dùng PCA. Vì vậy toạ độ của cbba so sánh trực tiếp được với 5 bộ còn lại: cùng một không gian. Config reduces ở đây chỉ chứa các dòng của cbba trong không gian chung đó. · 🇬🇧 This is the centerpiece. All 2,326,918 documents across the six ViLA datasets are embedded with Nemotron-3-Embed-8B (4096-D) and reduced in ONE joint fit — t-SNE and UMAP, at both 2-D and 3-D, with no PCA. cbba's coordinates are therefore directly comparable to the other five datasets: it is one shared space. This reduces config holds only cbba's rows of that space.
Long schema {id, dataset, language, method, dim, coords} — methods {tsne, umap}, dims {2, 3}, all 1,542,247 documents present at each (method × dim) → 6,168,988 rows. There are no `pca` rows.
Bộ corpus chung · The shared corpus (all six datasets, one joint fit):
6 · Trực quan hoá · Visualizations
Bản đồ toàn corpus chung · Global common-corpus maps
🇻🇳 Ba bản đồ dưới đây là một lần fit chung trên toàn bộ 2.326.918 văn bản của cả 6 bộ; cùng những hình này xuất hiện trên mọi card ViLA. · 🇬🇧 The three maps below are the single joint fit over all 2,326,918 documents of all six datasets; the same images appear on every ViLA card.
Riêng bộ cbba · This dataset, coloured by case category
🇻🇳 Toạ độ lấy trực tiếp từ config reduces (dim = 2), tô màu theo category. · 🇬🇧 Coordinates taken directly from the reduces config (dim = 2), coloured by category.
7 · Gợi ý phân tích · Analysis pointers
🇻🇳 Vì toạ độ nằm trong không gian chung, có thể: (1) dựng đồ thị kNN trên coords (hoặc trên vector 4096-D) để truy hồi bản án tương tự; (2) phân cụm (HDBSCAN/k-means) trên toạ độ 2-D/3-D và đối chiếu cụm với category/court_level; (3) đo hình học chéo bộ dữ liệu — cbba (bản án) so với luật (tvpl-vbpl, phapdien), hỏi đáp (hdpl), án lệ (anle) trong cùng một không gian. · 🇬🇧 Because the coordinates live in the shared space you can: (1) build a kNN graph over coords (or over the 4096-D vectors) for similar-judgment retrieval; (2) cluster (HDBSCAN/k-means) on the 2-D/3-D coordinates and compare clusters against category/court_level; (3) measure cross-dataset geometry — cbba judgments vs statutes (tvpl-vbpl, phapdien), Q&A (hdpl) and precedents (anle) in one common space. Join reduces.id == documents.doc_name to attach the label facets.
8 · Cách dùng · How to use
from datasets import load_dataset
# Văn bản (mặc định) · Documents (default)
docs = load_dataset("tmquan/cbba-toaan-gov-vn", "documents", split="train")
print(docs[0]["official_document_id"], docs[0]["category"], docs[0]["court_level"])
print(docs[0]["markdown"][:200])
# Vector 4096-D · 4096-D vectors
emb = load_dataset("tmquan/cbba-toaan-gov-vn", "embeddings", split="train")
print(emb[0]["embedding_model_id"], emb[0]["embedding_dim"]) # nvidia/Nemotron-3-Embed-8B-BF16 4096
# Toạ độ không gian chung · Shared-space coordinates (long)
red = load_dataset("tmquan/cbba-toaan-gov-vn", "reduces", split="train")
umap2d = red.filter(lambda r: r["method"] == "umap" and r["dim"] == 2)
print(umap2d[0]["id"], umap2d[0]["coords"]) # [x, y]
# Nối chéo theo id văn bản · Join on the document id
red_by_doc = {(r["id"], r["method"], r["dim"]): r["coords"] for r in umap2d}
row = docs[0]
xy = red_by_doc[(row["doc_name"], "umap", 2)]9 · Cách dữ liệu được tạo · Provenance & how the corpus was built
🇻🇳 Thu thập từ cổng công khai congbobanan.toaan.gov.vn; PDF đã ký được tải và phân giải bằng pypdf + một bộ "cmap healer" (sửa font); chỉ giữ văn bản có lớp text gốc trích xuất được (73,4 %), phần còn lại hoãn cho OCR. Thân bản án chuẩn hoá NFC và render sang markdown theo trang; siêu dữ liệu (category, instance_level, court_level, số hiệu, ngày) và trích dẫn được trích bằng regex kèm id câu + span ký tự; một lớp trích dẫn thứ hai đến từ bộ phân tích cấu trúc pháp luật. Sau đó embedding trên GB10 (Nemotron-3-Embed-8B, template passage: , L2-normalise) và giảm chiều chung với 5 bộ ViLA còn lại (t-SNE + UMAP, 2-D & 3-D, không PCA). · 🇬🇧 Harvested from the public congbobanan.toaan.gov.vn portal; signed PDFs are downloaded and parsed with pypdf plus a cmap healer; only documents with an extractable native text layer are kept (73.4%), the rest deferred to OCR. Bodies are NFC-normalised and rendered to per-page markdown; metadata (category, instance_level, court_level, number, date) and citations are regex-extracted with sentence ids + character spans, plus a second citation layer from the Vietnamese legal structure parser. Documents are then embedded on a GB10 (Nemotron-3-Embed-8B, passage: template, L2-normalised) and reduced jointly with the five other ViLA datasets (t-SNE + UMAP, 2-D & 3-D, no PCA).
Hạn chế đã biết · Known limitations
🇻🇳 (i) Độ phủ 73,4 % (còn 26,6 % chờ OCR). (ii) Lệch lĩnh vực: Marriage & Family chiếm ưu thế (≈47 %), rồi Civil, Criminal; Bankruptcy/Economic rất thưa; ~193k văn bản chưa gán category. (iii) Siêu dữ liệu & trích dẫn bằng regex có thể sai ở văn bản định dạng bất thường. (iv) precedent_number và confidence là cột đã khai kiểu nhưng toàn bộ null. (v) Dữ liệu vụ án thật — xem mục Riêng tư. · 🇬🇧 (i) 73.4% coverage (26.6% await OCR). (ii) Category skew: Marriage & Family dominates (≈47%), then Civil, Criminal; Bankruptcy/Economic are very sparse; ~193k documents have no category. (iii) Regex-extracted metadata & citations may be wrong on unusually formatted judgments. (iv) precedent_number and confidence are typed but all-null. (v) Real case data — see Personal / sensitive information.
Riêng tư & đạo đức · Personal / sensitive information
🇻🇳 Bản án được Tòa án nhân dân công bố công khai theo Nghị quyết 03/2017/NQ-HĐTP; toà đã ẩn danh một phần tên đương sự trước khi công bố. Văn bản vẫn có thể chứa thông tin cá nhân nhạy cảm. Hãy dùng ở mức tổng hợp/nghiên cứu và tuân thủ pháp luật Việt Nam về bảo vệ dữ liệu cá nhân. · 🇬🇧 Judgments are published by the People's Courts under Resolution 03/2017/NQ-HĐTP; the court partially anonymises party names before publication. The text can still contain sensitive personal information. Use at an aggregate/research level and comply with Vietnamese personal-data law.
10 · Giấy phép & trích dẫn · License & citation
🇻🇳 Bản án là văn bản công khai do Tòa án nhân dân tối cao công bố (Nghị quyết 03/2017/NQ-HĐTP). Bản phân phối lại này theo CC BY 4.0 kèm kỳ vọng: (1) ghi nhận nguồn congbobanan.toaan.gov.vn; (2) không trình bày bản trích xuất như văn bản pháp lý chính thức; (3) tôn trọng tính nhạy cảm của dữ liệu vụ án; (4) tuân thủ pháp luật Việt Nam. · 🇬🇧 Judgments are public documents published by the Supreme People's Court (Resolution 03/2017/NQ-HĐTP). This redistribution is CC BY 4.0 with the expectations that you (1) attribute congbobanan.toaan.gov.vn; (2) do not present the extraction as official legal text; (3) respect the sensitivity of case data; (4) comply with applicable Vietnamese law.
@misc{cbba_2026,
title = {Vietnamese Bản án Corpus (congbobanan.toaan.gov.vn)},
author = {TMQuan},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/tmquan/cbba-toaan-gov-vn}},
note = {Document-level mirror of published Vietnamese court judgments with 4096-D Nemotron-3-Embed-8B embeddings and shared-space (6-dataset joint) t-SNE / UMAP 2-D & 3-D projections over 1,542,247 natively-extractable documents.}
}
@misc{congbobanan_toaan_2026,
title = {Cổng công bố bản án, quyết định của Tòa án},
author = {{Công bố bản án — Tòa án nhân dân tối cao}},
year = {2026},
howpublished = {\url{https://congbobanan.toaan.gov.vn/}},
note = {Official judgment-publication portal of the Supreme People's Court of Vietnam. Judgments are public documents published under Resolution 03/2017/NQ-HDTP.}
}