tmquan/cbba-toaan-gov-vn
Vietnamese Bản án Corpus — congbobanan.toaan.gov.vn 🇻🇳 Tóm tắt. Bản án sơ thẩm/phúc thẩm/giám đốc thẩm/tái thẩm của Việt Nam, thu thập từ cổng công bố bản án congbobanan.toaan.gov.vn của Tòa án nhân dân tối cao. Ba cấu hình HF khoá theo doc_name/id: documents (nội dung siêu dữ liệu + trích dẫn), embeddings (vector 4096-D Nemotron-3-8B), reduces (toạ độ t-SNE/UMAP trong không gian chung 6 bộ dữ liệu). Tên cột và giá trị phân loại bằng tiếng Anh; chỉ nội dung pháp lý giữ tiếng… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/cbba-toaan-gov-vn.
Refresh global corpus map (now 7 datasets incl. htpl)
Refresh global corpus map (now 7 datasets incl. htpl)
viz: category scatters from shared 6-dataset reduce coords; drop masked corpus plots
Add masked shared-corpus-space plots (UMAP + t-SNE) to Visualizations
Unify to ViLA common-corpus standard: rename embeddings-nemotron-8b -> embeddings; replace wide reduce with shared-space `reduces` (tsne/umap x 2D/3D, no PCA); drop PCA plot; add global corpus maps; regen category scatters; rewrite card
Phase B: rename config embed -> embeddings-nemotron-8b (atomic server-side copy+delete); fix README frontmatter
documents: add structure-parser citations (citations, parent_acts, num_citations/articles/khoan/diem)
Redesign datacard to ViLA standard (per-config schema tables + Nemotron-8B/1B embedding standard)
Token-aware domain matching: LHST=divorce not Criminal, KHCN not Administrative; +49k classified
Normalize code to canonical hyphenated form (HSST -> HS-ST); 91,598 rows
Dual citation (HF redistribution + source portal), matching anle
Vietnamese bản án corpus with sentence-level structure + embedding + reduce layers
initial commit
