tmquan/thuvienphapluat-vn-hdpl
Hỏi đáp pháp luật Việt Nam — Vietnamese Legal Q&A VI — Bộ dữ liệu 109.543 bài hỏi đáp pháp luật Việt Nam từ thuvienphapluat.vn/hoi-dap-phap-luat: câu hỏi ↔ câu trả lời dài có trích dẫn văn bản, kèm embedding Nemotron-3-Embed-8B và toạ độ giảm chiều trong không gian chung của 6 bộ dữ liệu pháp luật. EN — 109,543 Vietnamese legal question-and-answer articles from thuvienphapluat.vn, each pairing a legal question with a citation-grounded long-form answer, embedded with… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/thuvienphapluat-vn-hdpl.
Refresh global corpus map (now 7 datasets incl. htpl)
Refresh global corpus map (now 7 datasets incl. htpl)
viz: per-dataset UMAP/t-SNE scatter from shared 6-dataset reduce (by area); drop masked corpus plots
Add masked shared-corpus-space plots (this dataset within the 6-dataset corpus) + README subsection
Unify to common-corpus standard: documents/embeddings/reduces (+pages); shared-space t-SNE/UMAP coords, global maps; drop nemotron-1b, old reduce, PCA
Split raw crawl HTML out of qa into its own pages table (unfold page_html_gz)
Fold pages into qa (page_html_gz column); shard qa to 30x<150MB; drop redundant plain embeddings config
Standardize to Nemotron-8B/1B (drop Qwen3 configs + parquets); filter reduce to Nemotron; add embeddings/reduce schema tables to the card
Restructure: clean qa (no reductions) + 4-model per-Q/A embeddings (Nemotron-3-Embed-8B/1B, Qwen3-Embedding-8B/0.6B) + long reduce.parquet (per-field 2D+3D PCA/t-SNE/UMAP)
hdpl Q&A: 109,543 pairs (added ~7.8k low-id + edge)
v8: question_embedding recomputed on the real detailed question (was the title); question/summary columns + coords + scatters corrected
fix: swap question<->summary columns (question is the detailed query, summary the title); embeddings/coords recompute to follow
v7.7: re-chunk parquets to small row groups + page index (fix HF dataset viewer TooBigContentError)
v7.6: tether alpha 0.005
v7.5: tether alpha 0.006 (balanced — flows visible, clusters discernible)
v7.4: tether lines much more transparent (alpha 0.004, thinner) — clusters foreground, flows as haze
v7.3: more-transparent tether lines (alpha 0.01)
v7.2: all Q&A tethered (opacity 1/N) + inverse-alpha Sankey ribbons
v7.1: add pages config (raw source html.gz); tether lines in front; all 27 areas in legend + card
hdpl Q&A v7: 101,724 rich Q&A + Nemotron-3-Embed-8B question/answer embeddings + PCA/t-SNE/UMAP projections + area→law Sankey
UMAP: legend outside plot, single column, full labels — figure fully visible
v6: 25,381 Q&A with REAL legal categories (breadcrumb backfill) + re-colored UMAP
v5: 25,265 clean Q&A (home-IP crawl) + refreshed 8B UMAP
v4: 16,002 clean Q&A (Mac-exit crawl) + refreshed 8B UMAP
v3: 9,346 clean Q&A (boilerplate-stripped answers) + refreshed 8B UMAP
Card: dual citation section + larger UMAP legend fonts
v2: 2571 Q&A (deduped) + refreshed 8B category-coloured UMAP + card
UMAP: colour by legal category (grey source feed)
Data card: embed Q<->A UMAP figure + embedding-map section
Refresh 8B Q<->A UMAP figure on full 2105 pairs
Vietnamese legal Q&A (hoi-dap-phap-luat) crawl
initial commit
