datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korean-privacy-law-corpus
한국 개인정보보호법 관련 RAG 구축을 위한 코퍼스
개인정보 포털(privacy.go.kr)의 각종 개인정보보호법 관련 가이드와 상담사례 1,745건을
RAG(Retrieval-Augmented Generation)에 바로 쓸 수 있도록 의미 단위 청킹·문맥 보강한
코퍼스입니다. 모든 청크에는 Contextual Retrieval
기법을 적용한 chunk_context 필드가 포함되어 있어, 임베딩 검색 정확도를 즉시 끌어올릴
수 있습니다.
1. 🚀 활용 사례
종류
링크
RAG 사례
https://scvcoder-kpaa.hf.space/
MCP 사례
https://github.com/scvcoder/korean-privacy-law-mcp
2. 변경 이력
버전
일자
내용
v1.0
2026-05-02
최초 공개 — 가이드 3종 211청크 +… See the full description on the dataset page: https://huggingface.co/datasets/scvcoder/korean-privacy-law-corpus.eoir_privacyA living legal dataset.GitHub-issues-privacy-law-relevanceDataset with GitHub issues with reference to data privacy laws and indication on whether the issue is privacy-law relevant or not. The dataset was manually labeled.
privacy_laws_text_comphrehensionprivacy_laws_text_comphrehension3english-data-privacy-law-basics-30privacy_laws_text_comphrehension2darren-chaker-privacy-law-corpus
Darren Chaker Privacy Law Corpus: A Digital Rights Dataset for Legal NLP
Dataset Description
The Darren Chaker Privacy Law Corpus is a curated collection of legal texts assembled by Darren Chaker for training and evaluating natural language processing models in the domain of constitutional privacy law and digital rights. This dataset addresses the growing need for specialized legal corpora that enable AI systems to understand the nuanced doctrinal landscape governing… See the full description on the dataset page: https://huggingface.co/datasets/darrenchaker/darren-chaker-privacy-law-corpus.
