CordwainerSmith/GolemGuard
GolemGuard: Hebrew Privacy Information Detection Corpus GolemGuard is a comprehensive Hebrew language dataset specifically designed for training and evaluating models for Personal Identifiable Information (PII) detection and masking. The dataset contains ~600MB of synthetic text data representing various document types and communication formats commonly found in Israeli professional and administrative contexts. Source Data Initial Data Collection and… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/GolemGuard.
Upload dataset_config.json with huggingface_hub
Upload test.jsonl with huggingface_hub
Upload train.jsonl with huggingface_hub
Upload GolemGuard.jsonl
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
initial commit
