CoolFace
Datasetpublic

CordwainerSmith/GolemGuard

GolemGuard: Hebrew Privacy Information Detection Corpus GolemGuard is a comprehensive Hebrew language dataset specifically designed for training and evaluating models for Personal Identifiable Information (PII) detection and masking. The dataset contains ~600MB of synthetic text data representing various document types and communication formats commonly found in Israeli professional and administrative contexts. Source Data Initial Data Collection and… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/GolemGuard.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes103downloads
10 commits on main
cb74a372y ago

Upload dataset_config.json with huggingface_hub

CordwainerSmith
ede9dc92y ago

Upload test.jsonl with huggingface_hub

CordwainerSmith
98276632y ago

Upload train.jsonl with huggingface_hub

CordwainerSmith
072767a2y ago

Upload GolemGuard.jsonl

CordwainerSmith
afe74772y ago

Update README.md

CordwainerSmith
b1085692y ago

Update README.md

CordwainerSmith
f137dc32y ago

Update README.md

CordwainerSmith
8e79bc62y ago

Update README.md

CordwainerSmith
b7006f92y ago

Update README.md

CordwainerSmith
880890f2y ago

initial commit

CordwainerSmith