KhalidAlharbi377/pii-detection-multisource-en-saudi-arabic
PII Detection Multisource EN + Saudi/Arabic 284,619 English examples. 2,088,335 labelled spans. 31 entity types. One label space. Four public PII datasets, merged into a single schema, plus Saudi and Arabic coverage that none of them had, plus material for two failure modes that matter when you run redaction in production. Built for OnKith, a privacy first voice assistant that transcribes speech and strips personal information on the device itself, before anything is allowed to… See the full description on the dataset page: https://huggingface.co/datasets/KhalidAlharbi377/pii-detection-multisource-en-saudi-arabic.
042
