CoolFace
Modelpublic

onnx-community/eu-pii-safeguard-ONNX

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes16downloads
Model Card

eu-pii-safeguard (ONNX)

This is an ONNX version of tabularisai/eu-pii-safeguard. It was automatically converted and uploaded using this Hugging Face Space.

Usage with Transformers.js

See the pipeline documentation for token-classification: https://huggingface.co/docs/transformers.js/api/pipelines#module_pipelines.TokenClassificationPipeline


<p align="center"> <img src="https://tabularis.ai/piitoppic.png" alt="ModernFinBERT" width="512"> </p>

<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/Discord%20button.png" width="200"/>

Multilingual Privacy Filter

Multilingual PII Detection Model for European Languages

A state-of-the-art multilingual model for detecting Personally Identifiable Information (PII) across 26 European languages (all EU official languages). It is designed for GDPR compliance, privacy-preserving AI applications, and secure handling of sensitive data in multilingual settings. This model enables enterprises, researchers, and data protection teams to identify and safeguard PII with high accuracy (โ‰ˆ98%) across diverse European contexts.

๐ŸŽฏ Model Performance

  • โ€”Global F1 Score: 97.02%
  • โ€”26 Languages Supported
  • โ€”42 PII Entity Types
  • โ€”Consistent 95%+ F1 across all languages

๐ŸŒ Supported Languages

๐Ÿ‡ง๐Ÿ‡ฌ Bulgarian โ€ข ๐Ÿ‡จ๐Ÿ‡ฟ Czech โ€ข ๐Ÿ‡ฉ๐Ÿ‡ฐ Danish โ€ข ๐Ÿ‡ฉ๐Ÿ‡ช German โ€ข ๐Ÿ‡ฌ๐Ÿ‡ท Greek โ€ข ๐Ÿ‡ฌ๐Ÿ‡ง English โ€ข ๐Ÿ‡ช๐Ÿ‡ธ Spanish โ€ข ๐Ÿ‡ช๐Ÿ‡ช Estonian โ€ข ๐Ÿ‡ซ๐Ÿ‡ฎ Finnish โ€ข ๐Ÿ‡ซ๐Ÿ‡ท French โ€ข ๐Ÿ‡ฎ๐Ÿ‡ช Irish โ€ข ๐Ÿ‡ญ๐Ÿ‡ท Croatian โ€ข ๐Ÿ‡ญ๐Ÿ‡บ Hungarian โ€ข ๐Ÿ‡ฎ๐Ÿ‡น Italian โ€ข ๐Ÿ‡ฑ๐Ÿ‡น Lithuanian โ€ข ๐Ÿ‡ฑ๐Ÿ‡ป Latvian โ€ข ๐Ÿ‡ฒ๐Ÿ‡น Maltese โ€ข ๐Ÿ‡ณ๐Ÿ‡ฑ Dutch โ€ข ๐Ÿ‡ต๐Ÿ‡ฑ Polish โ€ข ๐Ÿ‡ต๐Ÿ‡น Portuguese โ€ข ๐Ÿ‡ท๐Ÿ‡ด Romanian โ€ข ๐Ÿ‡ท๐Ÿ‡บ Russian โ€ข ๐Ÿ‡ธ๐Ÿ‡ฐ Slovak โ€ข ๐Ÿ‡ธ๐Ÿ‡ฎ Slovenian โ€ข ๐Ÿ‡ธ๐Ÿ‡ช Swedish โ€ข ๐Ÿ‡บ๐Ÿ‡ฆ Ukrainian

๐Ÿ” Detected PII Types

  • โ€”Personal: First/Last/Middle Names, Age, Gender, Ethnicity
  • โ€”Contact: Email, Phone, Address, City, Country, Postal Code
  • โ€”Financial: Credit Card, IBAN, Account Numbers, Salary
  • โ€”Identity: National ID, Passport, Driver License, Tax ID
  • โ€”Health: Medical Conditions, Health Insurance ID
  • โ€”Digital: IP Address, MAC Address, URL, Username, Password
  • โ€”And more: 42 total entity types

๐Ÿš€ Quick Start

python
from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch

# Load model and tokenizer
model_name = "tabularisai/eu-pii-safeguard"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)

# Example text (French)
text = "Bonjour, je suis Marie Dubois, email: marie@company.fr"

# Tokenize and predict
inputs = tokenizer(text, return_tensors="pt", truncation=True)
with torch.no_grad():
    outputs = model(**inputs)
    predictions = torch.argmax(outputs.logits, dim=-1)

# Get predictions
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
predicted_labels = [model.config.id2label[pred.item()] for pred in predictions[0]]

print("Detected PII:")
for token, label in zip(tokens, predicted_labels):
    if label != "O":
        print(f"  {label}: {token}")

๐Ÿ“Š Performance by Language

LanguageF1 ScoreLanguageF1 Score
Irish (ga)97.98%Dutch (nl)97.24%
Bulgarian (bg)97.80%Slovak (sk)97.21%
Italian (it)97.68%Swedish (sv)97.09%
Portuguese (pt)97.61%Russian (ru)97.04%
Slovenian (sl)97.51%Croatian (hr)96.93%
Czech (cs)97.51%Polish (pl)96.63%
Hungarian (hu)97.50%French (fr)96.59%
Estonian (et)97.41%Romanian (ro)96.54%
Latvian (lv)97.40%Danish (da)96.36%
English (en)97.36%German (de)96.22%
Spanish (es)97.34%Ukrainian (uk)96.09%
Finnish (fi)97.30%Maltese (mt)95.78%
Lithuanian (lt)97.24%Greek (el)95.42%

๐Ÿ’ผ Use Cases

  • โ€”๐Ÿ”’ Data Privacy: Automatically detect and anonymize PII before processing
  • โ€”โš–๏ธ GDPR Compliance: Ensure regulatory compliance across EU markets
  • โ€”๐Ÿ›ก๏ธ Security: Prevent data breaches by identifying sensitive information
  • โ€”๐Ÿ“Š Data Governance: Audit and catalog personal data in multilingual datasets

๐Ÿ—๏ธ Model Architecture

  • โ€”Base Model: XLM-RoBERTa-large
  • โ€”Task: Token Classification
  • โ€”Labels: 74 (B-/I- format for 42 entity types)
  • โ€”Max Length: 256 tokens

๐Ÿ”„ Community Feedback

We're actively seeking feedback from the community! Please:

  • โ€”๐Ÿ› Report issues or edge cases
  • โ€”๐Ÿ’ก Suggest improvements
  • โ€”๐Ÿงช Share your use cases and results
  • โ€”๐Ÿ“Š Contribute evaluation on new datasets

๐Ÿข About Tabularis AI

Developed by Tabularis AI - Building privacy-preserving AI solutions for enterprise data protection.


For questions, collaborations, or licensing inquiries: info@tabularis.ai