CoolFace
Modelpublic

OpenMed/OpenMed-PII-BioClinicalBERT-Base-110M-v1

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
3likes6.2kdownloads
Model Card

OpenMed-PII-BioClinicalBERT-110M-v1

PII Detection Model | 110M Parameters | Open Source

![F1 Score]() ![Precision]() ![Recall]()

Model Description

OpenMed-PII-BioClinicalBERT-110M-v1 is a transformer-based token classification model fine-tuned for Personally Identifiable Information (PII) detection in text. This model identifies and classifies 54 types of sensitive information including names, addresses, SSNs, medical record numbers, and more.

Key Features

  • High Accuracy: Achieves strong F1 scores across diverse PII categories
  • Comprehensive Coverage: Detects 50+ entity types spanning personal, financial, medical, and contact information
  • Privacy-Focused: Designed for de-identification and compliance with HIPAA, GDPR, and other privacy regulations
  • Production-Ready: Optimized for real-world text processing pipelines

Performance

Evaluated on a stratified 2,000-sample test set from NVIDIA Nemotron-PII:

MetricScore
Micro F10.9437
Precision0.9449
Recall0.9426
Macro F10.9462
Weighted F10.9434
Accuracy0.9925

Top 10 PII Models

Best Performing Entities

EntityF1PrecisionRecallSupport
tax_id1.0001.0001.00043
ssn0.9960.9931.000141
biometric_identifier0.9961.0000.991232
email0.9950.9950.995757
date_of_birth0.9950.9891.000273

Challenging Entities

These entity types have lower performance and may benefit from additional post-processing:

EntityF1PrecisionRecallSupport
fax_number0.8700.8100.940100
time0.8640.8930.838468
sexuality0.8370.8090.86783
gender0.8150.7690.867188
occupation0.6390.6540.625717

Supported Entity Types

This model detects 54 PII entity types organized into categories:

<details> <summary><strong>Identifiers</strong> (16 types)</summary>

EntityDescription
account_numberAccount Number
api_keyApi Key
bank_routing_numberBank Routing Number
certificate_license_numberCertificate License Number
credit_debit_cardCredit Debit Card
cvvCvv
employee_idEmployee Id
health_plan_beneficiary_numberHealth Plan Beneficiary Number
mac_addressMac Address
medical_record_numberMedical Record Number
...and 6 more

</details>

<details> <summary><strong>Personal Info</strong> (14 types)</summary>

EntityDescription
ageAge
biometric_identifierBiometric Identifier
blood_typeBlood Type
date_of_birthDate Of Birth
education_levelEducation Level
first_nameFirst Name
last_nameLast Name
genderGender
languageLanguage
occupationOccupation
...and 4 more

</details>

<details> <summary><strong>Contact Info</strong> (4 types)</summary>

EntityDescription
emailEmail
phone_numberPhone Number
fax_numberFax Number
urlUrl

</details>

<details> <summary><strong>Location</strong> (6 types)</summary>

EntityDescription
cityCity
coordinateCoordinate
countryCountry
countyCounty
stateState
street_addressStreet Address

</details>

<details> <summary><strong>Network Info</strong> (3 types)</summary>

EntityDescription
device_identifierDevice Identifier
ipv4Ipv4
ipv6Ipv6

</details>

<details> <summary><strong>Temporal</strong> (3 types)</summary>

EntityDescription
dateDate
date_timeDate Time
timeTime

</details>

<details> <summary><strong>Organization</strong> (1 types)</summary>

EntityDescription
company_nameCompany Name

</details>

Usage

Quick Start

python
from transformers import pipeline

# Load the PII detection pipeline
ner = pipeline("ner", model="openmed/OpenMed-PII-BioClinicalBERT-110M-v1", aggregation_strategy="simple")

text = """
Patient John Smith (DOB: 03/15/1985, SSN: 123-45-6789) was seen today.
Contact: john.smith@email.com, Phone: (555) 123-4567.
Address: 456 Oak Street, Boston, MA 02108.
"""

entities = ner(text)
for entity in entities:
    print(f"{entity['entity_group']}: {entity['word']} (score: {entity['score']:.3f})")

De-identification Example

python
def redact_pii(text, entities, placeholder='[REDACTED]'):
    """Replace detected PII with placeholders."""
    # Sort entities by start position (descending) to preserve offsets
    sorted_entities = sorted(entities, key=lambda x: x['start'], reverse=True)
    redacted = text
    for ent in sorted_entities:
        redacted = redacted[:ent['start']] + f"[{ent['entity_group']}]" + redacted[ent['end']:]
    return redacted

# Apply de-identification
redacted_text = redact_pii(text, entities)
print(redacted_text)

Batch Processing

python
from transformers import AutoModelForTokenClassification, AutoTokenizer
import torch

model_name = "openmed/OpenMed-PII-BioClinicalBERT-110M-v1"
model = AutoModelForTokenClassification.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name)

texts = [
    "Contact Dr. Jane Doe at jane.doe@hospital.org",
    "Patient SSN: 987-65-4321, MRN: 12345678",
]

inputs = tokenizer(texts, return_tensors='pt', padding=True, truncation=True)
with torch.no_grad():
    outputs = model(**inputs)
    predictions = torch.argmax(outputs.logits, dim=-1)

Training Details

Dataset

  • Source: NVIDIA Nemotron-PII
  • Format: BIO-tagged token classification
  • Labels: 106 total (53 entity types × 2 BIO tags + O)
  • Splits: 50K train / 5K validation / 45K test

Training Configuration

  • Max Sequence Length: 384 tokens
  • Label Strategy: First token only (label_all_tokens=False)
  • Framework: Hugging Face Transformers + Trainer API

Intended Use & Limitations

Intended Use

  • De-identification: Automated redaction of PII in clinical notes, medical records, and documents
  • Compliance: Supporting HIPAA, GDPR, and privacy regulation compliance
  • Data Preprocessing: Preparing datasets for research by removing sensitive information
  • Audit Support: Identifying PII in document collections

Limitations

⚠️ Important: This model is intended as an assistive tool, not a replacement for human review.

  • False Negatives: Some PII may not be detected; always verify critical applications
  • Context Sensitivity: Performance may vary with domain-specific terminology
  • Challenging Categories: occupation, time, and sexuality have lower F1 scores
  • Language: Primarily trained on English text

Citation

bibtex
@misc{openmed-pii-2026,
  title = {OpenMed-PII-BioClinicalBERT-110M-v1: PII Detection Model},
  author = {OpenMed Science},
  year = {2026},
  publisher = {Hugging Face},
  url = {https://huggingface.co/openmed/OpenMed-PII-BioClinicalBERT-110M-v1}
}

Links