temsa/OpenMed-PPSN-v5_1
OpenMed PPSN v5.1
Token classification model derived from OpenMed/OpenMed-PII-SuperClinical-Large-434M-v1 with B-PPSN / I-PPSN support for Irish PPSN detection.
Why v5.1
This iteration hardens PPSN behavior against known false positives on number-like strings (for example phone numbers and malformed ID-like tokens) while preserving non-PPSN behavior.
What this release contains
- Full model weights (
model.safetensors) with original OpenMed labels + PPSN labels. label_meta.jsonwith label mapping and provenance.- Eval artifacts:
eval_manual_irish_v5_1_large_v2_raw.jsoneval_hybrid_v5_1_large_v2_strict.jsonab_non_ppsn_v5_1.jsonqa_ppsn_regression_v6_validated.jsonleval_manual_qa_regression_v6_validated_v5_1_raw.jsoneval_hybrid_qa_regression_v6_validated_v5_1_strict.json
Key results
On irish_ppsn_eval_large_v2:
- Raw model-only PPSN performance:
P=0.8699 R=0.9956 F1=0.9285- Recommended strict hybrid mode (
--no-plausible-ppsn,--ppsn-min-score 0.6): P=1.0000 R=0.9997 F1=0.9999
Non-PPSN retention vs base OpenMed (ab_non_ppsn_v5_1.json):
- Synthetic non-PPSN F1 delta vs base:
+0.00049 - Real-set agreement F1 (candidate vs base):
1.0000 - Real entities per 1k chars delta vs base:
0.0000
Recommended production usage
Use strict hybrid PPSN post-processing (checksum-backed) for production masking. Raw model-only PPSN spans are less reliable on malformed numeric strings.
For a quick local smoke test of the packaged checkpoint, use the bundled word_aligned helper:
python3 inference_word_aligned.py \
--ppsn-min-score 0.4 \
--text "My PPSN is 1234567TW and I need help with my housing grant." \
--jsonInstall dependencies from pyproject.toml first if you are not already in an environment with transformers, torch, and regex.
License and attribution
- This derivative release is distributed under Apache-2.0, consistent with the base model license tag.
- Base model:
OpenMed/OpenMed-PII-SuperClinical-Large-434M-v1. - See
NOTICEfor attribution details.
<!-- portfolio-comparison:start -->
Portfolio Comparison
Updated: 2026-03-16.
Use this section for the fastest public comparison across the temsa PII masking portfolio.
- The first core table only includes public checkpoints that ship both comparable q8 accuracy and q8 CPU throughput.
- The first PPSN table only includes public artifacts that ship comparable PPSN accuracy and CPU throughput.
- Missing cells in the archive tables mean the older release did not ship that metric in its public bundle.
- DiffMask rows use the reconciled
clean_single_passharness that matches the deployed runtime. - GlobalPointer rows use the public raw-only span-matrix release bundle and its packaged q8 ONNX artifact.
- The same content is shipped as
PORTFOLIO_COMPARISON.mdinside each public model repo.
Irish Core PII: Comparable Public Checkpoints
Irish Core PII: Other Public Checkpoints
Finance-boundary q8 F1 is 1.0000 for OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc6, OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc7, OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc8, and all public IrishCore-DiffMask releases from rc1 to rc6. OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc5 ships 0.8750 on that public q8 suite.
PPSN-Only: Comparable Public Artifacts
PPSN-Only: Historical Public Checkpoints
If you need the strongest current raw-only Irish core model, start with IrishCore-GlobalPointer-135M-v1-rc4. If you need the fastest CPU-first raw-only line, compare it against IrishCore-DiffMask-135M-v1-rc6. If you need a PPSN-only artifact, compare the canonical fp32, fp16, and q8 variants of OpenMed-mLiteClinical-IrishPPSN-135M-v1 directly in the table above. <!-- portfolio-comparison:end -->
