CoolFace
Modelpublic

temsa/IrishCore-DiffMask-135M-v1-rc3

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes22downloads
Model Card

IrishCore-DiffMask-135M-v1-rc3

IrishCore-DiffMask-135M-v1-rc3 is a raw-only Irish PII masking model derived from OpenMed/OpenMed-PII-mLiteClinical-Base-135M-v1.

It is a small, scanner-free span extractor tuned for:

  • —PPSN
  • —ACCOUNT_NUMBER
  • —BANK_ROUTING_NUMBER
  • —CREDIT_DEBIT_CARD
  • —PASSPORT_NUMBER
  • —POSTCODE
  • —PHONE_NUMBER
  • —EMAIL
  • —FIRST_NAME
  • —LAST_NAME
  • —SWIFT_BIC

The main target is English plus Irish Gaelic text in citizen-support, public-sector, and HSE-style flows. The repo ships both the full transformers checkpoint and a dynamic q8 ONNX artifact for CPU deployment.

What "DiffMask" Means Here

This release is not a generative diffusion language model. It is a compact discriminative token-span model trained with a diffusion-style denoising schedule.

The short version:

  • —Base OpenMed: plain BIO token classification
  • —DiffMask: token-span extraction with token-presence and boundary heads
  • —DiffMask training: repeated masked denoising over the same sentence
  • —DiffMask inference: one forward pass, no iterative refinement, no text generation

Concretely:

  • —The encoder starts from the DistilBERT-family weights inside OpenMed/OpenMed-PII-mLiteClinical-Base-135M-v1.
  • —The model adds three task heads over the encoder hidden states:
  • —a per-label token-presence head
  • —a typed start-boundary head
  • —a typed end-boundary head
  • —During training, each input sentence is corrupted multiple times by replacing a random fraction of visible tokens with [MASK].
  • —The corruption level follows a short noise schedule from heavy masking to light masking.
  • —The same gold spans are learned at every noise level, and the losses are averaged across the denoising passes.
  • —At inference time there is no diffusion loop and no rewrite step: the model runs once and a score-only span decoder reconstructs spans from token scores plus typed boundaries.

So the "DLLM" aspect here is the training recipe: repeated masked denoising over text, not autoregressive generation.

What It Is Not

This model is not a full discrete diffusion language model in the LLaDA sense.

A true DLLM would usually have:

  • —timestep or noise conditioning inside the model
  • —iterative denoising at inference time
  • —multi-step sequence refinement at runtime
  • —text generation or full-sequence reconstruction as a first-class objective

This release does not do that.

Instead, it uses the diffusion idea only as a training-time robustness trick:

  • —corrupt the sentence with [MASK] at several noise levels
  • —train on the same target spans each time
  • —average those losses

At runtime, it behaves like a normal fast discriminative extractor.

Architecture

  • —Encoder: DistilBERT-size encoder from the OpenMed mLiteClinical 135M base
  • —Heads:
  • —token presence per released label
  • —typed start boundary per released label
  • —typed end boundary per released label
  • —Decoder:
  • —score-only span decoding from offsets, token continuity, label-specific thresholds, and typed boundaries
  • —no regex candidate extractor
  • —no checksum validator
  • —no scanner layer

The release behavior is fully defined by the weights plus the bundled decoder in common.py.

Training And Inference Flow

Training:

  1. 1.tokenize a sentence with gold BIO spans
  2. 2.convert spans into:
  3. 3.token-presence targets
  4. 4.typed start targets
  5. 5.typed end targets
  6. 6.create several noised copies of the same tokenized sentence by masking random visible tokens
  7. 7.run the same encoder+heads on each noised copy
  8. 8.average the losses across those denoising passes

Inference:

  1. 1.tokenize the raw text once
  2. 2.run a single forward pass
  3. 3.predict:
  4. 4.which labels are present on each token
  5. 5.where each labeled span starts
  6. 6.where each labeled span ends
  7. 7.decode spans with label-aware thresholds and boundary rules
  8. 8.replace the detected spans with placeholders such as [PII:PPSN]

There is no multi-step refinement loop in deployment.

How It Differs From The Original OpenMed Model

The original OpenMed/OpenMed-PII-mLiteClinical-Base-135M-v1 is a standard DistilBertForTokenClassification model:

  • —one encoder
  • —one token-classification head
  • —BIO labels such as B-email, I-email, B-phone_number
  • —generic token aggregation to recover spans

DiffMask changes two things:

  1. 1.Different supervision
  2. 2.base OpenMed learns only BIO token labels
  3. 3.DiffMask learns token presence plus typed span boundaries
  1. 1.Different training recipe
  2. 2.base OpenMed is trained as a standard token classifier
  3. 3.DiffMask is trained on multiple masked-noised views of the same sentence

That makes DiffMask better suited to structured Irish identifiers and mixed PII masking, while still keeping a small encoder and a fast CPU path.

How It Differs From rc5 And rc8

ModelCore ideaExternal scanner/validatorRuntime shape
rc5token classifier + repair logicyesheavier, decoder-assisted
rc8raw-only token-span modelnoone pass + span decoder
DiffMaskraw-only token-span model + denoising trainingnoone pass + span decoder

So DiffMask is closest to rc8 operationally, but it uses a stronger training recipe.

Why This Exists

The older rc5 release still depended on a repair-oriented decoder stack. The public rc8 release removed that external logic, but it regressed on several structured Irish identifiers. This release keeps the raw-only deployment shape while re-hardening the model on Irish numeric and mixed-PII cases.

rc3 is the next candidate after rc2. It keeps the stronger focusv3 checkpoint selected during local iteration, then applies a small decoder-profile retune for the published config:

  • —lower EMAIL token extend threshold to keep contiguous mailbox fragments together
  • —lower PASSPORT_NUMBER q8 threshold slightly to recover a mixed-message passport miss after dynamic quantization

The weights remain raw-only and scanner-free. The rc3 change is the checkpoint plus a stricter release-time decoder profile in config.json.

References

Direct implementation references:

  • —Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding https://arxiv.org/abs/1810.04805
  • —Sanh et al., DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter https://arxiv.org/abs/1910.01108
  • —Fu et al., Boundary Smoothing for Named Entity Recognition https://aclanthology.org/2022.acl-long.490/
  • —Wang et al., SPANNER: Named Entity Re-/Recognition as Span Prediction https://aclanthology.org/2021.acl-long.558/

Conceptual diffusion-style training references:

  • —Nie et al., LLaDA 2.0: Scaling Up Diffusion Language Models to 100B https://arxiv.org/abs/2512.15745
  • —Gong et al., Scaling Diffusion Language Models via Adaptation from Autoregressive Models https://arxiv.org/abs/2410.17891

These diffusion papers were used as architectural inspiration for the masked noising schedule. This release does not implement a generative text diffusion runtime.

Included Artifacts

  • —Full transformers checkpoint in the repo root
  • —Dynamic q8 ONNX export in onnx/model_quantized.onnx
  • —Unquantized ONNX export in onnx/model.onnx
  • —inference_mask.py for the full checkpoint
  • —inference_mask_onnx.py for the ONNX q8 path
  • —common.py, model.py, and multitask_model.py implementing the release decoder
  • —benchmark files in eval/

Artifact sizes:

  • —Full checkpoint: 514 MB (model.safetensors)
  • —Dynamic q8 ONNX: 393 MB (onnx/model_quantized.onnx)

How To Use It

Full checkpoint:

bash
uv run python inference_mask.py \
  --model temsa/IrishCore-DiffMask-135M-v1-rc3 \
  --min-score 0.5 \
  --text "My PPSN is 1234567TW, my Eircode is D02 X285, and my phone is 087 123 4567." \
  --json

Dynamic q8 ONNX:

bash
uv run python inference_mask_onnx.py \
  --model temsa/IrishCore-DiffMask-135M-v1-rc3 \
  --min-score 0.5 \
  --text "Please provide your passport NN5123456 and call me on 0851234567." \
  --json

Both scripts emit explicit placeholders like [PII:PPSN] in masked_text.

Q8 Comparison

Deployment-relevant comparison on CPU:

ModelCore F1Edge F1Finance F1Finance-boundary F1User PPSN F1GA weak PPSN F1Multilingual PPSN F1Hardening F1
rc5 ONNX q80.96690.97440.93620.87501.00001.00000.9333-
rc8 ONNX q80.97371.00001.00001.00001.00001.00000.91760.7059
IrishCore-DiffMask-135M-v1-rc3 ONNX q80.96641.00001.00001.00000.85711.00000.95911.0000

UAT replay exact suite used for the recent hardening pass:

ModelUAT replay exact F1PrecisionRecall
IrishCore-DiffMask-135M-v1-rc1 ONNX q80.45451.00000.2941
IrishCore-DiffMask-135M-v1-rc2 ONNX q80.82761.00000.7059
rc8 ONNX q80.36360.37500.3529
IrishCore-DiffMask-135M-v1-rc3 ONNX q80.90321.00000.8235

CPU throughput references:

Suite`rc5` q8`rc8` q8`IrishCore-DiffMask-135M-v1-rc3` q8
Irish core short-text path33.6193 ex/s257.3756 ex/s29.9676 ex/s
Multilingual PPSN short-text path35.5561 ex/s230.5181 ex/s54.2219 ex/s
Runtime profile source23.8338 ex/s179.4708 ex/s46.1519 ex/s

Notes:

  • —The rc5 speed references come from its published q8 end-to-end inference stack, which includes its older repair decoder.
  • —The rc8 and IrishCore-DiffMask-135M-v1-rc3 numbers use the same raw-only token-span ONNX path.
  • —A weight-only q4 ONNX experiment was also tried during development, but it was slower than q8 on this CPU and is not shipped.
  • —The user_raw_regression_cases_v1 suite is a legacy PPSN-only regression set. In rc3, the single counted false positive is 0871234567, which is now intentionally masked as PHONE_NUMBER rather than misread as PPSN.

Additional Training Data Used For This RC

Published training sources:

  • —temsa/OpenMed-Irish-CorePII-TrainMix-v1
  • —temsa/OpenMed-Irish-PPSN-Eircode-Spec-v1
  • —joelniklaus/mapa
  • —gretelai/synthetic_pii_finance_multilingual

Additional local synthetic hardening and replay sets used during checkpoint selection:

  • —irish_core_diffmask_v5_mix
  • —dllm_uat_replay_v1
  • —dllm_gap_patch_v4
  • —dllm_uat_patch_v3
  • —irish_core_diffmask_focus_v3

rc3 is based on the locally selected focusv3 checkpoint and then retuned with a narrower decoder profile for the public config.

Limits

  • —This is still a compact model. The hardest remaining errors are multilingual PPSN near-miss cases rather than Irish core numeric formats.
  • —The release path is intentionally scanner-free. If you need deterministic validation of individual identifier types, add that in your application layer.
  • —If you rely on release behavior, use the bundled inference scripts or import decode_token_presence_segments from common.py.
  • —Known remaining misses on the current UAT replay suite are the second phone number in the long Client Identity Services sentence (071 967 2616), R93 EC57 inside the longer allocation-centre block, and EPStamp4@enterprise.gov.ie.

License And Attribution

  • —Release license: Apache-2.0
  • —Base model: OpenMed/OpenMed-PII-mLiteClinical-Base-135M-v1
  • —The derivative release remains subject to the attribution terms of the upstream datasets listed above.
  • —See NOTICE, training_sources.json, and eval/benchmark_summary.json for provenance and benchmark details.

<!-- portfolio-comparison:start -->

Portfolio Comparison

Updated: 2026-03-16.

Use this section for the fastest public comparison across the temsa PII masking portfolio.

  • —The first core table only includes public checkpoints that ship both comparable q8 accuracy and q8 CPU throughput.
  • —The first PPSN table only includes public artifacts that ship comparable PPSN accuracy and CPU throughput.
  • —Missing cells in the archive tables mean the older release did not ship that metric in its public bundle.
  • —DiffMask rows use the reconciled clean_single_pass harness that matches the deployed runtime.
  • —GlobalPointer rows use the public raw-only span-matrix release bundle and its packaged q8 ONNX artifact.
  • —The same content is shipped as PORTFOLIO_COMPARISON.md inside each public model repo.

Irish Core PII: Comparable Public Checkpoints

RepoStackFull Core F1Q8 Core F1Q8 Multilingual PPSN F1Q8 Core ex/s
`temsa/IrishCore-GlobalPointer-ContextPII-4L-122M-v1-rc4`4-layer GlobalPointer distilled fast student1.00001.00000.9333299.0
`temsa/IrishCore-GlobalPointer-ContextPII-4L-122M-v1-rc3`4-layer GlobalPointer distilled fast student1.00001.00000.9333317.9
`temsa/IrishCore-GlobalPointer-ContextPII-4L-122M-v1-rc2`4-layer GlobalPointer distilled fast student1.00001.00000.9333292.5
`temsa/IrishCore-GlobalPointer-ContextPII-4L-122M-v1-rc1`4-layer GlobalPointer distilled fast student1.00001.00000.9333337.3
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc27`GlobalPointer raw-only + context labels1.00001.00000.9333270.0
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc25`GlobalPointer raw-only + context labels1.00001.00000.9333212.1
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc24`GlobalPointer raw-only + context labels1.00001.00000.9333278.9
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc23`GlobalPointer raw-only + context labels1.00001.00000.9333237.6
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc22`GlobalPointer raw-only + context labels1.00001.00000.9333106.8
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc21`GlobalPointer raw-only + context labels1.00001.00000.9333150.8
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc20`GlobalPointer raw-only + context labels1.00001.00000.9333181.9
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc19`GlobalPointer raw-only + context labels1.00001.00000.933373.1
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc18`GlobalPointer raw-only + context labels1.00001.00000.9333126.2
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc17`GlobalPointer raw-only + context labels1.00001.00000.9333125.5
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc16`GlobalPointer raw-only + context labels1.00001.00000.9333125.5
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc15`GlobalPointer raw-only + context labels1.00001.00000.9333125.5
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc14`GlobalPointer raw-only + context labels1.00001.00000.9333119.2
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc13`GlobalPointer raw-only + context labels1.00001.00000.9333126.1
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc12`GlobalPointer raw-only + context labels1.00001.00000.933373.6
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc11`GlobalPointer raw-only + context labels1.00001.00000.933394.1
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc10`GlobalPointer raw-only + context labels1.00001.00000.9333125.8
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc9`GlobalPointer raw-only + context labels1.00001.00000.9333119.8
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc8`GlobalPointer raw-only + context labels1.00001.00000.9333128.9
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc7`GlobalPointer raw-only + context labels1.00001.00000.933389.0
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc6`GlobalPointer raw-only + context labels1.00001.00000.933389.0
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc5`GlobalPointer raw-only + context labels1.00001.00000.933384.5
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc4`GlobalPointer raw-only + context labels0.99350.99350.933361.5
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc3`GlobalPointer raw-only + context labels0.99350.99350.933361.5
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc2`GlobalPointer raw-only + context labels0.99350.99350.922261.5
`temsa/IrishCore-GlobalPointer-ContextPII-135M-v1-rc1`GlobalPointer raw-only + context labels0.99350.99350.922261.5
`temsa/IrishCore-GlobalPointer-135M-v1-rc4`GlobalPointer raw-only span-matrix1.00001.00000.9333221.6
`temsa/IrishCore-GlobalPointer-135M-v1-rc3`GlobalPointer raw-only span-matrix1.00001.00000.9213204.9
`temsa/IrishCore-GlobalPointer-135M-v1-rc2`GlobalPointer raw-only span-matrix0.99340.99340.9326231.2
`temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc8`Raw-only token-span0.97370.97370.917646.1
`temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc7`Hybrid classifier + generated scanner spec1.00000.99341.000030.0
`temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc6`Hybrid classifier + repair decoders1.00000.99341.000029.5
`temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc5`Hybrid classifier + repair decoders0.97370.96690.933334.4
`temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc4`Hybrid classifier + repair decoders0.98700.97400.9600114.2
`temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc3`Hybrid classifier + repair decoders0.98060.96770.933344.9
`temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc2`Hybrid classifier + repair decoders0.95540.96150.7887119.1
`temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v1`Hybrid classifier baseline0.95300.93330.9882103.3
`temsa/IrishCore-DiffMask-135M-v1-rc6`DiffMask token-span, scanner-free0.98010.97330.9274130.3
`temsa/IrishCore-DiffMask-135M-v1-rc5`DiffMask token-span, scanner-free0.97330.97330.9379249.2
`temsa/IrishCore-DiffMask-135M-v1-rc4`DiffMask token-span, scanner-free0.97330.97330.937129.5
`temsa/IrishCore-DiffMask-135M-v1-rc3`DiffMask token-span, scanner-free0.96640.96640.959130.0
`temsa/IrishCore-DiffMask-135M-v1-rc2`DiffMask token-span, scanner-free0.96640.96640.9212247.1
`temsa/IrishCore-DiffMask-135M-v1-rc1`DiffMask token-span, scanner-free0.98010.99340.9412251.2

Irish Core PII: Other Public Checkpoints

RepoStackFull Core F1Q8 Core F1Q8 Multilingual PPSN F1Notes
`temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc1`Hybrid classifier prototype0.9487——Predates the public q8 artifact.

Finance-boundary q8 F1 is 1.0000 for OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc6, OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc7, OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc8, and all public IrishCore-DiffMask releases from rc1 to rc6. OpenMed-mLiteClinical-IrishCorePII-135M-v2-rc5 ships 0.8750 on that public q8 suite.

PPSN-Only: Comparable Public Artifacts

RepoArtifactIrish Large F1Multilingual PPSN F1User Raw F1QA v8 F1CPU ex/s
`temsa/OpenMed-mLiteClinical-IrishPPSN-135M-v1`fp32 canonical checkpoint0.89790.97040.80000.738557.4
`temsa/OpenMed-mLiteClinical-IrishPPSN-135M-v1-fp16`fp16 CPU/GPU artifact—0.97040.80000.738545.8
`temsa/OpenMed-mLiteClinical-IrishPPSN-135M-v1-q8`dynamic int8 CPU artifact—0.9040——132.1

PPSN-Only: Historical Public Checkpoints

RepoMain Published MetricsNotes
`temsa/OpenMed-PPSN-mLiteClinical-v1`same as canonical fp32 repo: multilingual 0.9704, user raw 0.8000Legacy alias; prefer temsa/OpenMed-mLiteClinical-IrishPPSN-135M-v1.
`temsa/OpenMed-PPSN-v6-raw-rc2`irishregv5 0.8750; userraw 0.8000; qav8 0.7385Raw PPSN-only research checkpoint; no packaged multilingual CPU benchmark row.
`temsa/OpenMed-PPSN-v5_1`irishlargev2 raw 0.9285; qa_v6 hybrid strict 1.0000Hybrid PPSN-only checkpoint; predates the canonical multilingual suite packaging.
`temsa/OpenMed-PPSN-v5`irishregv5 raw 0.8235; irishregv5 hybrid strict 1.0000Hybrid PPSN-only checkpoint; predates the canonical multilingual suite packaging.
`temsa/OpenMed-PPSN-v4`synthetic non-PPSN drift check onlyPredates the current PPSN eval suite; no packaged apples-to-apples multilingual CPU row.

If you need the strongest current raw-only Irish core model, start with IrishCore-GlobalPointer-135M-v1-rc4. If you need the fastest CPU-first raw-only line, compare it against IrishCore-DiffMask-135M-v1-rc6. If you need a PPSN-only artifact, compare the canonical fp32, fp16, and q8 variants of OpenMed-mLiteClinical-IrishPPSN-135M-v1 directly in the table above. <!-- portfolio-comparison:end -->