CoolFace
Modelpublic

DT4H/CardioBERTa.en_P_translations_only

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes12downloads
Model Card

DT4HCardioBERTaparentsentranslations_only

DT4H_CardioBERTa_parents_en_translations_only is a English biomedical terminology encoder for clinical concept normalization and entity linking. It is initialized from [DT4H/CardioBERTa.en] and specialized using CUI-supervised terminology pairs and metric learning.

Backbone

The backbone belongs to the CardioBERTa family from CardioLM - a multilingual suite of small language models for the cardiology domain. CardioBERTa comprises language-specific encoder models adapted to cardiology through continued pretraining on monolingual biomedical and cardiology-related corpora using Masked Language Modeling (MLM). The family covers Czech, Dutch, English, Italian, Romanian, Spanish and Swedish.

Training

LanguageEnglish (en)
Triplet collectiontranslations_only
Strategyparents
ObjectiveMulti-Similarity Loss
MiningAll triplets, margin 0.2
PoolingCLS
Epochs1
Batch size256
Learning rate2e-5
Max. length25

CUI-supervised terminology pairs enriched with parent-level ontology relations.

Terminology statistics

StrategyTripletsCUIsUnique termsUnique positivesTerms/CUIΔ terms
synonyms83,91483,914165,66183,4722.000
parents1,699,553477,290550,651432,5524.03+384,990
grandparents4,952,020477,293550,651485,60710.06+384,990

This model uses 1,699,553 triplets, covering 477,290 CUIs and 550,651 unique normalized terms.

The training terminology is not distributed with this repository because it contains resources subject to UMLS licensing conditions. Only aggregate statistics are released.

Intended use

The model is intended for terminology embedding, biomedical candidate retrieval, concept normalization and entity linking, particularly in cardiology and clinical NLP pipelines. It is not intended for direct clinical decision-making.

Usage

python
import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer

model_id = "DT4H/DT4H_CardioBERTa_parents_en_translations_only"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)

inputs = tokenizer(
    "clinical concept",
    return_tensors="pt",
    truncation=True,
    max_length=25,
)

with torch.no_grad():
    output = model(**inputs)

embedding = F.normalize(
    output.last_hidden_state[:, 0, :],
    p=2,
    dim=1,
)

Reference

Danu et al. CardioLM - a multilingual suite of small language models for the cardiology domain.

Developed within the DataTools4Heart (DT4H) project, Grant Agreement 101057849.