CoolFace
Modelpublic

kamalkraj/BioSimCSE-BioLinkBERT-BASE

sourceHugging Faceupdated 4y agoView on Hugging Face
1likes1.4kdownloads
README.md150 linesDownload Raw Back to root
1---2pipeline_tag: sentence-similarity3tags:4- sentence-transformers5- feature-extraction6- sentence-similarity7- transformers8widget:9- source_sentence: "The up-regulation of miR-146a was also detected in cervical cancer tissues."10  sentences: ["The expression of miR-146a has been found to be up-regulated in cervical cancer.", "Only concomitant ablation of ERK1 and ERK2 impairs tumor growth."]11  example_title: "BioNLP Example"12---13 14# kamalkraj/BioSimCSE-BioLinkBERT-BASE15 16This is a [sentence-transformers](https://www.SBERT.net) model: It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search.17 18<!--- Describe your model here -->19 20## Usage (Sentence-Transformers)21 22Using this model becomes easy when you have [sentence-transformers](https://www.SBERT.net) installed:23 24```25pip install -U sentence-transformers26```27 28Then you can use the model like this:29 30```python31from sentence_transformers import SentenceTransformer32sentences = ["This is an example sentence", "Each sentence is converted"]33 34model = SentenceTransformer('kamalkraj/BioSimCSE-BioLinkBERT-BASE')35embeddings = model.encode(sentences)36print(embeddings)37```38 39 40 41## Usage (HuggingFace Transformers)42Without [sentence-transformers](https://www.SBERT.net), you can use the model like this: First, you pass your input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings.43 44```python45from transformers import AutoTokenizer, AutoModel46import torch47 48 49#Mean Pooling - Take attention mask into account for correct averaging50def mean_pooling(model_output, attention_mask):51    token_embeddings = model_output[0] #First element of model_output contains all token embeddings52    input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()53    return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(input_mask_expanded.sum(1), min=1e-9)54 55 56# Sentences we want sentence embeddings for57sentences = ['This is an example sentence', 'Each sentence is converted']58 59# Load model from HuggingFace Hub60tokenizer = AutoTokenizer.from_pretrained('kamalkraj/BioSimCSE-BioLinkBERT-BASE')61model = AutoModel.from_pretrained('kamalkraj/BioSimCSE-BioLinkBERT-BASE')62 63# Tokenize sentences64encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')65 66# Compute token embeddings67with torch.no_grad():68    model_output = model(**encoded_input)69 70# Perform pooling. In this case, mean pooling.71sentence_embeddings = mean_pooling(model_output, encoded_input['attention_mask'])72 73print("Sentence embeddings:")74print(sentence_embeddings)75```76 77 78 79## Evaluation Results80 81<!--- Describe how your model was evaluated -->82 83For an automated evaluation of this model, see the *Sentence Embeddings Benchmark*: [https://seb.sbert.net](https://seb.sbert.net?model_name=kamalkraj/BioSimCSE-BioLinkBERT-BASE)84 85 86## Training87The model was trained with the parameters:88 89**DataLoader**:90 91`torch.utils.data.dataloader.DataLoader` of length 7708 with parameters:92```93{'batch_size': 128, 'sampler': 'torch.utils.data.sampler.RandomSampler', 'batch_sampler': 'torch.utils.data.sampler.BatchSampler'}94```95 96**Loss**:97 98`sentence_transformers.losses.MultipleNegativesRankingLoss.MultipleNegativesRankingLoss` with parameters:99  ```100  {'scale': 20.0, 'similarity_fct': 'cos_sim'}101  ```102 103Parameters of the fit()-Method:104```105{106    "epochs": 1,107    "evaluation_steps": 0,108    "evaluator": "NoneType",109    "max_grad_norm": 1,110    "optimizer_class": "<class 'torch.optim.adamw.AdamW'>",111    "optimizer_params": {112        "lr": 5e-05113    },114    "scheduler": "WarmupLinear",115    "steps_per_epoch": null,116    "warmup_steps": 771,117    "weight_decay": 0.01118}119```120 121 122## Full Model Architecture123```124SentenceTransformer(125  (0): Transformer({'max_seq_length': 128, 'do_lower_case': False}) with Transformer model: BertModel 126  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False})127)128```129 130## Citing & Authors131 132<!--- Describe where people can find more information -->133```bibtex134@inproceedings{kanakarajan-etal-2022-biosimcse,135    title = "{B}io{S}im{CSE}: {B}io{M}edical Sentence Embeddings using Contrastive learning",136    author = "Kanakarajan, Kamal raj  and137      Kundumani, Bhuvana  and138      Abraham, Abhijith  and139      Sankarasubbu, Malaikannan",140    booktitle = "Proceedings of the 13th International Workshop on Health Text Mining and Information Analysis (LOUHI)",141    month = dec,142    year = "2022",143    address = "Abu Dhabi, United Arab Emirates (Hybrid)",144    publisher = "Association for Computational Linguistics",145    url = "https://aclanthology.org/2022.louhi-1.10",146    pages = "81--86",147    abstract = "Sentence embeddings in the form of fixed-size vectors that capture the information in the sentence as well as the context are critical components of Natural Language Processing systems. With transformer model based sentence encoders outperforming the other sentence embedding methods in the general domain, we explore the transformer based architectures to generate dense sentence embeddings in the biomedical domain. In this work, we present BioSimCSE, where we train sentence embeddings with domain specific transformer based models with biomedical texts. We assess our model{'}s performance with zero-shot and fine-tuned settings on Semantic Textual Similarity (STS) and Recognizing Question Entailment (RQE) tasks. Our BioSimCSE model using BioLinkBERT achieves state of the art (SOTA) performance on both tasks.",148}149```150