CoolFace
Modelpublic

HassanCS/peptide_HLA_TCRa_TCRb__esm2_t6_8M_UR50D_up_to_epoch_10

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes19downloads
Model Card

SentenceTransformer based on facebook/esm2t68M_UR50D

This is a sentence-transformers model finetuned from facebook/esm2_t6_8M_UR50D. It maps sentences & paragraphs to a 320-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: facebook/esm2_t6_8M_UR50D <!-- at revision c731040fcd8d73dceaa04b0a8e6329b345b0f5df -->
  • —Maximum Sequence Length: 1026 tokens
  • —Output Dimensionality: 320 dimensions
  • —Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 1026, 'do_lower_case': False, 'architecture': 'EsmModel'})
  (1): Pooling({'word_embedding_dimension': 320, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("HassanCS/peptide_HLA_TCRa_TCRb__esm2_t6_8M_UR50D_up_to_epoch_10")
# Run inference
sentences = [
    'M E V T P S G T W L',
    'G Q A R V A Y Q V',
    'L L L A R A A S L S L',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 320]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.4448, 0.4761],
#         [0.4448, 1.0000, 0.8661],
#         [0.4761, 0.8661, 1.0000]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Semantic Similarity
MetricValue
pearson_cosine0.9678
spearman_cosine0.9651

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

Unnamed Dataset
  • —Size: 101,745 training samples
  • —Columns: <code>sentence1</code>, <code>sentence2</code>, and <code>score</code>
  • —Approximate statistics based on the first 1000 samples: | | sentence1 | sentence2 | score | |:--------|:----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------|:---------------------------------------------------------------| | type | string | string | float | | details | <ul><li>min: 9 tokens</li><li>mean: 11.56 tokens</li><li>max: 26 tokens</li></ul> | <ul><li>min: 10 tokens</li><li>mean: 11.47 tokens</li><li>max: 26 tokens</li></ul> | <ul><li>min: 0.03</li><li>mean: 0.3</li><li>max: 1.0</li></ul> |
  • —Samples: | sentence1 | sentence2 | score | |:-------------------------------|:-------------------------------|:---------------------------------| | <code>V Y G I R L E H F</code> | <code>A L G W V F V P V</code> | <code>0.4487163971048506</code> | | <code>G R I A F F L K Y</code> | <code>I P S I N V H H Y</code> | <code>0.2950890611282742</code> | | <code>I L A K F A H W L</code> | <code>Q T N P V T L Q Y</code> | <code>0.24611965364324553</code> |
  • —Loss: <code>CoSENTLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "pairwise_cos_sim"
  }

Evaluation Dataset

Unnamed Dataset
  • —Size: 11,305 evaluation samples
  • —Columns: <code>sentence1</code>, <code>sentence2</code>, and <code>score</code>
  • —Approximate statistics based on the first 1000 samples: | | sentence1 | sentence2 | score | |:--------|:-----------------------------------------------------------------------------------|:----------------------------------------------------------------------------------|:----------------------------------------------------------------| | type | string | string | float | | details | <ul><li>min: 10 tokens</li><li>mean: 11.48 tokens</li><li>max: 26 tokens</li></ul> | <ul><li>min: 9 tokens</li><li>mean: 11.47 tokens</li><li>max: 26 tokens</li></ul> | <ul><li>min: 0.02</li><li>mean: 0.31</li><li>max: 1.0</li></ul> |
  • —Samples: | sentence1 | sentence2 | score | |:---------------------------------|:---------------------------------|:---------------------------------| | <code>K S K R T P M G F</code> | <code>G A D G V G K S A L</code> | <code>0.35084774335438396</code> | | <code>T P R V T G G G A M</code> | <code>E I F D R Y G E E V</code> | <code>0.21484461945678654</code> | | <code>S R I M L L A P K</code> | <code>K I F G S L A F L</code> | <code>0.14818627314226981</code> |
  • —Loss: <code>CoSENTLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "pairwise_cos_sim"
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: epoch
  • —per_device_train_batch_size: 128
  • —per_device_eval_batch_size: 128
  • —learning_rate: 0.001
  • —weight_decay: 0.0001
  • —num_train_epochs: 10
  • —fp16: True
  • —load_best_model_at_end: True
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: epoch
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 128
  • —per_device_eval_batch_size: 128
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 1
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 0.001
  • —weight_decay: 0.0001
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 10
  • —max_steps: -1
  • —lr_scheduler_type: linear
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.0
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —bf16: False
  • —fp16: True
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: True
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —parallelism_config: None
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —project: huggingface
  • —trackio_space_id: trackio
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —hub_revision: None
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: no
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —liger_kernel_config: None
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: True
  • —prompts: None
  • —batch_sampler: batch_sampler
  • —multi_dataset_batch_sampler: proportional
  • —router_mapping: {}
  • —learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining LossValidation Lossall-dev_spearman_cosine
0.12581008.9433--
0.25162008.7471--
0.37743008.6225--
0.50314008.5602--
0.62895008.5013--
0.75476008.4547--
0.88057008.4103--
1.0795-8.36490.8817
1.00638008.3797--
1.13219008.3271--
1.257910008.3076--
1.383611008.2519--
1.509412008.2203--
1.635213008.2108--
1.761014008.1601--
1.886815008.1745--
2.01590-8.16040.9278
2.012616008.1477--
2.138417008.192--
2.264218008.1813--
2.389919008.1643--
2.515720008.1401--
2.641521008.1043--
2.767322008.1092--
2.893123008.0935--
3.02385-8.08520.9385
3.018924008.0631--
3.144725008.031--
3.270426008.0052--
3.396227007.9924--
3.522028007.9701--
3.647829007.958--
3.773630007.9537--
3.899431007.9321--
4.03180-7.99550.9511
4.025232007.9196--
4.150933007.9891--
4.276734007.9788--
4.402535007.9793--
4.528336007.9699--
4.654137007.9371--
4.779938007.9548--
4.905739007.9349--
5.03975-7.97890.9520
5.031440007.9036--
5.157241007.874--
5.283042007.853--
5.408843007.8523--
5.534644007.8475--
5.660445007.8436--
5.786246007.8402--
5.911947007.8097--
6.04770-7.91750.9592
6.037748007.8271--
6.163549007.8704--
6.289350007.8793--
6.415151007.8644--
6.540952007.8611--
6.666753007.8571--
6.792554007.8439--
6.918255007.845--
7.05565-7.92480.9585
7.044056007.795--
7.169857007.7812--
7.295658007.7793--
7.421459007.7608--
7.547260007.7587--
7.673061007.7586--
7.798762007.7612--
7.924563007.7674--
8.06360-7.88330.9629
8.050364007.7635--
8.176165007.7795--
8.301966007.806--
8.427767007.8075--
8.553568007.7789--
8.679269007.7805--
8.805070007.793--
8.930871007.7699--
9.07155-7.88740.9627
9.056672007.7367--
9.182473007.7384--
9.308274007.7197--
9.434075007.7189--
9.559776007.7204--
9.685577007.7208--
9.811378007.6914--
9.937179007.6861--
10.07950-7.86880.9651
  • —The bold row denotes the saved checkpoint.

Framework Versions

  • —Python: 3.10.19
  • —Sentence Transformers: 5.1.2
  • —Transformers: 4.57.1
  • —PyTorch: 2.9.1+cu128
  • —Accelerate: 1.11.0
  • —Datasets: 4.4.1
  • —Tokenizers: 0.22.1

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
CoSENTLoss
bibtex
@article{10531646,
    author={Huang, Xiang and Peng, Hao and Zou, Dongcheng and Liu, Zhiwei and Li, Jianxin and Liu, Kay and Wu, Jia and Su, Jianlin and Yu, Philip S.},
    journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing},
    title={CoSENT: Consistent Sentence Embedding via Similarity Ranking},
    year={2024},
    doi={10.1109/TASLP.2024.3402087}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->