CoolFace
Modelpublic

TuanNM171284/TuanNM171284-HaLong-embedding-medical-v5

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes14downloads
Model Card

SentenceTransformer based on hiieu/halong_embedding

This is a sentence-transformers model finetuned from hiieu/halong_embedding. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: hiieu/halong_embedding <!-- at revision b57776031035f70ed2030d2e35ecc533eb0f8f71 -->
  • —Maximum Sequence Length: 512 tokens
  • —Output Dimensionality: 768 dimensions
  • —Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'XLMRobertaModel'})
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("TuanNM171284/TuanNM171284-HaLong-embedding-medical-v5")
# Run inference
sentences = [
    'Nếu một con bò có hành vi bất thường, khó khăn khi di chuyển và bị giảm thể trọng, đó có thể là dấu hiệu của bệnh gì?',
    'Nếu một con bò có các biểu hiện như hành vi bất thường, khó khăn trong di chuyển và giảm thể trọng, đó có thể là dấu hiệu của bệnh bò điên (Bovine Spongiform Encephalopathy - BSE).',
    'Trong giai đoạn viêm tấy cấp tính khi bị giãn dây chằng, người bệnh tuyệt đối không được chườm nóng, xoa dầu nóng, rượu thuốc vì có thể làm tổn thương nặng hơn. Đồng thời, không vận động vùng tổn thương và không được tiêm kháng viêm trực tiếp vào vùng tổn thương.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000,  0.9184, -0.0623],
#         [ 0.9184,  1.0000, -0.0776],
#         [-0.0623, -0.0776,  1.0000]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
json
  {
      "truncate_dim": 768
  }
MetricValue
cosine_accuracy@10.7744
cosine_accuracy@30.9914
cosine_accuracy@50.9985
cosine_accuracy@101.0
cosine_precision@10.7744
cosine_precision@30.3305
cosine_precision@50.1997
cosine_precision@100.1
cosine_recall@10.7744
cosine_recall@30.9914
cosine_recall@50.9985
cosine_recall@101.0
cosine_ndcg@100.9136
cosine_mrr@100.8833
cosine_map@1000.8833
Information Retrieval
json
  {
      "truncate_dim": 512
  }
MetricValue
cosine_accuracy@10.7763
cosine_accuracy@30.9916
cosine_accuracy@50.999
cosine_accuracy@101.0
cosine_precision@10.7763
cosine_precision@30.3305
cosine_precision@50.1998
cosine_precision@100.1
cosine_recall@10.7763
cosine_recall@30.9916
cosine_recall@50.999
cosine_recall@101.0
cosine_ndcg@100.9145
cosine_mrr@100.8845
cosine_map@1000.8845
Information Retrieval
json
  {
      "truncate_dim": 256
  }
MetricValue
cosine_accuracy@10.7784
cosine_accuracy@30.9936
cosine_accuracy@50.9995
cosine_accuracy@101.0
cosine_precision@10.7784
cosine_precision@30.3312
cosine_precision@50.1999
cosine_precision@100.1
cosine_recall@10.7784
cosine_recall@30.9936
cosine_recall@50.9995
cosine_recall@101.0
cosine_ndcg@100.9158
cosine_mrr@100.8862
cosine_map@1000.8862
Information Retrieval
json
  {
      "truncate_dim": 128
  }
MetricValue
cosine_accuracy@10.7781
cosine_accuracy@30.9935
cosine_accuracy@50.9993
cosine_accuracy@101.0
cosine_precision@10.7781
cosine_precision@30.3312
cosine_precision@50.1999
cosine_precision@100.1
cosine_recall@10.7781
cosine_recall@30.9935
cosine_recall@50.9993
cosine_recall@101.0
cosine_ndcg@100.9157
cosine_mrr@100.886
cosine_map@1000.886
Information Retrieval
json
  {
      "truncate_dim": 64
  }
MetricValue
cosine_accuracy@10.7779
cosine_accuracy@30.9924
cosine_accuracy@50.999
cosine_accuracy@101.0
cosine_precision@10.7779
cosine_precision@30.3308
cosine_precision@50.1998
cosine_precision@100.1
cosine_recall@10.7779
cosine_recall@30.9924
cosine_recall@50.999
cosine_recall@101.0
cosine_ndcg@100.9154
cosine_mrr@100.8857
cosine_map@1000.8857

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

Unnamed Dataset
  • —Size: 5,808 training samples
  • —Columns: <code>anchor</code> and <code>positive</code>
  • —Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 8 tokens</li><li>mean: 23.52 tokens</li><li>max: 83 tokens</li></ul> | <ul><li>min: 10 tokens</li><li>mean: 71.2 tokens</li><li>max: 170 tokens</li></ul> |
  • —Samples: | anchor | positive | |:----------------------------------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Bệnh Addison là gì?</code> | <code>Bệnh Addison, còn được gọi là suy thượng thận nguyên phát, là tình trạng tuyến thượng thận không sản xuất đủ các hormone cortisol và aldosterone.</code> | | <code>Nguyên nhân chính gây ra bệnh Addison là gì?</code> | <code>Hầu hết các trường hợp bệnh Addison là do bệnh tự miễn, khi hệ thống miễn dịch tấn công nhầm tuyến thượng thận. Các nguyên nhân khác bao gồm nhiễm trùng kéo dài như bệnh lao, HIV, nhiễm nấm, và các tế bào ung thư lây lan đến tuyến thượng thận.</code> | | <code>Tôi thường xuyên mệt mỏi, yếu cơ, sụt cân và da bị sạm màu. Đây có thể là triệu chứng của bệnh gì?</code> | <code>Các triệu chứng như mệt mỏi mãn tính, yếu cơ, giảm cân, và da sẫm màu (nám, sạm đen, tàn nhang) có thể là dấu hiệu của bệnh Addison.</code> |
  • —Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "MultipleNegativesRankingLoss",
      "matryoshka_dims": [
          768,
          512,
          256,
          128,
          64
      ],
      "matryoshka_weights": [
          1,
          1,
          1,
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: epoch
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —gradient_accumulation_steps: 4
  • —learning_rate: 3e-05
  • —weight_decay: 0.01
  • —num_train_epochs: 10
  • —warmup_ratio: 0.05
  • —fp16: True
  • —load_best_model_at_end: True
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: epoch
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 4
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 3e-05
  • —weight_decay: 0.01
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 10
  • —max_steps: -1
  • —lr_scheduler_type: linear
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.05
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: False
  • —fp16: True
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: True
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —hub_revision: None
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —liger_kernel_config: None
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional
  • —router_mapping: {}
  • —learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining Lossdim_768_cosine_ndcg@10dim_512_cosine_ndcg@10dim_256_cosine_ndcg@10dim_128_cosine_ndcg@10dim_64_cosine_ndcg@10
0.4396200.7115-----
0.8791400.2283-----
1.046-0.87510.87510.87710.87450.8610
1.3077600.1449-----
1.7473800.0753-----
2.092-0.88650.88820.88940.88830.8810
2.17581000.0898-----
2.61541200.0419-----
3.0138-0.89260.89490.89870.89760.8944
3.04401400.0379-----
3.48351600.0292-----
3.92311800.0421-----
4.0184-0.90220.90420.90600.90550.9042
4.35162000.0261-----
4.79122200.0298-----
5.0230-0.90520.90750.90990.90890.9073
5.21982400.0264-----
5.65932600.0177-----
6.0276-0.91110.91230.91410.91380.9117
6.08792800.0236-----
6.52753000.023-----
6.96703200.0225-----
7.0322-0.91150.91270.91450.91480.9129
7.39563400.0163-----
7.83523600.0169-----
8.0368-0.91080.91220.91460.91410.9128
8.26373800.0187-----
8.70334000.0144-----
9.0414-0.91290.91360.91560.91500.9144
9.13194200.0176-----
9.57144400.0105-----
10.04600.0110.91360.91450.91580.91570.9154
  • —The bold row denotes the saved checkpoint.

Framework Versions

  • —Python: 3.12.11
  • —Sentence Transformers: 5.1.0
  • —Transformers: 4.55.2
  • —PyTorch: 2.8.0+cu126
  • —Accelerate: 1.10.0
  • —Datasets: 4.0.0
  • —Tokenizers: 0.21.4

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MatryoshkaLoss
bibtex
@misc{kusupati2024matryoshka,
    title={Matryoshka Representation Learning},
    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
    year={2024},
    eprint={2205.13147},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->