thang1943/multilingual-e5-large-v2
SentenceTransformer based on intfloat/multilingual-e5-large
This is a sentence-transformers model finetuned from intfloat/multilingual-e5-large on the dataset_full_fixed dataset. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: intfloat/multilingual-e5-large <!-- at revision 0dc5580a448e4284468b8909bae50fa925907bc5 -->
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 1024 dimensions
- Similarity Function: Cosine Similarity
- Training Dataset:
- dataset_full_fixed <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: XLMRobertaModel
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("multilingual-e5-large-v2")
# Run inference
sentences = [
'Hình minh họa Chào em, Bất kỳ\r\nnhững bất thường nào trên da cũng cần được đánh giá chuyên môn của BS bằng quan\r\nsát và khám trực tiếp mới đưa ra nhận định chính xác cao được. Bởi vì\r\nnhững nhận định của người bệnh mang tính chủ quan, dễ bị ảnh hưởng bởi các\r\nthông tin trên báo đài và mạng không chuyên, hơn nữa sang thương da trong bệnh\r\nda liễu đôi khi cũng muôn hình vạn trạng. Điển hình là những thông tin em cung cấp rõ ràng là không\r\ngiống với vẩy phấn hồng, đó có thể là nấm da, viêm da dị ứng, chàm...Không phải\r\nnhất thiết bệnh nào cũng sẽ có đầy đủ tất cả triệu chứng do đó khám BS chuyên\r\nkhoa Da liễu là an toàn và chính xác nhất để tìm ra bệnh và điều trị đúng.',
'Chào bác sĩ, \r\n\r\nNăm nay em 22 tuổi. Gần đây trên bắp đùi trong em xuất hiện những vết đỏ tập trung thành 1 vùng lớn tầm 2 cm có hình bầu dục, gần vùng bẹn và cũng có 2-3 vệt ở vùng rìa bụng. \r\n\r\nKhông ngứa, không đau. Hình dạng giống với vẩy phấn hồng nhưng không bong vẩy, trong mảng đỏ có lốm đốm và hột màu trắng, em để ý cách đây gần 3 tuần. Xin hỏi BS em bị bệnh gì vậy ạ? Xin cảm ơn.',
'Chào BS, \r\n\r\nXin BS tư vấn giúp về sức khỏe của em, vừa mới đi khám bệnh với kết quả như sau: WBC kết quả: 14, 24 k Glucose: 3, 9ml.\r\n\r\nTrước đây em có bị huyết áp cao nhưng nay cũng ổn định trong phạm vi cho phép, ngoài ra em còn bị dạng tắc nghẽn mạch máu ở van tim.\r\n\r\nMong BS tư vấn giúp bệnh của em cần ăn uống và điều trị như thế nào cho có kết quả tốt? Rất mong nhận được hồi âm của BS. Xin chân thành cám ơn!',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
Evaluation
Metrics
Information Retrieval
- Dataset:
dim_768 - Evaluated with <code>InformationRetrievalEvaluator</code>
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
datasetfullfixed
- Dataset: dataset_full_fixed at ef2e7fd
- Size: 54,755 training samples
- Columns: <code>positive</code> and <code>query</code>
- Approximate statistics based on the first 1000 samples: | | positive | query | |:--------|:-------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 29 tokens</li><li>mean: 235.31 tokens</li><li>max: 512 tokens</li></ul> | <ul><li>min: 5 tokens</li><li>mean: 81.27 tokens</li><li>max: 512 tokens</li></ul> |
- Samples: | positive | query | |:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Rách sụn chêm là một trong những chấn thương thường gặp khi chơi thể thao, bị ngã Chào em, Rách sụn chêm ngoài thường dễ hồi phục do vùng này có máu nuôi tốt, nếu rách đơn giản, nhỏ có thể xem xét điều trị bảo tồn với kháng viêm, giảm đau và vật lý trị liệu. Rách sụn chêm độ 3 thường sẽ biểu hiện triệu chứng, do đó, em nên khám lại ở bác sĩ chuyên khoa cơ xương khớp để đánh giá thêm vấn đề vận động của khớp, nhằm tránh nguy cơ thoái hoá khớp, teo cơ về sau. Trong trường hợp cho phép điều trị bảo tồn vẫn nên theo dõi sát kết hợp với vật lý trị liệu tích cực em nhé! Thân mến.</code> | <code>Chào BS, cách đây 2 tháng em có chơi đá bóng, sau khi đá về thì sáng hôm sau gối bị sưng. Em có chườm đá để giảm sưng, đến khi hết sưng đi chụp CHT thì có kết luận rách sụn chêm ngoài độ 3-4, sụn trong độ 2 và rách bán phần dây chằng chéo trước.BS khuyên nên phẫu thuật nhưng hiện tại em đi đứng làm việc nặng bình thường, vẫn có thể chơi thể thao. Xin BS cho em lời khuyên, em cám ơn ạ.</code> | | <code>Mô tả ngắn:<br>Tobcol – Dex có chứa tobramycin , dexamethason natri phosphat do Công ty dược 3/2 sản xuất được dùng trong các trường hợp viêm ở mắt có đáp ứng với steroid, nhiễm khuẩn do các vi khuẩn nhạy cảm với tobramycin gây ra như: Viêm mi mắt, viêm kết mạc, viêm túi lệ, viêm giác mạc, viêm màng bồ đào trước mãn tính, tổn thương giác mạc do hóa chất, tia xạ hay bỏng nhiệt do dị vật.<br>Thành phần:<br>Tobramycin: 0.3%<br>Dexamethason natri phosphat: 0.1%<br>Chỉ định:<br>Thuốc Tobcol – Dex được chỉ định dùng trong các trường hợp sau:<br>Dùng trong các trường hợp viêm ở mắt có đáp ứng với steroid, nhiễm khuẩn do các vi khuẩn nhạy cảm với tobramycin gây ra như: Viêm mi mắt , viêm kết mạc, viêm túi lệ, viêm giác mạc , viêm màng bồ đào trước mãn tính, tổn thương giác mạc do hóa chất, tia xạ hay bỏng nhiệt do dị vật...</code> | <code>Thuốc nhỏ mắt Tobcol-Dex Dược 3-2 điều trị Viêm mi mắt, viêm kết mạc (5ml)</code> | | <code>Mãn kinh khiến chị em "khô hạn" mất đi tự tị trong đời sống tình dục Chào bạn. Mãn kinh là tình trạng buồng trứng đã ngưng hoạt động, không còn khả năng thụ thai. Khi mãn kinh do tình trạng thiếu hụt nội tiết, niêm mạc âm đạo không còn tác động của nội tiết nên trở nên teo mỏng đi, dễ gây khô đau khi quan hệ vợ chồng. Bạn có thể dùng chất bôi trơn hỗ trợ như: gel KY, regelle bơm âm đạo để làm ẩm âm đạo. Nếu có triệu chứng mệt mỏi, mất ngủ… thì bạn có thể khám chuyên khoa phụ khoa để được tư vấn điều trị thuốc nội tiết thay thế.</code> | <code>Xin hỏi bác sĩ:Tôi năm nay 48 tuổi mãn kinh 2 năm nay rồi, nếu quan hệ với chồng có cần uống thuốc ngừa thai không, tháng sau chồng tôi sẽ về thăm tôi, chúng tôi xa nhau hơn 2 năm vì dịch, mong bác sĩ cho tôi lời khuyên. Cảm ơn bác sĩ nhiều !(Lương Vĩnh Phấn - 0932085...)</code> |
- Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim"
}Training Hyperparameters
Non-Default Hyperparameters
eval_strategy: epochper_device_train_batch_size: 5per_device_eval_batch_size: 1learning_rate: 1e-06num_train_epochs: 5lr_scheduler_type: constantwithwarmupwarmup_ratio: 0.1bf16: Truetf32: Falseload_best_model_at_end: Trueoptim: adamwtorchfusedbatch_sampler: no_duplicates
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: epochprediction_loss_only: Trueper_device_train_batch_size: 5per_device_eval_batch_size: 1per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 1e-06weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 5max_steps: -1lr_scheduler_type: constantwithwarmuplr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Truefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Falselocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Trueignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamwtorchfusedoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional
</details>
Training Logs
<details><summary>Click to expand</summary>
- The bold row denotes the saved checkpoint. </details>
Framework Versions
- Python: 3.10.16
- Sentence Transformers: 3.4.1
- Transformers: 4.49.0
- PyTorch: 2.6.0+cu124
- Accelerate: 1.5.2
- Datasets: 3.3.2
- Tokenizers: 0.21.0
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}MultipleNegativesRankingLoss
@misc{henderson2017efficient,
title={Efficient Natural Language Response Suggestion for Smart Reply},
author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
year={2017},
eprint={1705.00652},
archivePrefix={arXiv},
primaryClass={cs.CL}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
