meandyou200175/E5_v3_41_instruct_topic
SentenceTransformer based on meandyou200175/E5v3instruct_topic
This is a sentence-transformers model finetuned from meandyou200175/E5_v3_instruct_topic. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: meandyou200175/E5_v3_instruct_topic <!-- at revision e1cd18d29dcab90869d10fb264523bc44cbe8455 -->
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 1024 dimensions
- Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'XLMRobertaModel'})
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("meandyou200175/E5_v3_41_instruct_topic")
# Run inference
sentences = [
'task: classification | query: Từ vựng các loại biển báo giao thông\nBổ sung vốn từ ngay bạn nhé\n#giaoduc\n#hoctap\n#sinhvien\n#hoctienganh\n#tuyensinh\n#luyenthi\n#truonghoc\n#giaovien\n#daihoc\n#giaoducsom',
'Học tập - Kỹ năng',
'Học tập - Kỹ năng',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.4971, 0.4971],
# [0.4971, 1.0000, 1.0000],
# [0.4971, 1.0000, 1.0000]])<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
Evaluation
Metrics
Information Retrieval
- Evaluated with <code>InformationRetrievalEvaluator</code>
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 209,302 training samples
- Columns: <code>anchor</code> and <code>positive</code>
- Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:-------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 12 tokens</li><li>mean: 102.33 tokens</li><li>max: 495 tokens</li></ul> | <ul><li>min: 3 tokens</li><li>mean: 6.59 tokens</li><li>max: 28 tokens</li></ul> |
- Samples: | anchor | positive | |:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------| | <code>task: classification \| query: Kho đạn triều Nguyễn hoả dược khố</code> | <code>Lịch sử</code> | | <code>task: classification \| query: PIQU Nguồn: klapty</code> | <code>Âm nhạc</code> | | <code>task: classification \| query: Bãi Dâu Vũng Tàu<br>Bãi Dâu Vũng Tàu tọa lạc ở đường Trần Phú, Thành phố Vũng Tàu - Đây là một trong những con đường khá lớn, nổi tiếng tại Vũng Tàu nên bạn có thể dễ dàng tìm thấy nó.<br>Theo người dân địa phương nơi đây kể lại, bãi Dâu có tên gọi cũ là bãi Vũng Mây, tên gọi này được đặt dựa trên khung cảnh thiên nhiên được bao phủ rất nhiều mây rừng. Nơi này khá kín gió và bạn có thể thấy nhiều mỏm đá lớn nhô ra ở ngoài biển ở hai đầu bãi.<br>Đặc biệt, bãi Dâu nổi tiếng với vẻ đẹp hoang sơ của thiên nhiên, nó vốn chưa được nhiều người biết đến và chưa được khai thác nhiều. Cũng chính vì điều đó mà bãi Dâu đã sở hữu một đặc trưng về diện mạo hoang sơ mà không phải bất kỳ bãi biển nào tại Vũng Tàu cũng có.<br><br>#vietnam360 #yoolife #dulich #vietnam #34tinhthanh #vr360vietnam #vr360thanhphohochiminh #vr360baidauvungtau #ba</code> | <code>Danh lam thắng cảnh</code> |
- Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim",
"gather_across_devices": false
}Evaluation Dataset
Unnamed Dataset
- Size: 2,115 evaluation samples
- Columns: <code>anchor</code> and <code>positive</code>
- Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 12 tokens</li><li>mean: 98.78 tokens</li><li>max: 308 tokens</li></ul> | <ul><li>min: 3 tokens</li><li>mean: 7.06 tokens</li><li>max: 42 tokens</li></ul> |
- Samples: | anchor | positive | |:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------------------------------| | <code>task: classification \| query: Hội trường gác 2 9 Nguyễn Đình Chiểu, Hà Nội chật kín dự. Sự choán, chiếm thời gian phát biểu yếu nhân Hội Nhà văn, Nguyễn Quang Thiều, Trần Đăng Khoa, Nguyễn Bình Phương, Hữu Thỉnh… mắt sách Nhà văn chữ tình gởi tác giả niên Trình Quang Phú toát yếu hội thảo sách. Cuốn sách 400 trang in đẹp 25 chân dung thảy. Tác giả Trình, Trình Quang Phú- hiếm. Trước lầm cầu sông Sài Gòn tướng Trịnh Minh Thế. Chả phải. Mà Trình. Một dịp hợp, viết đọc cố hương gốc tổ Trình xứ Thanh. Ông đại tá an ninh, văn báo, nhiếp ảnh, doanh nhân thành chủ ngơi Tập đoàn Sao Việt đất Tuy Hòa. Và chức hiện đương Viện trưởng Viện Nghiên cứu phát triển trực Liên hiệp Hội KHKT Việt Nam. Tất tật đều… trúng cả! Tôi đương nhắc tắc nhẽ diễn giả. Tất thảy luyến láy duyên độc đáo chi tiết bầu thành công ký tạm gọi tiểu sử này.</code> | <code>Thời sự</code> | | <code>task: sentence similarity \| query: luyện tập một cách thường xuyên để đạt tới những phẩm chất hay trình độ ở một mức nào đó</code> | <code>rèn luyện thân thể</code> | | <code>task: sentence similarity \| query: gắn thêm từng mảnh trên bề mặt, thường để trang trí</code> | <code>mũ dát ngọc</code> |
- Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim",
"gather_across_devices": false
}Training Hyperparameters
Non-Default Hyperparameters
eval_strategy: stepsper_device_train_batch_size: 32per_device_eval_batch_size: 32learning_rate: 2e-05num_train_epochs: 5warmup_ratio: 0.1bf16: Truebatch_sampler: no_duplicates
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 32per_device_eval_batch_size: 32per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 2e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 5max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Truefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}parallelism_config: Nonedeepspeed: Nonelabel_smoothing_factor: 0.0optim: adamwtorchfusedoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsehub_revision: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseliger_kernel_config: Noneeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportionalrouter_mapping: {}learning_rate_mapping: {}
</details>
Training Logs
<details><summary>Click to expand</summary>
</details>
Framework Versions
- Python: 3.12.6
- Sentence Transformers: 5.1.2
- Transformers: 4.56.0
- PyTorch: 2.8.0+cu129
- Accelerate: 1.10.1
- Datasets: 4.4.1
- Tokenizers: 0.22.0
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}MultipleNegativesRankingLoss
@misc{henderson2017efficient,
title={Efficient Natural Language Response Suggestion for Smart Reply},
author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
year={2017},
eprint={1705.00652},
archivePrefix={arXiv},
primaryClass={cs.CL}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
