Turbo-AI/gte-base-v0__trim_vocab-1024
SentenceTransformer based on Turbo-AI/gte-multilingual-base_trimvocab
This is a sentence-transformers model finetuned from Turbo-AI/gte-multilingual-base__trim_vocab. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: Turbo-AI/gte-multilingual-base__trim_vocab <!-- at revision b49c5d1f2703a4f4725bcd54c2307348ae6b2381 -->
- Maximum Sequence Length: 1024 tokens
- Output Dimensionality: 768 tokens
- Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 1024, 'do_lower_case': False}) with Transformer model: NewModel
(1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("Turbo-AI/gte-base-v0__trim_vocab-1024")
# Run inference
sentences = [
'Trách nhiệm của nhân viên quản lý chất lượng tại phòng xét nghiệm trong quản lý chất lượng bệnh viện đa khoa được quy định ra sao?',
'Trách nhiệm của nhân viên quản lý chất lượng tại phòng xét nghiệm\n1. Tổng hợp, tham mưu cho trưởng phòng xét nghiệm trong triển khai các nội dung về quản lý chất lượng xét nghiệm.\n2. Xây dựng kế hoạch và nội dung quản lý chất lượng xét nghiệm của phòng, trình lãnh đạo phòng xét nghiệm xem xét, quyết định để trình lãnh đạo cơ sở khám bệnh, chữa bệnh xem xét, phê duyệt.\n3. Tổ chức thực hiện chương trình nội kiểm và tham gia chương trình ngoại kiểm để theo dõi, giám sát, đánh giá chất lượng công tác xét nghiệm và phát hiện, đề xuất giải pháp can thiệp kịp thời nhằm quản lý những trường hợp sai sót, có nguy cơ sai sót trong các quy trình xét nghiệm.\n4. Thu thập, tổng hợp, phân tích dữ liệu, quản lý và bảo mật thông tin liên quan đến hoạt động phòng xét nghiệm.\n5. Phối hợp và hỗ trợ các khoa hoặc phòng liên quan khác trong việc triển khai quản lý chất lượng xét nghiệm.\n6. Tổng kết, báo cáo định kỳ hằng tháng, quý và năm về hoạt động và kết quả quản lý chất lượng xét nghiệm với trưởng phòng xét nghiệm, trưởng phòng (hoặc tổ trưởng) quản lý chất lượng bệnh viện và lãnh đạo cơ sở khám bệnh, chữa bệnh.\n7. Là đầu mối tham mưu để thực hiện các công việc liên quan với các tổ chức đánh giá, cấp chứng nhận phòng xét nghiệm đạt tiêu chuẩn quốc gia hoặc tiêu chuẩn quốc tế.',
'1. Phòng tiếp khách đối ngoại tại trụ sở cơ quan đại diện, văn phòng trực thuộc hay nhà riêng có quốc kỳ Việt Nam và treo ảnh hoặc đặt tượng Chủ tịch Hồ Chí Minh phía sau nơi ngồi tiếp khách của người chủ trì tiếp khách.\n2. Treo ảnh hoặc đặt tượng Chủ tịch Hồ Chí Minh ở chính giữa, quốc kỳ Việt Nam treo trên cột cờ đặt phía bên trái ảnh hoặc tượng nếu nhìn từ phía đối diện. Đỉnh của ảnh hoặc tượng Chủ tịch Hồ Chí Minh không cao hơn đỉnh ngôi sao vàng trong quốc kỳ Việt Nam khi treo trên cột cờ. (Xem hình 5 trong Phụ lục).',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
Evaluation
Metrics
Information Retrieval
- Evaluated with <code>InformationRetrievalEvaluator</code>
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 609,413 training samples
- Columns: <code>anchor</code> and <code>positive</code>
- Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:----------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 8 tokens</li><li>mean: 25.66 tokens</li><li>max: 74 tokens</li></ul> | <ul><li>min: 26 tokens</li><li>mean: 265.94 tokens</li><li>max: 1024 tokens</li></ul> |
- Samples: | anchor | positive | |:---------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Việc tuần tra, canh gác bảo vệ đê Điều trong mùa lũ được thực hiện như thế nào?</code> | <code>Thông tư này hướng dẫn tuần tra, canh gác bảo vệ đê Điều trong mùa lũ đối với các tuyến đê sông được phân loại, phân cấp theo quy định tại Điều 4 của Luật Đê Điều.</code> | | <code>Cách thức bảo vệ tuyến đê sông trong mùa lũ được quy định như thế nào?</code> | <code>Thông tư này hướng dẫn tuần tra, canh gác bảo vệ đê Điều trong mùa lũ đối với các tuyến đê sông được phân loại, phân cấp theo quy định tại Điều 4 của Luật Đê Điều.</code> | | <code>Các tuyến đê sông được phân loại, phân cấp thì được bảo vệ ra sao?</code> | <code>Thông tư này hướng dẫn tuần tra, canh gác bảo vệ đê Điều trong mùa lũ đối với các tuyến đê sông được phân loại, phân cấp theo quy định tại Điều 4 của Luật Đê Điều.</code> |
- Loss: <code>CachedMultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim"
}Training Hyperparameters
Non-Default Hyperparameters
eval_strategy: stepsper_device_train_batch_size: 4096per_device_eval_batch_size: 4096num_train_epochs: 5warmup_ratio: 0.05bf16: Trueload_best_model_at_end: Truebatch_sampler: no_duplicates
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 4096per_device_eval_batch_size: 4096per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 5max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.05warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Truefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Trueignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Falsehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseeval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falsebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional
</details>
Training Logs
<details><summary>Click to expand</summary>
</details>
Framework Versions
- Python: 3.10.6
- Sentence Transformers: 3.3.0.dev0
- Transformers: 4.45.2
- PyTorch: 2.4.1+cu118
- Accelerate: 0.34.0
- Datasets: 2.21.0
- Tokenizers: 0.20.2
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}CachedMultipleNegativesRankingLoss
@misc{gao2021scaling,
title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
year={2021},
eprint={2101.06983},
archivePrefix={arXiv},
primaryClass={cs.LG}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
