seroe/bge-m3-turkish-triplet-matryoshka
BGE-M3 Türkçe Triplet Matryoshka
This is a sentence-transformers model finetuned from BAAI/bge-m3 on the vodex-turkish-triplets dataset. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
⚠️ Domain-Specific Warning
This model was fine-tuned on Turkish data specifically sourced from the telecommunications domain. While it performs well on telecom-related tasks such as mobile services, billing, campaigns, and subscription details, it may not generalize well to other domains. Please assess its performance carefully before applying it outside of telecommunications use cases.
Model Description
- Model Type: Sentence Transformer
- Base model: BAAI/bge-m3 <!-- at revision 5617a9f61b028005a4858fdac845db406aefb181 -->
- Maximum Sequence Length: 8192 tokens
- Output Dimensionality: 1024 dimensions
- Similarity Function: Cosine Similarity
- Training Dataset:
- vodex-turkish-triplets
- Language: tr
- License: apache-2.0
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 8192, 'do_lower_case': False}) with Transformer model: XLMRobertaModel
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("seroe/bge-m3-turkish-triplet-matryoshka")
# Run inference
sentences = [
"Vodafone Net'in internet hız garantisi var mı?",
'Vodafone Net, internet hızını garanti etmemekte, bu hız abonenin hattının uygunluğuna ve santrale olan mesafeye bağlı olarak değişiklik göstermektedir.',
'Vodafone Net, tüm abonelerine en az 100 Mbps hız garantisi vermektedir.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
Evaluation
Metrics
Triplet
- Datasets:
tr-triplet-dev-1024dandall-nli-test-1024d - Evaluated with <code>TripletEvaluator</code> with these parameters:
{
"truncate_dim": 1024
}Triplet
- Datasets:
tr-triplet-dev-768dandall-nli-test-768d - Evaluated with <code>TripletEvaluator</code> with these parameters:
{
"truncate_dim": 768
}Triplet
- Datasets:
tr-triplet-dev-512dandall-nli-test-512d - Evaluated with <code>TripletEvaluator</code> with these parameters:
{
"truncate_dim": 512
}Triplet
- Datasets:
tr-triplet-dev-256dandall-nli-test-256d - Evaluated with <code>TripletEvaluator</code> with these parameters:
{
"truncate_dim": 256
}Triplet
- Dataset:
tr-triplet-dev-1024d - Evaluated with <code>TripletEvaluator</code> with these parameters:
{
"truncate_dim": 1024
}Triplet
- Dataset:
tr-triplet-dev-768d - Evaluated with <code>TripletEvaluator</code> with these parameters:
{
"truncate_dim": 768
}Triplet
- Dataset:
tr-triplet-dev-512d - Evaluated with <code>TripletEvaluator</code> with these parameters:
{
"truncate_dim": 512
}Triplet
- Dataset:
tr-triplet-dev-256d - Evaluated with <code>TripletEvaluator</code> with these parameters:
{
"truncate_dim": 256
}<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
vodex-turkish-triplets
- Dataset: vodex-turkish-triplets at 0c9fab0
- Size: 70,941 training samples
- Columns: <code>query</code>, <code>positive</code>, and <code>negative</code>
- Approximate statistics based on the first 1000 samples: | | query | positive | negative | |:--------|:----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | type | string | string | string | | details | <ul><li>min: 4 tokens</li><li>mean: 13.58 tokens</li><li>max: 46 tokens</li></ul> | <ul><li>min: 11 tokens</li><li>mean: 26.32 tokens</li><li>max: 61 tokens</li></ul> | <ul><li>min: 10 tokens</li><li>mean: 20.54 tokens</li><li>max: 45 tokens</li></ul> |
- Samples: | query | positive | negative | |:-----------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Kampanya tarihleri ve katılım şartları</code> | <code>Kampanya, 11 Ekim 2018'de başlayıp 29 Ekim 2018'de sona erecek. Katılımcıların belirli bilgileri doldurması ve Vodafone Müzik pass veya Video pass sahibi olmaları gerekiyor.</code> | <code>Kampanya, sadece İstanbul'daki kullanıcılar için geçerli olup, diğer şehirlerden katılım mümkün değildir.</code> | | <code>Taahhüt süresi dolmadan başka bir kampanyaya geçiş yapılırsa ne olur?</code> | <code>Eğer abone taahhüt süresi dolmadan başka bir kampanyaya geçerse, bu durumda önceki kampanya süresince sağlanan indirimler ve diğer faydalar, iptal tarihinden sonraki fatura ile tahsil edilecektir.</code> | <code>Aboneler, taahhüt süresi dolmadan başka bir kampanyaya geçtiklerinde, yeni kampanyadan faydalanmak için ek bir ücret ödemek zorundadırlar.</code> | | <code>FreeZone üyeliğimi nasıl sorgulayabilirim?</code> | <code>Üyeliğinizi sorgulamak için FREEZONESORGU yazarak 1525'e SMS gönderebilirsiniz.</code> | <code>Üyeliğinizi sorgulamak için Vodafone mağazasına gitmeniz gerekmektedir.</code> |
- Loss: <code>MatryoshkaLoss</code> with these parameters:
{
"loss": "CachedMultipleNegativesRankingLoss",
"matryoshka_dims": [
1024,
768,
512,
256
],
"matryoshka_weights": [
1,
1,
1,
1
],
"n_dims_per_step": -1
}Evaluation Dataset
vodex-turkish-triplets
- Dataset: vodex-turkish-triplets at 0c9fab0
- Size: 3,941 evaluation samples
- Columns: <code>query</code>, <code>positive</code>, and <code>negative</code>
- Approximate statistics based on the first 1000 samples: | | query | positive | negative | |:--------|:----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------|:---------------------------------------------------------------------------------| | type | string | string | string | | details | <ul><li>min: 4 tokens</li><li>mean: 13.26 tokens</li><li>max: 36 tokens</li></ul> | <ul><li>min: 12 tokens</li><li>mean: 26.55 tokens</li><li>max: 62 tokens</li></ul> | <ul><li>min: 9 tokens</li><li>mean: 20.4 tokens</li><li>max: 40 tokens</li></ul> |
- Samples: | query | positive | negative | |:-----------------------------------------------------------------------|:----------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------| | <code>Vodafone Net'e geçiş yaparken bağlantı ücreti var mı?</code> | <code>Vodafone Net'e geçişte 264 TL bağlantı ücreti bulunmaktadır ve bu ücret 24 ay boyunca aylık 11 TL olarak faturalandırılmaktadır.</code> | <code>Vodafone Net'e geçişte bağlantı ücreti yoktur ve tüm işlemler ücretsizdir.</code> | | <code>Bağımsız akıllı cihaz kampanyalarının detayları nelerdir?</code> | <code>Kampanyalar, farklı cihaz modelleri için aylık ödeme planları sunmaktadır.</code> | <code>Vodafone'un kampanyaları, sadece internet paketleri ile ilgilidir.</code> | | <code>Fibermax hizmeti iptal edilirse ne gibi sonuçlar doğar?</code> | <code>İptal işlemi taahhüt süresi bitmeden yapılırsa, indirimler ve ücretsiz hizmet bedelleri ödenmelidir.</code> | <code>Fibermax hizmeti iptal edildiğinde, kullanıcıdan hiçbir ücret talep edilmez.</code> |
- Loss: <code>MatryoshkaLoss</code> with these parameters:
{
"loss": "CachedMultipleNegativesRankingLoss",
"matryoshka_dims": [
1024,
768,
512,
256
],
"matryoshka_weights": [
1,
1,
1,
1
],
"n_dims_per_step": -1
}Training Hyperparameters
Non-Default Hyperparameters
eval_strategy: stepsper_device_train_batch_size: 2048per_device_eval_batch_size: 256weight_decay: 0.01num_train_epochs: 2lr_scheduler_type: cosinewarmup_ratio: 0.05bf16: Truebatch_sampler: no_duplicates
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 2048per_device_eval_batch_size: 256per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.01adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 2max_steps: -1lr_scheduler_type: cosinelr_scheduler_kwargs: {}warmup_ratio: 0.05warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Truefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}tp_size: 0fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional
</details>
Training Logs
Framework Versions
- Python: 3.10.12
- Sentence Transformers: 4.2.0.dev0
- Transformers: 4.51.3
- PyTorch: 2.7.0+cu126
- Accelerate: 1.6.0
- Datasets: 3.6.0
- Tokenizers: 0.21.1
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}MatryoshkaLoss
@misc{kusupati2024matryoshka,
title={Matryoshka Representation Learning},
author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
year={2024},
eprint={2205.13147},
archivePrefix={arXiv},
primaryClass={cs.LG}
}CachedMultipleNegativesRankingLoss
@misc{gao2021scaling,
title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
year={2021},
eprint={2101.06983},
archivePrefix={arXiv},
primaryClass={cs.LG}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
