CoolFace
Modelpublic

seroe/bge-m3-turkish-triplet-matryoshka

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes70downloads
Model Card

BGE-M3 Türkçe Triplet Matryoshka

This is a sentence-transformers model finetuned from BAAI/bge-m3 on the vodex-turkish-triplets dataset. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

⚠️ Domain-Specific Warning

This model was fine-tuned on Turkish data specifically sourced from the telecommunications domain. While it performs well on telecom-related tasks such as mobile services, billing, campaigns, and subscription details, it may not generalize well to other domains. Please assess its performance carefully before applying it outside of telecommunications use cases.

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: BAAI/bge-m3 <!-- at revision 5617a9f61b028005a4858fdac845db406aefb181 -->
  • —Maximum Sequence Length: 8192 tokens
  • —Output Dimensionality: 1024 dimensions
  • —Similarity Function: Cosine Similarity
  • —Training Dataset:
  • —vodex-turkish-triplets
  • —Language: tr
  • —License: apache-2.0

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 8192, 'do_lower_case': False}) with Transformer model: XLMRobertaModel 
  (1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("seroe/bge-m3-turkish-triplet-matryoshka")
# Run inference
sentences = [
    "Vodafone Net'in internet hız garantisi var mı?",
    'Vodafone Net, internet hızını garanti etmemekte, bu hız abonenin hattının uygunluğuna ve santrale olan mesafeye bağlı olarak değişiklik göstermektedir.',
    'Vodafone Net, tüm abonelerine en az 100 Mbps hız garantisi vermektedir.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Triplet
json
  {
      "truncate_dim": 1024
  }
Metrictr-triplet-dev-1024dall-nli-test-1024d
cosine_accuracy0.60870.9508
Triplet
json
  {
      "truncate_dim": 768
  }
Metrictr-triplet-dev-768dall-nli-test-768d
cosine_accuracy0.61740.9533
Triplet
json
  {
      "truncate_dim": 512
  }
Metrictr-triplet-dev-512dall-nli-test-512d
cosine_accuracy0.63030.9546
Triplet
json
  {
      "truncate_dim": 256
  }
Metrictr-triplet-dev-256dall-nli-test-256d
cosine_accuracy0.60160.9546
Triplet
json
  {
      "truncate_dim": 1024
  }
MetricValue
cosine_accuracy0.9566
Triplet
json
  {
      "truncate_dim": 768
  }
MetricValue
cosine_accuracy0.9571
Triplet
json
  {
      "truncate_dim": 512
  }
MetricValue
cosine_accuracy0.9589
Triplet
json
  {
      "truncate_dim": 256
  }
MetricValue
cosine_accuracy0.9604

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

vodex-turkish-triplets
  • —Dataset: vodex-turkish-triplets at 0c9fab0
  • —Size: 70,941 training samples
  • —Columns: <code>query</code>, <code>positive</code>, and <code>negative</code>
  • —Approximate statistics based on the first 1000 samples: | | query | positive | negative | |:--------|:----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | type | string | string | string | | details | <ul><li>min: 4 tokens</li><li>mean: 13.58 tokens</li><li>max: 46 tokens</li></ul> | <ul><li>min: 11 tokens</li><li>mean: 26.32 tokens</li><li>max: 61 tokens</li></ul> | <ul><li>min: 10 tokens</li><li>mean: 20.54 tokens</li><li>max: 45 tokens</li></ul> |
  • —Samples: | query | positive | negative | |:-----------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Kampanya tarihleri ve katılım şartları</code> | <code>Kampanya, 11 Ekim 2018'de başlayıp 29 Ekim 2018'de sona erecek. Katılımcıların belirli bilgileri doldurması ve Vodafone Müzik pass veya Video pass sahibi olmaları gerekiyor.</code> | <code>Kampanya, sadece İstanbul'daki kullanıcılar için geçerli olup, diğer şehirlerden katılım mümkün değildir.</code> | | <code>Taahhüt süresi dolmadan başka bir kampanyaya geçiş yapılırsa ne olur?</code> | <code>Eğer abone taahhüt süresi dolmadan başka bir kampanyaya geçerse, bu durumda önceki kampanya süresince sağlanan indirimler ve diğer faydalar, iptal tarihinden sonraki fatura ile tahsil edilecektir.</code> | <code>Aboneler, taahhüt süresi dolmadan başka bir kampanyaya geçtiklerinde, yeni kampanyadan faydalanmak için ek bir ücret ödemek zorundadırlar.</code> | | <code>FreeZone üyeliğimi nasıl sorgulayabilirim?</code> | <code>Üyeliğinizi sorgulamak için FREEZONESORGU yazarak 1525'e SMS gönderebilirsiniz.</code> | <code>Üyeliğinizi sorgulamak için Vodafone mağazasına gitmeniz gerekmektedir.</code> |
  • —Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "CachedMultipleNegativesRankingLoss",
      "matryoshka_dims": [
          1024,
          768,
          512,
          256
      ],
      "matryoshka_weights": [
          1,
          1,
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Evaluation Dataset

vodex-turkish-triplets
  • —Dataset: vodex-turkish-triplets at 0c9fab0
  • —Size: 3,941 evaluation samples
  • —Columns: <code>query</code>, <code>positive</code>, and <code>negative</code>
  • —Approximate statistics based on the first 1000 samples: | | query | positive | negative | |:--------|:----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------|:---------------------------------------------------------------------------------| | type | string | string | string | | details | <ul><li>min: 4 tokens</li><li>mean: 13.26 tokens</li><li>max: 36 tokens</li></ul> | <ul><li>min: 12 tokens</li><li>mean: 26.55 tokens</li><li>max: 62 tokens</li></ul> | <ul><li>min: 9 tokens</li><li>mean: 20.4 tokens</li><li>max: 40 tokens</li></ul> |
  • —Samples: | query | positive | negative | |:-----------------------------------------------------------------------|:----------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------| | <code>Vodafone Net'e geçiş yaparken bağlantı ücreti var mı?</code> | <code>Vodafone Net'e geçişte 264 TL bağlantı ücreti bulunmaktadır ve bu ücret 24 ay boyunca aylık 11 TL olarak faturalandırılmaktadır.</code> | <code>Vodafone Net'e geçişte bağlantı ücreti yoktur ve tüm işlemler ücretsizdir.</code> | | <code>Bağımsız akıllı cihaz kampanyalarının detayları nelerdir?</code> | <code>Kampanyalar, farklı cihaz modelleri için aylık ödeme planları sunmaktadır.</code> | <code>Vodafone'un kampanyaları, sadece internet paketleri ile ilgilidir.</code> | | <code>Fibermax hizmeti iptal edilirse ne gibi sonuçlar doğar?</code> | <code>İptal işlemi taahhüt süresi bitmeden yapılırsa, indirimler ve ücretsiz hizmet bedelleri ödenmelidir.</code> | <code>Fibermax hizmeti iptal edildiğinde, kullanıcıdan hiçbir ücret talep edilmez.</code> |
  • —Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "CachedMultipleNegativesRankingLoss",
      "matryoshka_dims": [
          1024,
          768,
          512,
          256
      ],
      "matryoshka_weights": [
          1,
          1,
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: steps
  • —per_device_train_batch_size: 2048
  • —per_device_eval_batch_size: 256
  • —weight_decay: 0.01
  • —num_train_epochs: 2
  • —lr_scheduler_type: cosine
  • —warmup_ratio: 0.05
  • —bf16: True
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: steps
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 2048
  • —per_device_eval_batch_size: 256
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 1
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 5e-05
  • —weight_decay: 0.01
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 2
  • —max_steps: -1
  • —lr_scheduler_type: cosine
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.05
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: True
  • —fp16: False
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: False
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —tp_size: 0
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamw_torch
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional

</details>

Training Logs

EpochStepTraining LossValidation Losstr-triplet-dev-1024d_cosine_accuracytr-triplet-dev-768d_cosine_accuracytr-triplet-dev-512d_cosine_accuracytr-triplet-dev-256d_cosine_accuracyall-nli-test-1024d_cosine_accuracyall-nli-test-768d_cosine_accuracyall-nli-test-512d_cosine_accuracyall-nli-test-256d_cosine_accuracy
-1-1--0.60870.61740.63030.6016----
0.34291210.6773.49880.87640.88070.88760.8950----
0.6857246.59472.72190.93450.93530.94110.9419----
1.0286365.7772.46410.95840.95790.96020.9617----
1.3714485.37272.52690.95310.95430.95760.9546----
1.7143605.14852.44400.95660.95710.95890.9604----
-1-1------0.95080.95330.95460.9546

Framework Versions

  • —Python: 3.10.12
  • —Sentence Transformers: 4.2.0.dev0
  • —Transformers: 4.51.3
  • —PyTorch: 2.7.0+cu126
  • —Accelerate: 1.6.0
  • —Datasets: 3.6.0
  • —Tokenizers: 0.21.1

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MatryoshkaLoss
bibtex
@misc{kusupati2024matryoshka,
    title={Matryoshka Representation Learning},
    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
    year={2024},
    eprint={2205.13147},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
CachedMultipleNegativesRankingLoss
bibtex
@misc{gao2021scaling,
    title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
    author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
    year={2021},
    eprint={2101.06983},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->