CoolFace
Modelpublic

KayaTechAI/BGE-M3-0.56B-Fine-Tuned-Telecom-Technical-Documents-Retrieval-Embedding-With-Config

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes70downloads
Model Card

SentenceTransformer based on BAAI/bge-m3

This is a sentence-transformers model finetuned from BAAI/bge-m3 on the telecom-technical-documents-retrieval-embedding-dataset dataset. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: BAAI/bge-m3 <!-- at revision 5617a9f61b028005a4858fdac845db406aefb181 -->
  • Maximum Sequence Length: 8192 tokens
  • Output Dimensionality: 1024 dimensions
  • Similarity Function: Cosine Similarity
  • Training Dataset:
  • telecom-technical-documents-retrieval-embedding-dataset <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 8192, 'do_lower_case': False, 'architecture': 'XLMRobertaModel'})
  (1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("KayaTechAI/BGE-M3-0.56B-Fine-Tuned-Telecom-Technical-Documents-Retrieval-Embedding-With-Config")
# Run inference
sentences = [
    'What is the provisioning scope for the eMLPP service?',
    'eMLPP is provisioned per subscriber.',
    'The main objective is to verify that the User Equipment (UE) tracks channel variations and selects the optimal transport format for frequency non-selective scheduling.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000,  0.7451, -0.1130],
#         [ 0.7451,  1.0000, -0.1339],
#         [-0.1130, -0.1339,  1.0000]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
json
  {
      "truncate_dim": 1024
  }
MetricValue
cosine_accuracy@10.7864
cosine_accuracy@30.9008
cosine_accuracy@50.9336
cosine_accuracy@100.9556
cosine_precision@10.7864
cosine_precision@30.3003
cosine_precision@50.1867
cosine_precision@100.0956
cosine_recall@10.7864
cosine_recall@30.9008
cosine_recall@50.9336
cosine_recall@100.9556
cosine_ndcg@100.8753
cosine_mrr@100.849
cosine_map@1000.851
Information Retrieval
json
  {
      "truncate_dim": 768
  }
MetricValue
cosine_accuracy@10.7884
cosine_accuracy@30.9016
cosine_accuracy@50.9312
cosine_accuracy@100.9572
cosine_precision@10.7884
cosine_precision@30.3005
cosine_precision@50.1862
cosine_precision@100.0957
cosine_recall@10.7884
cosine_recall@30.9016
cosine_recall@50.9312
cosine_recall@100.9572
cosine_ndcg@100.8769
cosine_mrr@100.8507
cosine_map@1000.8526
Information Retrieval
json
  {
      "truncate_dim": 512
  }
MetricValue
cosine_accuracy@10.7868
cosine_accuracy@30.9016
cosine_accuracy@50.932
cosine_accuracy@100.9564
cosine_precision@10.7868
cosine_precision@30.3005
cosine_precision@50.1864
cosine_precision@100.0956
cosine_recall@10.7868
cosine_recall@30.9016
cosine_recall@50.932
cosine_recall@100.9564
cosine_ndcg@100.876
cosine_mrr@100.8497
cosine_map@1000.8514
Information Retrieval
json
  {
      "truncate_dim": 256
  }
MetricValue
cosine_accuracy@10.7812
cosine_accuracy@30.8964
cosine_accuracy@50.928
cosine_accuracy@100.9552
cosine_precision@10.7812
cosine_precision@30.2988
cosine_precision@50.1856
cosine_precision@100.0955
cosine_recall@10.7812
cosine_recall@30.8964
cosine_recall@50.928
cosine_recall@100.9552
cosine_ndcg@100.8721
cosine_mrr@100.845
cosine_map@1000.8467
Information Retrieval
json
  {
      "truncate_dim": 128
  }
MetricValue
cosine_accuracy@10.7736
cosine_accuracy@30.8868
cosine_accuracy@50.9212
cosine_accuracy@100.95
cosine_precision@10.7736
cosine_precision@30.2956
cosine_precision@50.1842
cosine_precision@100.095
cosine_recall@10.7736
cosine_recall@30.8868
cosine_recall@50.9212
cosine_recall@100.95
cosine_ndcg@100.8648
cosine_mrr@100.8372
cosine_map@1000.839
Information Retrieval
json
  {
      "truncate_dim": 64
  }
MetricValue
cosine_accuracy@10.7456
cosine_accuracy@30.8644
cosine_accuracy@50.9016
cosine_accuracy@100.9344
cosine_precision@10.7456
cosine_precision@30.2881
cosine_precision@50.1803
cosine_precision@100.0934
cosine_recall@10.7456
cosine_recall@30.8644
cosine_recall@50.9016
cosine_recall@100.9344
cosine_ndcg@100.8427
cosine_mrr@100.8131
cosine_map@1000.8154

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

telecom-technical-documents-retrieval-embedding-dataset
  • Dataset: telecom-technical-documents-retrieval-embedding-dataset at 3ebf34a
  • Size: 127,731 training samples
  • Columns: <code>anchor</code> and <code>positive</code>
  • Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:----------------------------------------------------------------------------------|:----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 8 tokens</li><li>mean: 22.46 tokens</li><li>max: 75 tokens</li></ul> | <ul><li>min: 5 tokens</li><li>mean: 31.38 tokens</li><li>max: 95 tokens</li></ul> |
  • Samples: | anchor | positive | |:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>What is the estimated Transmit power considered sufficient for achieving 95% Downlink coverage with a single Base Station?</code> | <code>Approximately 14 dBm Transmit power is considered sufficient.</code> | | <code>What is the primary goal of the Nominal Accuracy requirement?</code> | <code>The primary goal of the Nominal Accuracy requirement is to ensure good accuracy when signal conditions are ideal.</code> | | <code>What happens on the mobile station side if contention resolution fails because the G-RNTI value in the network's acknowledgement message differs from what the mobile station sent?</code> | <code>If the mobile station receives a PACKET UPLINK ACK/NACK message with a G-RNTI value different from the one it included in its first RLC data blocks, it signifies a contention resolution failure, and the mobile station will not transmit a PACKET CONTROL ACKNOWLEDGEMENT.</code> |
  • Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "MultipleNegativesRankingLoss",
      "matryoshka_dims": [
          1024,
          768,
          512,
          256,
          128,
          64
      ],
      "matryoshka_weights": [
          1,
          1,
          1,
          1,
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Training Hyperparameters

Non-Default Hyperparameters
  • eval_strategy: epoch
  • per_device_train_batch_size: 32
  • per_device_eval_batch_size: 32
  • gradient_accumulation_steps: 16
  • learning_rate: 2e-05
  • num_train_epochs: 4
  • lr_scheduler_type: cosine
  • warmup_ratio: 0.1
  • bf16: True
  • tf32: True
  • load_best_model_at_end: True
  • batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • overwrite_output_dir: False
  • do_predict: False
  • eval_strategy: epoch
  • prediction_loss_only: True
  • per_device_train_batch_size: 32
  • per_device_eval_batch_size: 32
  • per_gpu_train_batch_size: None
  • per_gpu_eval_batch_size: None
  • gradient_accumulation_steps: 16
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 2e-05
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1.0
  • num_train_epochs: 4
  • max_steps: -1
  • lr_scheduler_type: cosine
  • lr_scheduler_kwargs: {}
  • warmup_ratio: 0.1
  • warmup_steps: 0
  • log_level: passive
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • save_safetensors: True
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • no_cuda: False
  • use_cpu: False
  • use_mps_device: False
  • seed: 42
  • data_seed: None
  • jit_mode_eval: False
  • use_ipex: False
  • bf16: True
  • fp16: False
  • fp16_opt_level: O1
  • half_precision_backend: auto
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: True
  • local_rank: 0
  • ddp_backend: None
  • tpu_num_cores: None
  • tpu_metrics_debug: False
  • debug: []
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_prefetch_factor: None
  • past_index: -1
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: True
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_min_num_params: 0
  • fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • fsdp_transformer_layer_cls_to_wrap: None
  • accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamwtorchfused
  • optim_args: None
  • adafactor: False
  • group_by_length: False
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • use_legacy_prediction_loop: False
  • push_to_hub: False
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_private_repo: None
  • hub_always_push: False
  • hub_revision: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_inputs_for_metrics: False
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • fp16_backend: auto
  • push_to_hub_model_id: None
  • push_to_hub_organization: None
  • mp_parameters:
  • auto_find_batch_size: False
  • full_determinism: False
  • torchdynamo: None
  • ray_scope: last
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • include_tokens_per_second: False
  • include_num_input_tokens_seen: False
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: False
  • use_liger_kernel: False
  • liger_kernel_config: None
  • eval_use_gather_object: False
  • average_tokens_across_devices: False
  • prompts: None
  • batch_sampler: no_duplicates
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining Lossdim_1024_cosine_ndcg@10dim_768_cosine_ndcg@10dim_512_cosine_ndcg@10dim_256_cosine_ndcg@10dim_128_cosine_ndcg@10dim_64_cosine_ndcg@10
0.0401101.8502------
0.0802201.5404------
0.1202301.0689------
0.1603400.7949------
0.2004500.6734------
0.2405600.5544------
0.2806700.4663------
0.3206800.4347------
0.3607900.4077------
0.40081000.3298------
0.44091100.3129------
0.48101200.3046------
0.52101300.2918------
0.56111400.2973------
0.60121500.3111------
0.64131600.2382------
0.68141700.2729------
0.72141800.2779------
0.76151900.2543------
0.80162000.2585------
0.84172100.2546------
0.88182200.2263------
0.92182300.2343------
0.96192400.2131------
1.02500.20420.86780.86620.86260.85780.84630.8150
1.04012600.1439------
1.08022700.1587------
1.12022800.1613------
1.16032900.1346------
1.20043000.1622------
1.24053100.1655------
1.28063200.1395------
1.32063300.1409------
1.36073400.1415------
1.40083500.1319------
1.44093600.1387------
1.48103700.1238------
1.52103800.1314------
1.56113900.123------
1.60124000.1374------
1.64134100.1456------
1.68144200.1571------
1.72144300.1267------
1.76154400.1338------
1.80164500.124------
1.84174600.1337------
1.88184700.1051------
1.92184800.13------
1.96194900.1293------
2.05000.11270.87350.87460.87360.86750.85650.8347
2.04015100.0809------
2.08025200.0784------
2.12025300.0783------
2.16035400.0733------
2.20045500.0919------
2.24055600.0806------
2.28065700.0885------
2.32065800.0857------
2.36075900.0866------
2.40086000.0819------
2.44096100.0844------
2.48106200.0958------
2.52106300.0814------
2.56116400.0764------
2.60126500.0754------
2.64136600.0863------
2.68146700.0806------
2.72146800.0702------
2.76156900.0809------
2.80167000.089------
2.84177100.084------
2.88187200.0954------
2.92187300.0737------
2.96197400.0683------
3.07500.080.87710.87570.87620.87260.86680.8414
3.04017600.0541------
3.08027700.063------
3.12027800.0772------
3.16037900.0722------
3.20048000.0666------
3.24058100.0703------
3.28068200.0734------
3.32068300.0675------
3.36078400.0633------
3.40088500.0623------
3.44098600.0646------
3.48108700.0601------
3.52108800.0711------
3.56118900.0774------
3.60129000.0656------
3.64139100.0652------
3.68149200.0637------
3.72149300.0641------
3.76159400.063------
3.80169500.066------
3.84179600.0589------
3.88189700.0608------
3.92189800.0723------
3.96199900.0683------
4.010000.06570.87530.87690.87600.87210.86480.8427
  • The bold row denotes the saved checkpoint.

Framework Versions

  • Python: 3.12.12
  • Sentence Transformers: 5.2.3
  • Transformers: 4.55.4
  • PyTorch: 2.10.0+cu128
  • Accelerate: 1.13.0
  • Datasets: 3.6.0
  • Tokenizers: 0.21.4

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MatryoshkaLoss
bibtex
@misc{kusupati2024matryoshka,
    title={Matryoshka Representation Learning},
    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
    year={2024},
    eprint={2205.13147},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->