CoolFace
Modelpublic

sartifyllc/bge-base-swahili-matryoshka

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes83downloads
Model Card

BGE base Swahili Matryoshka

This is a sentence-transformers model finetuned from BAAI/bge-base-en-v1.5. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: BAAI/bge-base-en-v1.5 <!-- at revision a5beb1e3e68b9ab74eb54cfd186867f64f240e1a -->
  • —Maximum Sequence Length: 512 tokens
  • —Output Dimensionality: 768 tokens
  • —Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown -->
  • —Language: en
  • —License: apache-2.0

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': True}) with Transformer model: BertModel 
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sartifyllc/bge-base-swahili-matryoshka")
# Run inference
sentences = [
    'Kuna masuala ya sera.',
    'Masuala ya sera ya mbinu nyingi na maombi.',
    'Mwanamke mwenye makunyanzi sana akishikilia miwani yake na kutembea kwenye barabara ya jiji.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
MetricValue
cosine_accuracy@10.268
cosine_accuracy@30.35
cosine_accuracy@50.3859
cosine_accuracy@100.4332
cosine_precision@10.268
cosine_precision@30.1167
cosine_precision@50.0772
cosine_precision@100.0433
cosine_recall@10.268
cosine_recall@30.35
cosine_recall@50.3859
cosine_recall@100.4332
cosine_ndcg@100.3461
cosine_mrr@100.3188
cosine_map@1000.3252
Information Retrieval
MetricValue
cosine_accuracy@10.2655
cosine_accuracy@30.346
cosine_accuracy@50.3811
cosine_accuracy@100.4291
cosine_precision@10.2655
cosine_precision@30.1153
cosine_precision@50.0762
cosine_precision@100.0429
cosine_recall@10.2655
cosine_recall@30.346
cosine_recall@50.3811
cosine_recall@100.4291
cosine_ndcg@100.3425
cosine_mrr@100.3154
cosine_map@1000.3217
Information Retrieval
MetricValue
cosine_accuracy@10.2576
cosine_accuracy@30.3379
cosine_accuracy@50.3717
cosine_accuracy@100.4195
cosine_precision@10.2576
cosine_precision@30.1126
cosine_precision@50.0743
cosine_precision@100.042
cosine_recall@10.2576
cosine_recall@30.3379
cosine_recall@50.3717
cosine_recall@100.4195
cosine_ndcg@100.3339
cosine_mrr@100.3071
cosine_map@1000.3134
Information Retrieval
MetricValue
cosine_accuracy@10.2449
cosine_accuracy@30.3219
cosine_accuracy@50.3557
cosine_accuracy@100.4029
cosine_precision@10.2449
cosine_precision@30.1073
cosine_precision@50.0711
cosine_precision@100.0403
cosine_recall@10.2449
cosine_recall@30.3219
cosine_recall@50.3557
cosine_recall@100.4029
cosine_ndcg@100.3191
cosine_mrr@100.2929
cosine_map@1000.2992
Information Retrieval
MetricValue
cosine_accuracy@10.2194
cosine_accuracy@30.2918
cosine_accuracy@50.3235
cosine_accuracy@100.3699
cosine_precision@10.2194
cosine_precision@30.0973
cosine_precision@50.0647
cosine_precision@100.037
cosine_recall@10.2194
cosine_recall@30.2918
cosine_recall@50.3235
cosine_recall@100.3699
cosine_ndcg@100.2897
cosine_mrr@100.2647
cosine_map@1000.271

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

Unnamed Dataset
  • —Size: 282,883 training samples
  • —Columns: <code>positive</code> and <code>anchor</code>
  • —Approximate statistics based on the first 1000 samples: | | positive | anchor | |:--------|:---------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 5 tokens</li><li>mean: 20.1 tokens</li><li>max: 78 tokens</li></ul> | <ul><li>min: 6 tokens</li><li>mean: 38.64 tokens</li><li>max: 184 tokens</li></ul> |
  • —Samples: | positive | anchor | |:----------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Alingoja mtu huyo mwingine arudi.</code> | <code>Ca'daan alingoja hadi alasiri yote mtu huyo atoke tena.</code> | | <code>Sheria hiyo huanzisha mfululizo wa ukaguzi wa majaribio.</code> | <code>Sheria hiyo pia inatoa sheria ya kudhibiti kwa ajili ya mashirika fulani ambayo yanahitaji kuandaa taarifa za kifedha za mashirika yote na kuzisimamisha kwa wakaguzi wa jumla.</code> | | <code>Mbwa anakimbia na kuruka nje.</code> | <code>Mbwa mwenye rangi ya kahawia anaruka na kukimbia shambani.</code> |
  • —Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "MultipleNegativesRankingLoss",
      "matryoshka_dims": [
          768,
          512,
          256,
          128,
          64
      ],
      "matryoshka_weights": [
          1,
          1,
          1,
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: epoch
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —gradient_accumulation_steps: 16
  • —learning_rate: 2e-05
  • —num_train_epochs: 4
  • —lr_scheduler_type: cosine
  • —warmup_ratio: 0.1
  • —bf16: True
  • —tf32: True
  • —load_best_model_at_end: True
  • —optim: adamwtorchfused
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: epoch
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 16
  • —eval_accumulation_steps: None
  • —learning_rate: 2e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 4
  • —max_steps: -1
  • —lr_scheduler_type: cosine
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: True
  • —fp16: False
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: True
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: True
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: False
  • —hub_always_push: False
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —dispatch_batches: None
  • —split_batches: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional

</details>

Training Logs

<details><summary>Click to expand</summary>

EpochStepTraining Lossdim_128_cosine_map@100dim_256_cosine_map@100dim_512_cosine_map@100dim_64_cosine_map@100dim_768_cosine_map@100
0.0181109.7089-----
0.0362209.2806-----
0.0543308.8905-----
0.0724407.9651-----
0.0905507.4201-----
0.1086606.8346-----
0.1267706.515-----
0.1448806.2009-----
0.1629905.8256-----
0.18101005.549-----
0.19911105.1667-----
0.21721205.2684-----
0.23531305.0678-----
0.25341404.9183-----
0.27151504.844-----
0.28961604.5427-----
0.30771704.3324-----
0.32581804.4963-----
0.34391904.1704-----
0.36202004.1285-----
0.38002104.0235-----
0.39812204.0738-----
0.41622303.9132-----
0.43432403.9682-----
0.45242503.7542-----
0.47052603.6508-----
0.48862703.7596-----
0.50672803.5596-----
0.52482903.5077-----
0.54293003.3831-----
0.56103103.4-----
0.57913203.296-----
0.59723303.3646-----
0.61533403.3533-----
0.63343503.2171-----
0.65153603.2324-----
0.66963703.1544-----
0.68773803.3393-----
0.70583903.0864-----
0.72394003.1069-----
0.74204103.0722-----
0.76014203.1446-----
0.77824303.0847-----
0.79634403.0331-----
0.81444503.0197-----
0.83254602.9667-----
0.85064702.8331-----
0.86874802.9333-----
0.88684902.8714-----
0.90495002.8578-----
0.92305102.9689-----
0.94115202.7977-----
0.95925302.9832-----
0.97735402.9761-----
0.99545502.7711-----
0.9990552-0.27720.29540.30520.24450.3080
1.01355602.7194-----
1.03165702.8489-----
1.04975802.6559-----
1.06785902.6239-----
1.08596002.7081-----
1.10396102.6581-----
1.12206202.7709-----
1.14016302.6191-----
1.15826402.6712-----
1.17636502.5445-----
1.19446602.5264-----
1.21256702.5782-----
1.23066802.5652-----
1.24876902.6229-----
1.26687002.5557-----
1.28497102.5251-----
1.30307202.4555-----
1.32117302.5335-----
1.33927402.5027-----
1.35737502.3569-----
1.37547602.4255-----
1.39357702.4626-----
1.41167802.363-----
1.42977902.4-----
1.44788002.3317-----
1.46598102.2922-----
1.48408202.4086-----
1.50218302.3166-----
1.52028402.3401-----
1.53838502.1951-----
1.55648602.214-----
1.57458702.1859-----
1.59268802.3605-----
1.61078902.2528-----
1.62889002.2759-----
1.64699102.1458-----
1.66509202.187-----
1.68319302.3406-----
1.70129402.2151-----
1.71939502.2971-----
1.73749602.2736-----
1.75559702.2329-----
1.77369802.2602-----
1.79179902.2402-----
1.809810002.1971-----
1.827810102.1642-----
1.845910202.1274-----
1.864010302.1833-----
1.882110402.156-----
1.900210502.1252-----
1.918310602.161-----
1.936410702.1267-----
1.954510802.2017-----
1.972610902.3044-----
1.990711002.161-----
1.99981105-0.29280.30850.31650.26320.3204
2.008811102.0594-----
2.026911202.2277-----
2.045011302.1591-----
2.063111402.0396-----
2.081211502.1007-----
2.099311602.0705-----
2.117411702.0894-----
2.135511802.0677-----
2.153611902.0893-----
2.171712001.984-----
2.189812101.9206-----
2.207912202.132-----
2.226012302.0457-----
2.244112402.1428-----
2.262212502.1116-----
2.280312601.993-----
2.298412702.0181-----
2.316512801.9742-----
2.334612902.081-----
2.352713001.9107-----
2.370813101.9507-----
2.388913201.9844-----
2.407013302.0035-----
2.425113401.9121-----
2.443213502.0057-----
2.461313601.9323-----
2.479413701.9216-----
2.497513801.995-----
2.515613901.9285-----
2.533714001.8886-----
2.551714101.8298-----
2.569814201.8452-----
2.587914301.9488-----
2.606014401.8928-----
2.624114502.0101-----
2.642214601.7591-----
2.660314701.9177-----
2.678414801.9329-----
2.696514901.8978-----
2.714615001.9589-----
2.732715101.9744-----
2.750815201.9272-----
2.768915301.9234-----
2.787015401.9667-----
2.805115501.853-----
2.823215601.9191-----
2.841315701.8083-----
2.859415801.8543-----
2.877515901.9091-----
2.895616001.8079-----
2.913716101.8992-----
2.931816201.8742-----
2.949916301.9313-----
2.968016401.9832-----
2.986116501.9037-----
2.99881657-0.29820.31300.32110.26970.3247
3.004216601.7924-----
3.022316701.9677-----
3.040416801.9123-----
3.058516901.7691-----
3.076617001.8822-----
3.094717101.8543-----
3.112817201.8127-----
3.130917301.8844-----
3.149017401.911-----
3.167117501.7695-----
3.185217601.8134-----
3.203317701.7794-----
3.221417801.8851-----
3.239517901.8381-----
3.257618001.9184-----
3.275618101.8074-----
3.293718201.8236-----
3.311818301.8203-----
3.329918401.8874-----
3.348018501.7457-----
3.366118601.7933-----
3.384218701.759-----
3.402318801.8514-----
3.420418901.8163-----
3.438519001.8299-----
3.456619101.8112-----
3.474719201.7446-----
3.492819301.8314-----
3.510919401.742-----
3.529019501.7519-----
3.547119601.722-----
3.565219701.7454-----
3.583319801.7875-----
3.601419901.7596-----
3.619520001.8348-----
3.637620101.6954-----
3.655720201.7334-----
3.673820301.8318-----
3.691920401.7982-----
3.710020501.7987-----
3.728120601.8402-----
3.746220701.8569-----
3.764320801.8285-----
3.782420901.8652-----
3.800521001.7731-----
3.818621101.8697-----
3.836721201.6953-----
3.854821301.7493-----
3.872921401.8031-----
3.891021501.7053-----
3.909121601.8436-----
3.927221701.7572-----
3.945321801.7797-----
3.963421901.8827-----
3.981522001.8678-----
3.99592208-0.29920.31340.32170.2710.3252
  • —The bold row denotes the saved checkpoint. </details>

Framework Versions

  • —Python: 3.10.12
  • —Sentence Transformers: 3.0.1
  • —Transformers: 4.41.2
  • —PyTorch: 2.1.2+cu121
  • —Accelerate: 0.31.0
  • —Datasets: 2.19.1
  • —Tokenizers: 0.19.1

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MatryoshkaLoss
bibtex
@misc{kusupati2024matryoshka,
    title={Matryoshka Representation Learning}, 
    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
    year={2024},
    eprint={2205.13147},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply}, 
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->