CoolFace
Modelpublic

b00l26/embeddinggemma-300m-finetune-news-finalize

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes108downloads
Model Card

SentenceTransformer based on google/embeddinggemma-300m

This is a sentence-transformers model finetuned from google/embeddinggemma-300m on the train dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: google/embeddinggemma-300m <!-- at revision 57c266a740f537b4dc058e1b0cda161fd15afa75 -->
  • —Maximum Sequence Length: 256 tokens
  • —Output Dimensionality: 768 dimensions
  • —Similarity Function: Cosine Similarity
  • —Training Dataset:
  • —train <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 256, 'do_lower_case': False, 'architecture': 'Gemma3TextModel'})
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Dense({'in_features': 768, 'out_features': 3072, 'bias': False, 'activation_function': 'torch.nn.modules.linear.Identity'})
  (3): Dense({'in_features': 3072, 'out_features': 768, 'bias': False, 'activation_function': 'torch.nn.modules.linear.Identity'})
  (4): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("b00l26/embeddinggemma-300m-finetune-news-finalize")
# Run inference
queries = [
    "M\u1ef9 \u0111i\u1ec1u_tra tu\u00e2n_th\u1ee7 tho\u1ea3_thu\u1eadn th\u01b0\u01a1ng_m\u1ea1i Trung_Qu\u1ed1c . V\u0103n_ph\u00f2ng \u0110\u1ea1i_di\u1ec7n Th\u01b0\u01a1ng_m\u1ea1i M\u1ef9 ( USTR ) \u0111i\u1ec1u_tra tu\u00e2n_th\u1ee7 tho\u1ea3_thu\u1eadn th\u01b0\u01a1ng_m\u1ea1i Trung_Qu\u1ed1c k\u00fd_k\u1ebft 2020 .",
]
documents = [
    "Trung_Quốc đồng_ý đàm_phán thương_mại ' ' Mỹ . Trung_Quốc đồng_ý tiến_hành vòng đàm_phán thương_mại Mỹ   .",
    'Thời_điểm đẹp Hà_Giang ngắm tam_giác mạch ? . Du_khách Hà_Nội ngắm hoa tam_giác mạch Hà_Giang , tư_vấn hoa nở_rộ , rực_rỡ .',
    'Bất_động_sản Nam trở_lại đường_đua sáp_nhập địa_giới hành_chính . ( Dân_trí ) - Trong cung lõi TPHCM khan_hiếm khu_vực lân_cận Bình_Dương , Long_An Đồng_Nai ( cũ ) dồi_dào , đa_dạng phân khúc .',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 768] [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[ 0.6707,  0.0895, -0.0779]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
MetricValue
cosine_accuracy@10.0003
cosine_accuracy@30.3094
cosine_accuracy@50.427
cosine_accuracy@100.5576
cosine_precision@10.0003
cosine_precision@30.1187
cosine_precision@50.1154
cosine_precision@100.0891
cosine_recall@10.0001
cosine_recall@30.1817
cosine_recall@50.2839
cosine_recall@100.41
cosine_ndcg@100.2256
cosine_mrr@100.1799
cosine_map@1000.1642

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

train
  • —Dataset: train
  • —Size: 26,696 training samples
  • —Columns: <code>anchor</code> and <code>positive</code>
  • —Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 23 tokens</li><li>mean: 72.82 tokens</li><li>max: 141 tokens</li></ul> | <ul><li>min: 27 tokens</li><li>mean: 73.87 tokens</li><li>max: 132 tokens</li></ul> |
  • —Samples: | anchor | positive | |:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>NamĐịnh : Thuê thần đèn didời 500 đường điện . ( Dântrí ) - Tiếc 2 tầng xâydựng khảnăng dỡ , Thuế mua chi nửa tỷ đồng thuê thần đèn dichuyển đất .</code> | <code>Mụcsởthị mỏng VươngquốcAnh rao 1,3 triệu USD . Ngôi nằm ga tàuđiệnngầm GoldhawkRoad tây London , rộng 6 mét , 5 tầng tổng diệntích 1.034 métvuông . Tuy mỏng , rao giá 950.000 bảng Anh ( 1,3 triệu USD ) .</code> | | <code>Trangtrí đámcưới , đámhỏi kiêng sen trắng cành trúc ? . Theo chuyêngia vănhoá , khônggian chùachiền , đám hiếu ngược đámcưới đámcưới . Tuynhiên , , quanniệm dần .</code> | <code>Đạigia TPHCM chi tiền tỷ tổchức đámcưới bảomẫu . ( Dântrí ) - Vì đem nữ bảomẫu giađình đámcưới , nữ đạigia TPHCM chi tiền tổchức tiệc cưới bảomẫu .</code> | | <code>Đệ phunhân Pháp chỉtrích miệtthị biểutình . Những nổitiếng chínhtrịgia cánh tả Pháp bàytỏ phẫnnộ BrigitteMacron cụmtừ lũ đànbà ngungốc hoạtđộng nữquyền .</code> | <code>Đệ phunhân Pháp ámảnh đồn giớitính . Auziere , gái của Đệ nhất phunhân Pháp , nói rằng mẹ loâu sâusắc xoay quanh ̃ng đồn thất thiệt về giới tính .</code> |
  • —Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "cos_sim",
      "gather_across_devices": false
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: steps
  • —per_device_train_batch_size: 512
  • —per_device_eval_batch_size: 64
  • —gradient_accumulation_steps: 4
  • —num_train_epochs: 40
  • —warmup_ratio: 0.1
  • —bf16: True
  • —dataloader_num_workers: 2
  • —dataloader_pin_memory: False
  • —gradient_checkpointing: True
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: steps
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 512
  • —per_device_eval_batch_size: 64
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 4
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 5e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 40
  • —max_steps: -1
  • —lr_scheduler_type: linear
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: True
  • —fp16: False
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 2
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: False
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —tp_size: 0
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamw_torch
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: False
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —gradient_checkpointing: True
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional
  • —router_mapping: {}
  • —learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining Losscosine_ndcg@10
0.0755118.0826-
0.75471012.68350.1933
1.4528206.97620.2223
2.1509304.54350.2279
2.9057402.72980.2299
3.6038501.32550.2302
4.3019600.94390.2312
5.0700.70380.2243
5.7547800.7720.2274
6.4528900.60220.2313
7.15091000.50640.2304
7.90571100.50720.2340
8.60381200.4430.2339
9.30191300.40510.2255
10.01400.36590.2398
10.75471500.40880.2318
11.45281600.34340.2318
12.15091700.32710.2374
12.90571800.310.2319
13.60381900.27060.2311
14.30192000.27390.2373
15.02100.22550.2226
15.75472200.26390.2364
16.45282300.23030.2303
17.15092400.2020.2368
17.90572500.2160.2297
18.60382600.19510.2347
19.30192700.18350.2376
20.02800.15680.2336
20.75472900.18780.2333
21.45283000.15690.2363
22.15093100.1420.2372
22.90573200.15150.2328
23.60383300.14070.2315
24.30193400.13440.2343
25.03500.10830.2306
25.75473600.13790.2300
26.45283700.11620.2324
27.15093800.10680.2313
27.90573900.11070.2279
28.60384000.10820.2303
29.30194100.10380.2304
30.04200.08150.2280
30.75474300.11150.2278
31.45284400.09530.2280
32.15094500.08620.2268
32.90574600.09010.2263
33.60384700.09240.2262
34.30194800.0880.2260
35.04900.06910.2254
35.75475000.09810.2256
36.45285100.08360.2253
37.15095200.07630.2256

Framework Versions

  • —Python: 3.10.19
  • —Sentence Transformers: 5.2.0
  • —Transformers: 4.51.0
  • —PyTorch: 2.1.2+cu118
  • —Accelerate: 0.34.2
  • —Datasets: 3.3.2
  • —Tokenizers: 0.21.4

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->