CoolFace
Modelpublic

YesayaAlvinK/indobert-bible-search

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes42downloads
Model Card

SentenceTransformer based on indobenchmark/indobert-base-p1

This is a sentence-transformers model finetuned from indobenchmark/indobert-base-p1. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: indobenchmark/indobert-base-p1 <!-- at revision c2cd0b51ddce6580eb35263b39b0a1e5fb0a39e2 -->
  • —Maximum Sequence Length: 512 tokens
  • —Output Dimensionality: 768 dimensions
  • —Similarity Function: Cosine Similarity
  • —Supported Modality: Text <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'Lagipula, kepada siapakah aku memperhambakan diri? Bukankah kepada anaknya? Sebagaimana aku memperhambakan diri kepada ayahmu, demikianlah aku memperhambakan diri kepadamu."',
    'Lagi pula, kepada siapakah aku akan mengabdi? Bukankah kepada anaknya? Seperti aku mengabdi kepada ayahmu, demikianlah aku akan berlaku kepadamu.”',
    'Pasanglah telingamu dan datanglah kepada-Ku dengarlah supaya jiwamu akan hidup. Aku akan mengadakan perjanjian yang kekal denganmu, menurut kebaikan-Ku yang teguh kepada Daud.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8390, 0.0139],
#         [0.8390, 1.0000, 0.0512],
#         [0.0139, 0.0512, 1.0000]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

Unnamed Dataset
  • —Size: 62,204 training samples
  • —Columns: <code>sentence0</code> and <code>sentence1</code>
  • —Approximate statistics based on the first 100 samples: | | sentence0 | sentence1 | |:---------|:-----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | type | string | string | | modality | text | text | | details | <ul><li>min: 11 tokens</li><li>mean: 32.89 tokens</li><li>max: 90 tokens</li></ul> | <ul><li>min: 7 tokens</li><li>mean: 32.76 tokens</li><li>max: 101 tokens</li></ul> |
  • —Samples: | sentence0 | sentence1 | |:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Dan dengan berani Yesaya mengatakan: "Aku telah berkenan ditemukan mereka yang tidak mencari Aku, Aku telah menampakkan diri kepada mereka yang tidak menanyakan Aku."</code> | <code>Kemudian Yesaya dengan berani berkata atas nama Allah, “Orang yang tidak mencari Aku akan menemukan Aku. Aku menyatakan diri-Ku kepada orang yang tidak menanyakan Aku.”</code> | | <code>Pada pergantian tahun, pada waktu raja-raja biasanya maju berperang, maka Daud menyuruh Yoab maju beserta orang-orangnya dan seluruh orang Israel. Mereka memusnahkan bani Amon dan mengepung kota Raba, sedang Daud sendiri tinggal di Yerusalem.</code> | <code>Pada pergantian tahun, saat raja-raja keluar, Daud mengirim Yoab beserta anak buahnya dan semua orang Israel untuk memusnahkan orang Amon dan mengepung kota Raba, sedangkan Daud tinggal di Yerusalem.</code> | | <code>Sesuatu apapun yang beragi tidak boleh kamu makan; kamu makanlah roti yang tidak beragi di segala tempat kediamanmu."</code> | <code>Kamu tidak boleh makan apa pun yang beragi. Di seluruh tempat tinggalmu, kamu harus makan roti tidak beragi.’”</code> |
  • —Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "cos_sim",
      "gather_across_devices": false,
      "directions": [
          "query_to_doc"
      ],
      "partition_mode": "joint",
      "hardness_mode": null,
      "hardness_strength": 0.0
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —per_device_train_batch_size: 16
  • —num_train_epochs: 5
  • —per_device_eval_batch_size: 16
  • —multi_dataset_batch_sampler: round_robin
All Hyperparameters

<details><summary>Click to expand</summary>

  • —per_device_train_batch_size: 16
  • —num_train_epochs: 5
  • —max_steps: -1
  • —learning_rate: 5e-05
  • —lr_scheduler_type: linear
  • —lr_scheduler_kwargs: None
  • —warmup_steps: 0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —optim_target_modules: None
  • —gradient_accumulation_steps: 1
  • —average_tokens_across_devices: True
  • —max_grad_norm: 1
  • —label_smoothing_factor: 0.0
  • —bf16: False
  • —fp16: False
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —use_liger_kernel: False
  • —liger_kernel_config: None
  • —use_cache: False
  • —neftune_noise_alpha: None
  • —torch_empty_cache_steps: None
  • —auto_find_batch_size: False
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —include_num_input_tokens_seen: no
  • —log_level: passive
  • —log_level_replica: warning
  • —disable_tqdm: False
  • —project: huggingface
  • —trackio_space_id: None
  • —trackio_bucket_id: None
  • —trackio_static_space_id: None
  • —per_device_eval_batch_size: 16
  • —prediction_loss_only: True
  • —eval_on_start: False
  • —eval_do_concat_batches: True
  • —eval_use_gather_object: False
  • —eval_accumulation_steps: None
  • —include_for_metrics: []
  • —batch_eval_metrics: False
  • —save_only_model: False
  • —save_on_each_node: False
  • —enable_jit_checkpoint: False
  • —push_to_hub: False
  • —hub_private_repo: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_always_push: False
  • —hub_revision: None
  • —load_best_model_at_end: False
  • —ignore_data_skip: False
  • —restore_callback_states_from_checkpoint: False
  • —full_determinism: False
  • —seed: 42
  • —data_seed: None
  • —use_cpu: False
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —parallelism_config: None
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —dataloader_prefetch_factor: None
  • —remove_unused_columns: True
  • —label_names: None
  • —train_sampling_strategy: random
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —ddp_static_graph: None
  • —ddp_backend: None
  • —ddp_timeout: 1800
  • —fsdp: None
  • —fsdp_config: None
  • —deepspeed: None
  • —debug: []
  • —skip_memory_metrics: True
  • —do_predict: False
  • —resume_from_checkpoint: None
  • —warmup_ratio: None
  • —local_rank: -1
  • —prompts: None
  • —batch_sampler: batch_sampler
  • —multi_dataset_batch_sampler: round_robin
  • —router_mapping: {}
  • —learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining Loss
0.12865000.0770
0.257210000.0190
0.385815000.0195
0.514420000.0209
0.643025000.0219
0.771630000.0207
0.900235000.0229
1.028840000.0174
1.157445000.0120
1.286050000.0097
1.414655000.0125
1.543260000.0115
1.671865000.0108
1.800470000.0107
1.929075000.0070
2.057680000.0046
2.186285000.0032
2.314890000.0049
2.443495000.0058
2.5720100000.0040
2.7006105000.0038
2.8292110000.0032
2.9578115000.0023
3.0864120000.0029
3.2150125000.0019
3.3436130000.0023
3.4722135000.0021
3.6008140000.0017
3.7294145000.0014
3.8580150000.0013
3.9866155000.0017
4.1152160000.0016
4.2438165000.0006
4.3724170000.0011
4.5010175000.0012
4.6296180000.0010
4.7582185000.0006
4.8868190000.0021

Training Time

  • —Training: 2.1 hours

Framework Versions

  • —Python: 3.12.13
  • —Sentence Transformers: 5.6.0
  • —Transformers: 5.12.1
  • —PyTorch: 2.11.0+cu128
  • —Accelerate: 1.14.0
  • —Datasets: 4.0.0
  • —Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MultipleNegativesRankingLoss
bibtex
@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->