along26/all-MiniLM-Malaysian-Multi
SentenceTransformer based on sentence-transformers/all-MiniLM-L6-v2
This is a sentence-transformers model finetuned from sentence-transformers/all-MiniLM-L6-v2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: sentence-transformers/all-MiniLM-L6-v2 <!-- at revision c9745ed1d9f207416be6d2e6f8de32d1f16199bf -->
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 384 dimensions
- Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'BertModel'})
(1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("along26/mpnet_manglish-sentence-transformer")
# Run inference
sentences = [
'Why have critics accused Najib Razak of mishandling the economy and what evidence supports these claims?',
'Mengapa pengkritik menuduh Najib Razak salah mengendalikan ekonomi dan bukti apa yang menyokong dakwaan ini?',
"How does adding more reactant or product affect the equilibrium position of a chemical reaction? Explain using Le Chatelier's principle with at least three different examples.",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000, -0.8678, 0.9288],
# [-0.8678, 1.0000, -0.8804],
# [ 0.9288, -0.8804, 1.0000]])<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 139,404 training samples
- Columns: <code>sentence0</code>, <code>sentence1</code>, and <code>sentence_2</code>
- Approximate statistics based on the first 1000 samples: | | sentence0 | sentence1 | sentence_2 | |:--------|:-------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------| | type | string | string | string | | details | <ul><li>min: 12 tokens</li><li>mean: 221.26 tokens</li><li>max: 512 tokens</li></ul> | <ul><li>min: 21 tokens</li><li>mean: 262.59 tokens</li><li>max: 512 tokens</li></ul> | <ul><li>min: 13 tokens</li><li>mean: 241.29 tokens</li><li>max: 512 tokens</li></ul> |
- Samples: | sentence0 | sentence1 | sentence_2 | |:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Suppose there are 6 guests at a party. Among them, some are friends with each other and some are strangers. What is the minimum number of guests who must be friends with each other or at least strangers with each other, in order to guarantee that there are 3 guests who are all friends or 3 guests who are all strangers?</code> | <code>Eh, imagine got 6 people go party lah. Some of them kakis, some don't know each other one. How many people must be friends or at least don't know each other, so can confirm got 3 people all kakis or 3 people all don't know each other? Aiyoh, very headache leh!</code> | <code>Why is there still a lack of emphasis on science, technology, engineering, and mathematics (STEM) education in Malaysia?</code> | | <code>A photon at point A is entangled with a photon at point B. Using the quantum teleportation protocol, how can you transfer the quantum state of the photon at point A to point C, without physically moving the photon from point A to point C, given that a shared entangled pair of photons is available between point A and point C? Provide a step-by-step explanation of the procedure.</code> | <code>Foton pada titik A terikat dengan foton pada titik B. Menggunakan protokol teleportasi kuantum, bagaimana anda boleh memindahkan keadaan kuantum foton pada titik A ke titik C, tanpa memindahkan foton secara fizikal dari titik A ke titik C, diberikan bahawa sepasang foton terjerat yang dikongsi tersedia antara titik A dan titik C? Berikan penjelasan langkah demi langkah tentang prosedur.</code> | <code>Civil society groups and activists in Malaysia have expressed concern about the government's handling of the 1MDB scandal and the prosecution of those involved for several reasons. The 1MDB scandal involves allegations of massive corruption and money laundering at the state-owned investment fund, with billions of dollars allegedly misappropriated and used for personal gain.<br><br>Firstly, some civil society groups and activists have criticized the government's investigation and prosecution of those involved in the 1MDB scandal as being politically motivated and lacking in transparency. They argue that the government has selectively targeted certain individuals for prosecution, while others with close ties to the ruling party have been left untouched.<br><br>Secondly, there are concerns about the slow pace of the investigations and prosecutions, which have dragged on for years without any significant progress. Some civil society groups and activists have accused the government of deliberately slow...</code> | | <code>To solve this problem, we can use the generating functions method. Let's represent each number in the set {1, 2, 3, ..., 10} as a variable x raised to the power of that number. The generating function for this set is:<br><br>G(x) = x^1 + x^2 + x^3 + ... + x^10<br><br>Now, we want to find the coefficient of x^15 in the expansion of G(x)^3, as this will represent the number of ways to select 3 numbers from the set such that their sum is 15.<br><br>G(x)^3 = (x^1 + x^2 + x^3 + ... + x^10)^3<br><br>Expanding G(x)^3, we are looking for the coefficient of x^15. We can do this by finding the possible combinations of terms that multiply to x^15:<br><br>1. x^1 x^5 x^9<br>2. x^2 x^4 x^9<br>3. x^3 x^4 x^8<br>4. x^3 x^5 x^7<br>5. x^4 x^5 x^6<br><br>Now, let's count the number of ways each combination can be formed:<br><br>1. x^1 x^5 x^9: There is only 1 way to form this combination.<br>2. x^2 x^4 x^9: There is only 1 way to form this combination.<br>3. x^3 x^4 x^8: There is only 1 way to form this combination.<br>4. x^3 x^5 ...</code> | <code>Untuk menyelesaikan masalah ini, kita boleh menggunakan kaedah fungsi penjanaan. Mari kita wakili setiap nombor dalam set {1, 2, 3, ..., 10} sebagai pembolehubah x dinaikkan kepada kuasa nombor itu. Fungsi penjanaan untuk set ini ialah:<br><br>G(x) = x^1 + x^2 + x^3 + ... + x^10<br><br>Sekarang, kita ingin mencari pekali x^15 dalam pengembangan G(x)^3, kerana ini akan mewakili bilangan cara untuk memilih 3 nombor daripada set supaya jumlahnya ialah 15.<br><br>G(x)^3 = (x^1 + x^2 + x^3 + ... + x^10)^3<br><br>Mengembangkan G(x)^3, kita sedang mencari pekali bagi x^15. Kita boleh melakukan ini dengan mencari kemungkinan gabungan istilah yang didarab kepada x^15:<br><br>1. x^1 x^5 x^9<br>2. x^2 x^4 x^9<br>3. x^3 x^4 x^8<br>4. x^3 x^5 x^7<br>5. x^4 x^5 x^6<br><br>Sekarang, mari kita kira bilangan cara setiap gabungan boleh dibentuk:<br><br>1. x^1 x^5 x^9: Terdapat hanya 1 cara untuk membentuk gabungan ini.<br>2. x^2 x^4 x^9: Terdapat hanya 1 cara untuk membentuk gabungan ini.<br>3. x^3 x^4 x^8: Terdapat hanya 1 cara u...</code> | <code>Why does the Malaysian government still insist on implementing the controversial National Security Council Act, which gives the government excessive power to declare a security area and restrict civil liberties?</code> |
- Loss: <code>TripletLoss</code> with these parameters:
{
"distance_metric": "TripletDistanceMetric.EUCLIDEAN",
"triplet_margin": 5
}Training Hyperparameters
Non-Default Hyperparameters
per_device_train_batch_size: 16per_device_eval_batch_size: 16multi_dataset_batch_sampler: round_robin
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: noprediction_loss_only: Trueper_device_train_batch_size: 16per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 3max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}parallelism_config: Nonedeepspeed: Nonelabel_smoothing_factor: 0.0optim: adamwtorchfusedoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthproject: huggingfacetrackio_space_id: trackioddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsehub_revision: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: noneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseliger_kernel_config: Noneeval_use_gather_object: Falseaverage_tokens_across_devices: Trueprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robinrouter_mapping: {}learning_rate_mapping: {}
</details>
Training Logs
Framework Versions
- Python: 3.12.12
- Sentence Transformers: 5.1.2
- Transformers: 4.57.1
- PyTorch: 2.8.0+cu126
- Accelerate: 1.11.0
- Datasets: 4.0.0
- Tokenizers: 0.22.1
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}TripletLoss
@misc{hermans2017defense,
title={In Defense of the Triplet Loss for Person Re-Identification},
author={Alexander Hermans and Lucas Beyer and Bastian Leibe},
year={2017},
eprint={1703.07737},
archivePrefix={arXiv},
primaryClass={cs.CV}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
