seongil-dn/gte-further-filtered-neg4-bs96
SentenceTransformer based on Alibaba-NLP/gte-multilingual-base
This is a sentence-transformers model finetuned from Alibaba-NLP/gte-multilingual-base. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: Alibaba-NLP/gte-multilingual-base <!-- at revision 7fc06782350c1a83f88b15dd4b38ef853d3b8503 -->
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 768 tokens
- Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: NewModel
(1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("seongil-dn/gte-further-filtered-neg4-bs96")
# Run inference
sentences = [
'안압지의 역사적 이름은 무엇인가요?',
'1980년, 안압지에서 발굴된 토기 파편 등으로 신라시대에 이 곳이 월지(月池)라고 불렸다는 사실이 확인되었다. 이는 신라 왕궁인 반월성(半月城)과 가까이 있었기 때문이며, 임해전의 이름도 본디 월지궁이었다고 한다. 조선시대에는 폐허가 된 이곳에 기러기와 오리들이 날아들자 조선의 묵객들이 안압지(雁鴨池)라는 이름을 붙였다. 《삼국사기》에 동궁을 임해전(臨海殿), 즉 바다에 면한 건물이라고 불렀다는 기록이 있으며, 여기에서 안압지는 바다를 상징한다.',
'안압지라는 명칭은 조선 초기에 간행된 《동국여지승람》과 《동경잡기》등에 나타나고 있다.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 99,239 training samples
- Columns: <code>anchor</code>, <code>positive</code>, and <code>negative</code>
- Approximate statistics based on the first 1000 samples: | | anchor | positive | negative | |:--------|:---------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------| | type | string | string | string | | details | <ul><li>min: 8 tokens</li><li>mean: 18.0 tokens</li><li>max: 32 tokens</li></ul> | <ul><li>min: 16 tokens</li><li>mean: 155.17 tokens</li><li>max: 512 tokens</li></ul> | <ul><li>min: 2 tokens</li><li>mean: 132.62 tokens</li><li>max: 512 tokens</li></ul> |
- Samples: | anchor | positive | negative | |:-----------------------------------------|:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:----------------------------------| | <code>야노 타쿠지의 동생은 어떤 성격을 가지고 있나요?</code> | <code>중반부터 추가된 새로운 전사. 예전 유스케들을 감싸 볼트에게 살해당한 야노 타쿠지의 동생. 강한 파워 파이터 이기도 하며 형의 사후 도로테 박사들에 의해 라이브맨을 지원하는 훈련을 받고 있었다. 타케시라는 동생이 있으나 극중에서 이들 사이의 관계는 뚜렷하지 않은 편이다. 권투가 뛰어나고 조금 무모하고 말투도 난폭하다. 특히 형들의 적인 볼트가 관련되면 냉정할 수 없게 되지만 실력이 따라주지 않는 경우가 많아서 초기에는 멤버 3명의 발목을 잡기도 했다. 하지만 곤란한 사람은 내버려 둘수없는 상냥한 성격이다.</code> | <code>노노의 남동생이며 게임을 좋아한다.</code> | | <code>야노 타쿠지의 동생은 어떤 성격을 가지고 있나요?</code> | <code>중반부터 추가된 새로운 전사. 예전 유스케들을 감싸 볼트에게 살해당한 야노 타쿠지의 동생. 강한 파워 파이터 이기도 하며 형의 사후 도로테 박사들에 의해 라이브맨을 지원하는 훈련을 받고 있었다. 타케시라는 동생이 있으나 극중에서 이들 사이의 관계는 뚜렷하지 않은 편이다. 권투가 뛰어나고 조금 무모하고 말투도 난폭하다. 특히 형들의 적인 볼트가 관련되면 냉정할 수 없게 되지만 실력이 따라주지 않는 경우가 많아서 초기에는 멤버 3명의 발목을 잡기도 했다. 하지만 곤란한 사람은 내버려 둘수없는 상냥한 성격이다.</code> | <code>야스히코의 여동생.</code> | | <code>야노 타쿠지의 동생은 어떤 성격을 가지고 있나요?</code> | <code>중반부터 추가된 새로운 전사. 예전 유스케들을 감싸 볼트에게 살해당한 야노 타쿠지의 동생. 강한 파워 파이터 이기도 하며 형의 사후 도로테 박사들에 의해 라이브맨을 지원하는 훈련을 받고 있었다. 타케시라는 동생이 있으나 극중에서 이들 사이의 관계는 뚜렷하지 않은 편이다. 권투가 뛰어나고 조금 무모하고 말투도 난폭하다. 특히 형들의 적인 볼트가 관련되면 냉정할 수 없게 되지만 실력이 따라주지 않는 경우가 많아서 초기에는 멤버 3명의 발목을 잡기도 했다. 하지만 곤란한 사람은 내버려 둘수없는 상냥한 성격이다.</code> | <code>그의 동생인 고기석 또한 만화가이다.</code> |
- Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim"
}Training Hyperparameters
Non-Default Hyperparameters
per_device_train_batch_size: 96adam_epsilon: 1e-07warmup_ratio: 0.05bf16: Truebatch_sampler: no_duplicates
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: noprediction_loss_only: Trueper_device_train_batch_size: 96per_device_eval_batch_size: 8per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-07max_grad_norm: 1.0num_train_epochs: 3max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.05warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Truefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Truedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Falsehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseeval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseeval_use_gather_object: Falsebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional
</details>
Training Logs
<details><summary>Click to expand</summary>
</details>
Framework Versions
- Python: 3.10.13
- Sentence Transformers: 3.2.1
- Transformers: 4.44.2
- PyTorch: 2.4.0+cu121
- Accelerate: 1.1.1
- Datasets: 2.21.0
- Tokenizers: 0.19.1
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}MultipleNegativesRankingLoss
@misc{henderson2017efficient,
title={Efficient Natural Language Response Suggestion for Smart Reply},
author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
year={2017},
eprint={1705.00652},
archivePrefix={arXiv},
primaryClass={cs.CL}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
