CoolFace
Modelpublic

Leejy0-0/sbert-korean-triplet-mnr-v1

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes75downloads
Model Card

πŸ” Model Overview

This model is a Korean Sentence-BERT fine-tuned with Triplet and MultipleNegativesRankingLoss. It is optimized for semantic similarity search in Korean construction accident reports.

  • β€”Architecture: SBERT (Korean BERT backbone)
  • β€”Training: Triplet + MNR
  • β€”Language: Korean
  • β€”Use cases: semantic search, STS, dense retrieval

SentenceTransformer

This is a sentence-transformers model trained. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • β€”Model Type: Sentence Transformer <!-- - Base model: Unknown -->
  • β€”Maximum Sequence Length: 512 tokens
  • β€”Output Dimensionality: 768 dimensions
  • β€”Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: BertModel 
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the πŸ€— Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    '21λ…„ 12μ›” 11일(ν† ) 17μ‹œ 20λΆ„κ²½, B/H μš΄μ „μ› κΉ€**κ°€ μž‘μ—…μ™„λ£Œ ν›„ μš΄μ „μ„μ—μ„œ λ‚΄λ €μ˜€λ˜ 쀑 발이 λ―Έλ„λŸ¬μ§€λ©΄μ„œ 볡곡판 μƒλΆ€λ‘œ λ„˜μ–΄μ§ (높이 1m) B/H μš΄μ „μ„ μ΄νƒˆ μ‹œ μ΄λ™λ°©λ²•μ˜ λΆˆλŸ‰μœΌλ‘œ 인해 볡곡판 μƒλΆ€λ‘œ λ„˜μ–΄μ§μš΄μ „μ„ λ°œνŒμ„ ν†΅ν•œ ν•˜μ°¨κ°€ μ•„λ‹Œ λΈ”λ ˆμ΄λ“œ κ²½μ‚¬νŒμ„ 이용',
    '1μΈ΅ μ‹€λ‚΄ 보 μΈν…Œλ¦¬μ–΄ν•„λ¦„ μž‘μ—… 쀑, κ·Όλ‘œμžκ°€ 우마 닀리 체결 μƒνƒœλ₯Ό μ •ν™•ν•˜κ²Œ ν™•μΈν•˜μ§€ μ•Šμ€ μƒνƒœμ—μ„œ 우마 μœ„μ—μ„œ μž‘μ—…μ„ ν•˜λ‹€κ°€ 우마 닀리가 μ ‘ν˜€ λ„˜μ–΄μ§€λ©΄μ„œ μ˜†μ— 있던 우마 λͺ¨μ„œλ¦¬μ— μ½” 뢀뢄을 λΆ€λ”ͺνžˆλŠ” 사고 μš°λ§ˆλ‹€λ¦¬λ‘œ μΈν•œ λ„˜μ–΄μ§€λ©΄μ„œ λͺ¨μ„œλ¦¬μ— μ½”λ₯Ό λΆ€λ”ͺ힘',
    '2022λ…„ 05μ›” 28일(ν† μš”μΌ) 14μ‹œκ²½ μšΈμ‚°κ΄‘μ—­μ‹œ 동ꡬ 일산동에 μ†Œμž¬ν•œ β€œμŠ€νƒ€λ²…μŠ€ μšΈμ‚° μΌμ‚°λΉ„μΉ˜ D/T 신좕곡사” ν˜„μž₯μ—μ„œ 당사 직영 일용근둜자인 졜**(μž¬ν•΄μž)κ°€ 이동식 Aν˜• 사닀리λ₯Ό μ΄μš©ν•˜μ—¬ 1μΈ΅ 벽체 상뢀 타이 ν•€ 제거 μž‘μ—…μ„ μ§„ν–‰ν•˜λŠ” κ³Όμ •μ—μ„œ 쀑심을 μžƒκ³  사닀리와 ν•¨κ»˜ 1.2M μ •λ„μ˜ λ†’μ΄μ—μ„œ λ„˜μ–΄μ Έ 우츑 νŒ” λΆ€λΆ„μ—μž¬ν•΄λ₯Ό μž…λŠ” 사고가 λ°œν–‰ν•˜κ²Œ λ˜μ—ˆμŠ΅λ‹ˆλ‹€. 사고 λ°œμƒ ν›„ μšΈμ‚°λŒ€ν•™κ΅λ³‘μ›μ„ λ°©λ¬Έν•˜μ—¬ μ§„λ£Œ 및 치료λ₯Ό λ°›μ•˜μŠ΅λ‹ˆλ‹€. 이동식사닀리λ₯Ό μ΄μš©ν•œ μž‘μ—…μ€‘ λ„˜μ–΄μ§',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

Unnamed Dataset
  • β€”Size: 28,003 training samples
  • β€”Columns: <code>sentence0</code> and <code>sentence1</code>
  • β€”Approximate statistics based on the first 1000 samples: | | sentence0 | sentence1 | |:--------|:-----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 8 tokens</li><li>mean: 51.87 tokens</li><li>max: 334 tokens</li></ul> | <ul><li>min: 8 tokens</li><li>mean: 54.17 tokens</li><li>max: 325 tokens</li></ul> |
  • β€”Samples: | sentence0 | sentence1 | |:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>μ² κ·Ό κ°€κ³΅μž‘μ—…μ€‘ 철근을 κ°€κ³΅λŒ€μ— μ˜¬λ¦¬λŠ” κ³Όμ •μ—μ„œ 2인1μ‘° μž‘μ—…μ€‘ 1인이 μ„ ν–‰ 철근을 λ‚΄λ €λ†“μŒμœΌλ‘œ μΈν•˜μ—¬ 였λ₯Έμ† μ•½μ§€ 손가락이 μ² κ·Όκ³Ό 지면에 λΆ€λ”ͺ힘 λ°œμƒ μ² κ·Ό κ°€κ³΅μž‘μ—… 쀑 철근을 κ°€κ³΅λŒ€μ— μ˜¬λ¦¬λŠ” κ³Όμ •μ—μ„œ 2인1μ‘° μž‘μ—…μ€‘ 1인이 μ„ ν–‰ 철근을 λ‚΄λ €λ†“μŒμœΌλ‘œ μΈν•˜μ—¬ 였λ₯Έμ† μ•½μ§€ 손가락에 철근에 λΆ€λ”ͺ힘 λ°œμƒ</code> | <code>μ² κ·Ό μ ˆλ‹¨ μž‘μ—… 쀑 κ³ μ†μ ˆλ‹¨κΈ°μ— 손가락 λ² μž„ μž‘μ—…μžμ˜ μž‘μ—…λΆ€μ£Όμœ„μ— μ˜ν•œ 손가락 λ² μž„μ‚¬κ³  λ°œμƒ</code> | | <code>전일 타섀 ν›„ 슬라브 λ¨Ήλ§€κΉ€ μž‘μ—…μ„ μœ„ν•΄ ν—ˆλ¦¬λ₯Ό κ΅½ν˜€ ν†΅λ‘œ μΈκ·Όμ—μ„œ μž‘μ—…μ€‘ 슬라브 양생을 μœ„ν•œ μ‚΄μˆ˜λ‘œ λ°”λ‹₯이 λ―Έλ„λŸ¬μ›Œ κ· ν˜•μ„ μžƒκ³  λ„˜μ–΄μ§€λ©° 수직 철근에 우츑 턱에 열상 λ°œμƒ 1. 이동 쀑 주의λ ₯ λΆ€μ‘±2. 콘크리트 양생을 μœ„ν•΄ μ§€λ©΄ μ‚΄μˆ˜λ‘œ μΈν•œ λ―Έλ„λŸΌ ν˜„μƒ</code> | <code>μ§€ν•˜1μΈ΅ μ‹œμŠ€ν…œλ™λ°”λ¦¬ κ°€μƒˆ ν•΄μ²΄μž‘μ—…ν›„ μ˜†μœΌλ‘œ 이동쀑(높이2.8m) μ§€λ©΄μœΌλ‘œ 떨어짐 μ‹œμŠ€ν…œ λ™λ°”λ¦¬μ˜ κ°€μƒˆ 해체 ν•˜κ³  이동 쀑 μ•ˆμ „κ³ λ¦¬ 미체결둜 떨어짐</code> | | <code>μ² κ·Ό μ ˆλ‹¨κΈ° 상뢀 μž‘μ—… 쀑 λ°œμ„ ν—›λ””λ””λ©΄μ„œ μ ˆλ‹¨κΈ°κ³„μ— κ°€μŠ΄μ„ λΆ€λ”ͺ히며 κ°ˆλΉ„λΌˆ 골절 μ ˆλ‹¨κΈ° 상뢀에 자재λ₯Ό 내리기 μœ„ν•΄ ν•œ λ°œμ„ 단에 올리고 ν•œλ°œλ‘œ μ§€μ§€ν•˜λ˜ 쀑, 자재λ₯Ό λ‚΄λ¦¬λŠ” λ™μž‘ 쀑 발이 λ―Έλ„λŸ¬μ Έ ν—›λ””λ””λ©° 곡ꡬ에 κ°€μŠ΄μ„ λΆ€λ”ͺ힘</code> | <code>μ•„μΉ¨ TBM 쑰회 ν›„ 지상2μΈ΅ μ™ΈλΆ€ μ£Όμ°¨μž₯μ—μ„œ 철근가곡 μž‘μ—…μ„ ν•˜κΈ° μœ„ν•˜μ—¬, κ°€κ³΅μž‘μ—… 쀑 λΆ€μ£Όμ˜λ‘œ μΈν•˜μ—¬ (우츑 κ²€μ§€) 손가락을 μ ˆκ³‘κΈ°κ³„μ— ν˜‘μ°©ν•˜μ—¬ 사고가 λ°œμƒν•¨ μ•„μΉ¨ TBM 쑰회 ν›„ 지상2μΈ΅ μ™ΈλΆ€ μ£Όμ°¨μž₯μ—μ„œ 철근가곡 μž‘μ—…μ„ ν•˜κΈ° μœ„ν•˜μ—¬, κ°€κ³΅μž‘μ—… 쀑 λΆ€μ£Όμ˜λ‘œ μΈν•˜μ—¬ (우츑 κ²€μ§€) 손가락을 μ ˆκ³‘κΈ°κ³„μ— ν˜‘μ°©ν•˜μ—¬ 사고가 λ°œμƒν•¨</code> |
  • β€”Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "cos_sim"
  }

Training Hyperparameters

Non-Default Hyperparameters
  • β€”per_device_train_batch_size: 32
  • β€”per_device_eval_batch_size: 32
  • β€”num_train_epochs: 1
  • β€”multi_dataset_batch_sampler: round_robin
All Hyperparameters

<details><summary>Click to expand</summary>

  • β€”overwrite_output_dir: False
  • β€”do_predict: False
  • β€”eval_strategy: no
  • β€”prediction_loss_only: True
  • β€”per_device_train_batch_size: 32
  • β€”per_device_eval_batch_size: 32
  • β€”per_gpu_train_batch_size: None
  • β€”per_gpu_eval_batch_size: None
  • β€”gradient_accumulation_steps: 1
  • β€”eval_accumulation_steps: None
  • β€”torch_empty_cache_steps: None
  • β€”learning_rate: 5e-05
  • β€”weight_decay: 0.0
  • β€”adam_beta1: 0.9
  • β€”adam_beta2: 0.999
  • β€”adam_epsilon: 1e-08
  • β€”max_grad_norm: 1
  • β€”num_train_epochs: 1
  • β€”max_steps: -1
  • β€”lr_scheduler_type: linear
  • β€”lr_scheduler_kwargs: {}
  • β€”warmup_ratio: 0.0
  • β€”warmup_steps: 0
  • β€”log_level: passive
  • β€”log_level_replica: warning
  • β€”log_on_each_node: True
  • β€”logging_nan_inf_filter: True
  • β€”save_safetensors: True
  • β€”save_on_each_node: False
  • β€”save_only_model: False
  • β€”restore_callback_states_from_checkpoint: False
  • β€”no_cuda: False
  • β€”use_cpu: False
  • β€”use_mps_device: False
  • β€”seed: 42
  • β€”data_seed: None
  • β€”jit_mode_eval: False
  • β€”use_ipex: False
  • β€”bf16: False
  • β€”fp16: False
  • β€”fp16_opt_level: O1
  • β€”half_precision_backend: auto
  • β€”bf16_full_eval: False
  • β€”fp16_full_eval: False
  • β€”tf32: None
  • β€”local_rank: 0
  • β€”ddp_backend: None
  • β€”tpu_num_cores: None
  • β€”tpu_metrics_debug: False
  • β€”debug: []
  • β€”dataloader_drop_last: False
  • β€”dataloader_num_workers: 0
  • β€”dataloader_prefetch_factor: None
  • β€”past_index: -1
  • β€”disable_tqdm: False
  • β€”remove_unused_columns: True
  • β€”label_names: None
  • β€”load_best_model_at_end: False
  • β€”ignore_data_skip: False
  • β€”fsdp: []
  • β€”fsdp_min_num_params: 0
  • β€”fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • β€”tp_size: 0
  • β€”fsdp_transformer_layer_cls_to_wrap: None
  • β€”accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • β€”deepspeed: None
  • β€”label_smoothing_factor: 0.0
  • β€”optim: adamw_torch
  • β€”optim_args: None
  • β€”adafactor: False
  • β€”group_by_length: False
  • β€”length_column_name: length
  • β€”ddp_find_unused_parameters: None
  • β€”ddp_bucket_cap_mb: None
  • β€”ddp_broadcast_buffers: False
  • β€”dataloader_pin_memory: True
  • β€”dataloader_persistent_workers: False
  • β€”skip_memory_metrics: True
  • β€”use_legacy_prediction_loop: False
  • β€”push_to_hub: False
  • β€”resume_from_checkpoint: None
  • β€”hub_model_id: None
  • β€”hub_strategy: every_save
  • β€”hub_private_repo: None
  • β€”hub_always_push: False
  • β€”gradient_checkpointing: False
  • β€”gradient_checkpointing_kwargs: None
  • β€”include_inputs_for_metrics: False
  • β€”include_for_metrics: []
  • β€”eval_do_concat_batches: True
  • β€”fp16_backend: auto
  • β€”push_to_hub_model_id: None
  • β€”push_to_hub_organization: None
  • β€”mp_parameters:
  • β€”auto_find_batch_size: False
  • β€”full_determinism: False
  • β€”torchdynamo: None
  • β€”ray_scope: last
  • β€”ddp_timeout: 1800
  • β€”torch_compile: False
  • β€”torch_compile_backend: None
  • β€”torch_compile_mode: None
  • β€”include_tokens_per_second: False
  • β€”include_num_input_tokens_seen: False
  • β€”neftune_noise_alpha: None
  • β€”optim_target_modules: None
  • β€”batch_eval_metrics: False
  • β€”eval_on_start: False
  • β€”use_liger_kernel: False
  • β€”eval_use_gather_object: False
  • β€”average_tokens_across_devices: False
  • β€”prompts: None
  • β€”batch_sampler: batch_sampler
  • β€”multi_dataset_batch_sampler: round_robin

</details>

Training Logs

EpochStepTraining Loss
0.57085001.0036

Framework Versions

  • β€”Python: 3.11.12
  • β€”Sentence Transformers: 3.4.1
  • β€”Transformers: 4.51.3
  • β€”PyTorch: 2.6.0+cu124
  • β€”Accelerate: 1.6.0
  • β€”Datasets: 3.5.1
  • β€”Tokenizers: 0.21.1

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->