CoolFace
Modelpublic

iambestfeed/phobert-base-v2-finetuned-filtered_data_wseg-lr5e-06-1-epochs-bs-16

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes95downloads
Model Card

SentenceTransformer based on vinai/phobert-base-v2

This is a sentence-transformers model finetuned from vinai/phobert-base-v2 on the vnexpress-data-similarity dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: vinai/phobert-base-v2 <!-- at revision e2375d266bdf39c6e8e9a87af16a5da3190b0cc8 -->
  • —Maximum Sequence Length: 256 tokens
  • —Output Dimensionality: 768 dimensions
  • —Similarity Function: Cosine Similarity
  • —Training Dataset:
  • —vnexpress-data-similarity <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 256, 'do_lower_case': False}) with Transformer model: RobertaModel 
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("iambestfeed/phobert-base-v2-finetuned-filtered_data_wseg-lr5e-06")
# Run inference
sentences = [
    'Động_đất mạnh 5,6 độ ở tỉnh Tây_Java của Indonesia khiến ít_nhất 44 người thiệt_mạng , nhiều nhà_cửa tại thủ_đô Jakarta cũng bị rung_chuyển . - VnExpress',
    'Cơ_quan Khí_tượng , Khí_hậu và Địa_vật_lý Indonesia ( BMKG ) cho biết trận động_đất mạnh 5,6 độ xảy ra lúc 13h21 hôm_nay , tâm chấn nằm trên đất_liền ở độ sâu 10 km tại tỉnh Tây_Java , cách thủ_đô Jakarta khoảng 100 km về phía nam . Trong khi đó , Cơ_quan Khảo_sát Địa_chất Mỹ ( USGS ) cho biết trận động_đất mạnh 5,4 độ . USGS xác_định những trận động_đất có tâm chấn ở độ sâu dưới 70 km là " tâm chấn nông " và chúng thường gây thiệt_hại nặng_nề hơn so với " tâm chấn sâu " . \n BMKG nói rằng không có nguy_cơ xảy ra sóng_thần sau trận động_đất . \n " Hàng trăm , thậm_chí có_thể hàng nghìn ngôi nhà bị hư_hại . 44 người đã thiệt_mạng " , Adam , phát_ngôn_viên của chính_quyền địa_phương tại thị_trấn Cianjur ở Tây_Java , nói với AFP. \n Herman_Suherman , lãnh_đạo thị_trấn Cianjur , trước đó nói trên truyền_hình rằng chỉ riêng một trong 4 bệnh_viện tại đây đã ghi_nhận gần 20 người chết và ít_nhất 300 người bị_thương . " Phần_lớn nạn_nhân bị_thương do mắc_kẹt dưới đống đổ_nát " , ông cho biết và cảnh_báo con_số thương_vong có_thể tăng thêm . \n Nhân_chứng cho biết nhiều tòa nhà tại Jakarta bị rung_chuyển , khiến không ít người phải sơ_tán khỏi các văn_phòng ở quận trung_tâm thủ_đô . Số khác nói rằng họ cảm_nhận được các tòa nhà rung lắc và chứng_kiến đồ_đạc di_chuyển . \n " Tôi đang làm_việc thì sàn nhà rung_chuyển , tôi có_thể cảm_nhận chấn_động rất rõ_ràng " , Mayadita_Waluyo , một luật_sư 22 tuổi làm_việc tại một tòa nhà_văn_phòng ở Jakarta , kể . " Lúc đó tôi chưa hình_dung được chuyện gì đã xảy ra , nhưng rung lắc ngày_càng mạnh hơn và kéo_dài khá lâu " . \n Waluyo cho hay cô và nhiều đồng_nghiệp trong văn_phòng sau đó đã hoảng_hốt chạy xuống cầu_thang bộ để thoát khỏi tòa nhà . \n Indonesia thường_xuyên trải qua các trận động_đất và phun trào núi_lửa do nằm trên " Vành_đai lửa " , một vòng_cung hoạt_động địa_chấn dữ_dội , nơi các mảng kiến_tạo va_chạm kéo_dài từ Nhật_Bản qua Đông_Nam_Á và lưu_vực Thái_Bình_Dương . \n Năm 2004 , trận động_đất 9,1 độ xảy ra ngoài khơi đảo Sumatra , gây ra sóng_thần giết chết 220.000 người khắp khu_vực , trong đó có khoảng 170.000 người ở Indonesia . \n Năm 2018 , một trận động_đất 7,5 độ và sóng_thần theo sau ở Palu trên đảo Sulawesi đã khiến hơn 4.300 người chết và mất_tích . \n Vũ_Anh ( Theo AFP )',
    'Nếu mua xe mới tôi phải vay thêm tiền mới đủ lăn bánh_xe cỡ A trong khi mua xe cũ thì có thế chọn xe cỡ B. ( Lan_Anh ) \n >> Xem thêm : cẩm_nang mua xe',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

vnexpress-data-similarity
  • —Dataset: vnexpress-data-similarity at e6cb542
  • —Size: 74,620 training samples
  • —Columns: <code>anchor</code> and <code>positive</code>
  • —Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 18 tokens</li><li>mean: 32.21 tokens</li><li>max: 104 tokens</li></ul> | <ul><li>min: 21 tokens</li><li>mean: 140.01 tokens</li><li>max: 256 tokens</li></ul> |
  • —Samples: | anchor | positive | |:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Gene quyếtđịnh phầnlớn chiều cao của trẻ , những thóiquen sinhhoạt , ănuống cũng ảnhhưởng đến quátrình pháttriển . - VnExpress</code> | <code>Gene quyếtđịnh phầnlớn chiều cao của trẻ , những thóiquen sinhhoạt , ănuống cũng ảnhhưởng đến quátrình pháttriển . <br> Lê Nguyễn ( Theo Timesofindia , WebMD ) <br> Vuilòng điền đầyđủ thôngtin để chúngtôi cóthể hỗtrợ bạn</code> | | <code>H & # 039 ; Hen Niê đưa bốmẹ và cháu gái cùng đi côngtác , VũThuPhương đưa con sang Anh nhậphọc , PhươngTrinh Jolie bế contrai mới sinh , Mạc VănKhoa xúcđộng khi được diễn chung với NSƯT HoàiLinh . - Ngôisao</code> | <code>PhongKiều <br> - Showbiz <br> - Thờitrang <br> - Làmđẹp <br> - Xem <br> - Ănchơi <br> - Lốisống <br> - Thểthao <br> - Thờicuộc <br> - Podcasts <br> - Thươngtrường <br> - Trắcnghiệm <br> - Video <br> - Ảnh <br> - Reviews & Deals <br> H ' Hen Niê đưa bốmẹ và cháu gái cùng đi côngtác , VũThuPhương đưa con sang Anh nhậphọc , PhươngTrinh Jolie bế contrai mới sinh , Mạc VănKhoa xúcđộng khi được diễn chung với NSƯT HoàiLinh . <br> PhongKiều</code> | | <code>Hạtgiống số một giành chiếnthắng với tỷsố 7 - 6 ( 1 ) , 6 - 3 ở trận raquân tại Halle Mởrộng 2019 . - VnExpress</code> | <code>Hạtgiống số một giành chiếnthắng với tỷsố 7 - 6 ( 1 ) , 6 - 3 ở trận raquân tại Halle Mởrộng 2019 . <br> QuangDũng</code> |
  • —Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "cos_sim"
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —per_device_train_batch_size: 16
  • —gradient_accumulation_steps: 5
  • —learning_rate: 5e-06
  • —warmup_ratio: 0.1
  • —save_safetensors: False
  • —fp16: True
  • —push_to_hub: True
  • —hub_model_id: iambestfeed/phobert-base-v2-finetuned-filtereddatawseg-lr5e-06
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: no
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 16
  • —per_device_eval_batch_size: 8
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 5
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 5e-06
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 3
  • —max_steps: -1
  • —lr_scheduler_type: linear
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: False
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: False
  • —fp16: True
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: True
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: False
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamw_torch
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: True
  • —resume_from_checkpoint: None
  • —hub_model_id: iambestfeed/phobert-base-v2-finetuned-filtereddatawseg-lr5e-06
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —dispatch_batches: None
  • —split_batches: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional

</details>

Training Logs

<details><summary>Click to expand</summary>

EpochStepTraining Loss
0.0215100.6363
0.0429200.5673
0.0644300.4673
0.0858400.3299
0.1073500.1847
0.1287600.0889
0.1502700.036
0.1716800.0107
0.1931900.0074
0.21451000.0049
0.23601100.0036
0.25741200.0036
0.27891300.0027
0.30031400.0018
0.32181500.0049
0.34321600.0037
0.36471700.0019
0.38611800.002
0.40761900.0014
0.42902000.0013
0.45052100.0013
0.47192200.0013
0.49342300.0027
0.51482400.0017
0.53632500.0012
0.55772600.0009
0.57922700.0008
0.60062800.0016
0.62212900.0014
0.64353000.0006
0.66503100.0018
0.68643200.0005
0.70793300.0006
0.72933400.0006
0.75083500.0005
0.77223600.0009
0.79373700.0012
0.81513800.0003
0.83663900.0005
0.85804000.0011
0.87954100.0005
0.90094200.0006
0.92244300.0012
0.94384400.0007
0.96534500.0004
0.98674600.0003
1.00644700.0007
1.02794800.0004
1.04934900.0006
1.07085000.0003
1.09225100.0008
1.11375200.0004
1.13515300.0013
1.15665400.0003
1.17805500.0003
1.19955600.0003
1.22095700.0003
1.24245800.0003
1.26385900.0002
1.28536000.0002
1.30676100.0004
1.32826200.0003
1.34966300.0002
1.37116400.0002
1.39256500.0003
1.41406600.0003
1.43546700.0002
1.45696800.0001
1.47836900.0002
1.49987000.0011
1.52127100.0006
1.54277200.0003
1.56417300.0002
1.58567400.0002
1.60707500.0003
1.62857600.0003
1.64997700.0002
1.67147800.0002
1.69287900.0002
1.71438000.0002
1.73578100.0002
1.75728200.0002
1.77868300.0002
1.80018400.0002
1.82158500.0001
1.84308600.0003
1.86448700.0003
1.88598800.0003
1.90738900.0003
1.92889000.0003
1.95029100.0002
1.97179200.0002
1.99319300.0003
2.01299400.0002
2.03439500.0006
2.05589600.0001
2.07729700.0002
2.09879800.0002
2.12019900.0003
2.141610000.0008
2.163010100.0001
2.184510200.0002
2.205910300.0002
2.227410400.0002
2.248810500.0001
2.270310600.0002
2.291710700.0001
2.313210800.0002
2.334610900.0004
2.356111000.0001
2.377511100.0002
2.399011200.0002
2.420411300.0002
2.441911400.0001
2.463311500.0001
2.484811600.0007
2.506211700.0003
2.527711800.0002
2.549111900.0001
2.570612000.0001
2.592012100.0002
2.613512200.0002
2.634912300.0001
2.656412400.0002
2.677812500.0001
2.699312600.0002
2.720712700.0002
2.742212800.0001
2.763612900.0001
2.785113000.0002
2.806513100.0001
2.828013200.0002
2.849413300.0002
2.870913400.0002
2.892313500.0002
2.913813600.0002
2.935213700.0001
2.956713800.0001
2.978113900.0002

</details>

Framework Versions

  • —Python: 3.10.12
  • —Sentence Transformers: 3.3.1
  • —Transformers: 4.47.0
  • —PyTorch: 2.5.1+cu121
  • —Accelerate: 1.2.1
  • —Datasets: 3.3.1
  • —Tokenizers: 0.21.0

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->