ngocnamk3er/GreenNode-Embedding-Large-VN-Retrieval-fine-tuned-v2
SentenceTransformer based on GreenNode/GreenNode-Embedding-Large-VN-Mixed-V1
This is a sentence-transformers model finetuned from GreenNode/GreenNode-Embedding-Large-VN-Mixed-V1. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: GreenNode/GreenNode-Embedding-Large-VN-Mixed-V1 <!-- at revision 1d3dddb3862292dab4bd3eddf0664c0335ad5843 -->
- Maximum Sequence Length: 8192 tokens
- Output Dimensionality: 1024 dimensions
- Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 8192, 'do_lower_case': False, 'architecture': 'XLMRobertaModel'})
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("ngocnamk3er/GreenNode-Embedding-Large-VN-Retrieval-fine-tuned-v2")
# Run inference
sentences = [
'Khóa học này có bao nhiêu buổi học ngữ pháp liên quan đến cụm động từ?',
'| **Ngữ pháp** | **60-75 phút** | Giới từ và Cụm giới từ (Bài tập Ứng dụng – Buổi 3) Giới từ và Cụm giới từ (Bài tập Ứng dụng – Buổi 4) **Ngữ pháp Ứng dụng (2026) | Chuyên đề 6** | Học từ vựng và Tập dịch lại toàn bộ bài trong buổi học Hoàn thành bài thi Online |\n| **4** | **Ôn tập TV ngày 1;2;3** | **15-20 phút** | **EZ Vocab độc quyền** | |\n| **Từ vựng mới** | **30 phút** | **EZ Vocab độc quyền** | |\n| **Từ vựng** | **60-75 phút** | Bài 21,22: Mở rộng kiến thức Anh văn 12: Chủ đề 11: Relationships (Buổi 1 và Buổi 2) **Nền tảng Từ vựng – Đọc hiểu | Chuyên đề 1** | Học từ vựng và Tập dịch lại toàn bộ bài trong buổi học |\n| **5** | **Ôn tập TV ngày 1,2,3,4** | **15-20 phút** | **EZ Vocab độc quyền** | |\n| **Từ vựng mới** | **30 phút** | **EZ Vocab độc quyền** | |\n| **Ngữ pháp** | **45-60 phút** | Cụm động từ (Buổi 1) Cụm động từ (Buổi 2) **Ngữ pháp Ứng dụng (2026) | Chuyên đề 7** | Học từ vựng và Tập dịch lại toàn bộ bài trong buổi học Hoàn thành bài thi Online |\n| **6** | **Ôn tập TV ngày 1,2,3,4,5** | **15 phút** | **EZ Vocab độc quyền** | |\n| **Từ vựng mới** | **30 phút** | **EZ Vocab độc quyền** | |\n| **Từ vựng** | **60-75 phút** | Bài 23,24: Mở rộng kiến thức Anh văn 12: Chủ đề 12: Problems (Buổi 1 và Buổi 2) **Nền tảng Từ vựng – Đọc hiểu | Chuyên đề 1** | Học từ vựng và Tập dịch lại toàn bộ bài trong buổi học |\n| **7** | **Ôn tập từ vựng** **Các ngày 1-6** | **20-25 phút** | **EZ Vocab độc quyền** | |\n| **Học từ vựng mới** | **30 phút** | **EZ Vocab độc quyền** | |\n| **Ngữ pháp** | **45-60 phút** | Cụm động từ (Buổi 3) Cụm động từ (Buổi 4) **Ngữ pháp Ứng dụng (2026) | Chuyên đề 7** | Học từ vựng và Tập dịch lại toàn bộ bài trong buổi học Hoàn thành bài thi Online |\n\n**GIAI ĐOẠN 1: TUẦN SỐ 05**\n\n| | | | | |\n| --- | --- | --- | --- | --- |\n| **Ngày** | **Kiến thức** | **Thời lượng** | **Khóa học/ Bài giảng đi kèm** | **Lưu ý/ Bài tập đi kèm** |\n| **1** | **Từ vựng** | **30-45 phút** | **EZ Vocab độc quyền** | Học theo thứ tự ưu tiên mà khóa đã sắp xếp |\n| **Ngữ pháp** | **60-75 phút** | Cụm động từ (Bài tập Ứng dụng – Buổi 1) Cụm động từ (Bài tập Ứng dụng – Buổi 2) **Ngữ pháp Ứng dụng (2026) | Chuyên đề 7** | Học từ vựng và Tập dịch lại toàn bộ bài trong buổi học Hoàn thành bài thi Online |\n| **2** | **Ôn tập TV ngày 1** | **15 phút** | **EZ Vocab độc quyền** | |\n| **Từ vựng mới** | **30 phút** | **EZ Vocab độc quyền** | |\n| **Từ vựng** | **60-75 phút** | Bài 25,26: Mở rộng kiến thức Anh văn 12: Chủ đề 13: Holidays and Tourism (Buổi 1 và Buổi 2) **Nền tảng Từ vựng – Đọc hiểu | Chuyên đề 1** | Học từ vựng và Tập dịch lại toàn bộ bài trong buổi học |',
'TACMP- TÍNH NĂNG BÀI GIẢNG DẠNG BÀI BÁO\nI. Mục tiêu\nCung cấp cho học viên một hình thức học từ vựng và luyện dịch thông qua bài báo, trích dẫn từ tạp chí, sách chuyên khảo, v.v.\n Tính năng giúp học viên:\nRèn luyện kỹ năng đọc hiểu học thuật.\nGhi nhớ và luyện tập từ vựng chuyên sâu.\nLuyện dịch song ngữ và phản xạ tiếng Anh.\nII. Các tính năng chính cho học viên\n1. 📰 Đọc bài báo – Học và ghi nhớ từ vựng\nHiển thị nội dung bài báo gốc.\nCho phép học viên đọc, tra cứu, và tương tác với văn bản.\n[[img_link: https://imagedoan.s3.ap-southeast-2.amazonaws.com/docx_extracted/4b875d5e-65a4-4713-a95d-54427904badc.png]]\n2. 📌 Ghi nhớ từ vựng\nTự động hiển thị bảng từ vựng cần lưu ý trong bài.\nCho phép:\nHighlight từ vựng trong nội dung bài báo (2 chiều với bảng từ).\nClick từ trong bảng từ vựng để scroll đến từ highlight trong bài tương ứng và ngược lại.\n[[img_link: https://imagedoan.s3.ap-southeast-2.amazonaws.com/docx_extracted/4d0ceb3c-4283-4a32-9558-f5a558ebc86a.png]]\n3. 🎮 Luyện tập chuyên sâu từ vựng (EzVocab)\nCho phép học viên luyện tập với các từ vựng trong bài báo thông qua:\nTrò chơi ghi nhớ.\nFlashcard tương tác.\nQuiz từ vựng – câu hỏi chọn nghĩa, điền từ, matching, v.v.\n[[img_link: https://imagedoan.s3.ap-southeast-2.amazonaws.com/docx_extracted/91a37724-fc5a-4b0f-9c59-33adc11d87ee.png]]\n4. Tập dịch bài báo (Tập dịch thông minh)\nCho phép luyện kỹ năng dịch với:\nChế độ ẩn/hiện câu tiếng Việt để kiểm tra khả năng hiểu và dịch của học viên.\nHiển thị từng câu song ngữ để học viên đối chiếu, phân tích.\n[[img_link: https://imagedoan.s3.ap-southeast-2.amazonaws.com/docx_extracted/5c402b78-f7c2-45ca-8e68-a4f11a15f971.png]]',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.6752, 0.0323],
# [0.6752, 1.0000, 0.1196],
# [0.0323, 0.1196, 1.0000]])<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
Evaluation
Metrics
Information Retrieval
- Dataset:
val_hit_at_k - Evaluated with <code>InformationRetrievalEvaluator</code>
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 1,032 training samples
- Columns: <code>sentence0</code> and <code>sentence1</code>
- Approximate statistics based on the first 1000 samples: | | sentence0 | sentence1 | |:--------|:-----------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 14 tokens</li><li>mean: 23.22 tokens</li><li>max: 45 tokens</li></ul> | <ul><li>min: 25 tokens</li><li>mean: 437.41 tokens</li><li>max: 2694 tokens</li></ul> |
- Samples: | sentence0 | sentence1 | |:---------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Làm thế nào để tôi có thể tiếp cận và học theo thứ tự ưu tiên mà khóa học đã sắp xếp?</code> | <code>\| 4 \| Ôn tập TV ngày 1;2;3 \| 15-20 phút \| EZ Vocab độc quyền \| \|<br>\| Từ vựng mới \| 30 phút \| EZ Vocab độc quyền \| \|<br>\| Từ vựng \| 60-75 phút \| Bài 3,4: Mở rộng kiến thức Anh văn 12: Chủ đề 2: Cultural diversity (Buổi 1 và Buổi 2) Nền tảng Từ vựng – Đọc hiểu \| Chuyên đề 1 \| Học từ vựng và Tập dịch lại toàn bộ bài trong buổi học \|<br>\| 5 \| Ôn tập TV ngày 1,2,3,4 \| 15-20 phút \| EZ Vocab độc quyền \| \|<br>\| Từ vựng mới \| 30 phút \| EZ Vocab độc quyền \| \|<br>\| Ngữ pháp \| 45-60 phút \| Từ loại (Bài tập Ứng dụng - Buổi 1) Từ loại (Bài tập Ứng dụng - Buổi 2) Ngữ pháp Ứng dụng (2026) \| Chuyên đề 1 \| Học từ vựng và Tập dịch lại toàn bộ bài trong buổi học Hoàn thành bài thi Online \|<br>\| 6 \| Ôn tập TV ngày 1,2,3,4,5 \| 15 phút \| EZ Vocab độc quyền \| \|<br>\| Từ vựng mới \| 30 phút \| EZ Vocab độc quyền \| \|<br>\| Từ vựng \| 60-75 phút \| Bài 5,6: Mở rộng kiến thức Anh văn 12: Chủ đề 3: Going green (Buổi ...</code> | | <code>Làm thế nào để em truy cập vào tài liệu chung của buổi học livestream ạ?</code> | <code>Câu hỏi: Link tài liệu của buổi live<br>Trả lời: Em vào khóa học xong vào chỗ tiện ích nhé. Ví dụ hôm nay Cô live khóa Pro3M plus, em sẽ vào khóa đang livestream là khoá A. Chọn mục tài liệu chỗ tiện ích (góc phải laptop), chọn tài liệu chung thì sẽ có file nhé.</code> | | <code>Em có cần lưu ý gì về số lượng thiết bị tối đa mà em có thể đăng nhập không ạ?</code> | <code>Câu hỏi: Một tài khoản có thể vào được nhiều máy không?<br>Trả lời: Em có thể đăng nhập vào 1 thiết bị điện thoại và 1 thiết bị máy tính, chỉ sử dụng trên 1 thiết bị tại cùng thời điểm thôi nha em. Mình chú ý chỉ được đăng nhập tối đa 3 thiết bị đề không bị khóa tài khoản em nhé.</code> |
- Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim",
"gather_across_devices": false
}Training Hyperparameters
Non-Default Hyperparameters
eval_strategy: stepsper_device_train_batch_size: 4per_device_eval_batch_size: 4fp16: Truemulti_dataset_batch_sampler: round_robin
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 4per_device_eval_batch_size: 4per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 3max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Truefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Falsehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseeval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseeval_use_gather_object: Falseprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robinrouter_mapping: {}learning_rate_mapping: {}
</details>
Training Logs
Framework Versions
- Python: 3.10.12
- Sentence Transformers: 5.2.0
- Transformers: 4.44.2
- PyTorch: 2.9.1+cu128
- Accelerate: 1.12.0
- Datasets: 4.4.1
- Tokenizers: 0.19.1
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}MultipleNegativesRankingLoss
@misc{henderson2017efficient,
title={Efficient Natural Language Response Suggestion for Smart Reply},
author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
year={2017},
eprint={1705.00652},
archivePrefix={arXiv},
primaryClass={cs.CL}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
