songphucn7/me5-checkthat-task1-v1.1
SentenceTransformer based on intfloat/multilingual-e5-large
This is a sentence-transformers model finetuned from intfloat/multilingual-e5-large. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for retrieval.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: intfloat/multilingual-e5-large <!-- at revision 3d7cfbdacd47fdda877c5cd8a79fbcc4f2a574f3 -->
- Maximum Sequence Length: 256 tokens
- Output Dimensionality: 1024 dimensions
- Similarity Function: Cosine Similarity
- Supported Modality: Text <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'XLMRobertaModel'})
(1): Pooling({'embedding_dimension': 1024, 'pooling_mode': 'mean', 'include_prompt': True})
(2): Normalize({})
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("songphucn7/me5-checkthat-task1-v1.1")
# Run inference
sentences = [
'query: Erstaunlich, dass Ärzte eine Selbstverständlichkeit verlangen müssen. The Lancet, eine der bekanntesten ärztlichen Fachmagazine des Globus: ARR (Absolute[!] Risk Reduction) für Erwachsene[!] - AstraZeneca: 1,3% - Moderna–NIH: 1,2% - BioNTech: 0,84%',
'passage: It uses the relative risk (RR)—ie, the ratio of attack rates with and without a vaccine—which is expressed as 1–RR.\n\nRanking by reported efficacy gives relative risk reductions of 95% for the Pfizer–BioNTech, 94% for the Moderna–NIH, 91% for the Gamaleya, 67% for the J&J, and 67% for the AstraZeneca–Oxford vaccines.\nHowever, RRR should be seen against the background risk of being infected and becoming ill with COVID-19, which varies between populations and over time.\nAlthough the RRR considers only participants who could benefit from the vaccine, the absolute risk reduction (ARR), which is the difference between attack rates with and without a vaccine, considers the whole population.\nARRs tend to be ignored because they give a much less impressive effect size than RRRs: 1·3% for the AstraZeneca–Oxford, 1·2% for the Moderna–NIH, 1·2% for the J&J, 0·93% for the Gamaleya, and 0·84% for the Pfizer–BioNTech vaccines.\nARR is also used to derive an estimate of vaccine effectiveness, which is the number needed to vaccinate (NNV) to prevent one more case of COVID-19 as 1/ARR.\nNNVs bring a different perspective: 81 for the Moderna–NIH, 78 for the AstraZeneca–Oxford, 108 for the Gamaleya, 84 for the J&J, and 119 for the Pfizer–BioNTech vaccines.\nThe explanation lies in the combination of vaccine efficacy and different background risks of COVID-19 across studies: 0·9% for the Pfizer–BioNTech, 1% for the Gamaleya, 1·4% for the Moderna–NIH, 1·8% for the J&J, and 1·9% for the AstraZeneca–Oxford vaccines.',
'passage: title: Meningitis due to cerebrospinal fluid leak after nasal swab testing for COVID‐19 abstract: Data sharing is not applicable to this article as no new data were created or analyzed in this study.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000, 0.7269, -0.1077],
# [ 0.7269, 1.0000, -0.1972],
# [-0.1077, -0.1972, 1.0000]])<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
Evaluation
Metrics
Information Retrieval
- Dataset:
10-percent-dev-split - Evaluated with <code>InformationRetrievalEvaluator</code>
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 17,319 training samples
- Columns: <code>sentence0</code> and <code>sentence1</code>
- Approximate statistics based on the first 1000 samples: | | sentence0 | sentence1 | |:--------|:------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 27 tokens</li><li>mean: 59.78 tokens</li><li>max: 129 tokens</li></ul> | <ul><li>min: 18 tokens</li><li>mean: 204.56 tokens</li><li>max: 256 tokens</li></ul> |
- Samples: | sentence0 | sentence1 | |:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>query: Cureus \| Consistent Use of Ivermectin as Prophylaxis for COVID-19 Resulted in a 92% Cut in COVID-19 Death Rate in a Dose‑Response Fashion: Findings from a Prospective Observational Study of a Tightly Regulated Population of 88,012 Subjects</code> | <code>passage: title: Regular Use of Ivermectin as Prophylaxis for COVID-19 Led Up to a 92% Reduction in COVID-19 Mortality Rate in a Dose-Response Manner: Results of a Prospective Observational Study of a Strictly Controlled Population of 88,012 Subjects abstract: Background We have previously demonstrated that ivermectin used as prophylaxis for coronavirus disease 2019 (COVID-19), irrespective of the regularity, in a strictly controlled citywide program in Southern Brazil (Itajaí, Brazil), was associated with reductions in COVID-19 infection, hospitalization, and mortality rates. In this study, our objective was to determine if the regular use of ivermectin impacted the level of protection from COVID-19 and related outcomes, reinforcing the efficacy of ivermectin through the demonstration of a dose-response effect. Methods This exploratory analysis of a prospective observational study involved a program that used ivermectin at a dose of 0. 2 mg/kg/day for two consecutive days, every 15 day...</code> | | <code>query: Getting back to sport choices after a sharp lateral ankle sprain injury: unveiling the PAASS structure—an international multidisciplinary consensus 👀👀👇👇🦶🦶</code> | <code>passage: title: Return to sport decisions after an acute lateral ankle sprain injury: introducing the PAASS framework—an international multidisciplinary consensus abstract: Background Despite being the most commonly incurred sports injury with a high recurrence rate, there are no guidelines to inform return to sport (RTS) decisions following acute lateral ankle sprain injuries. We aimed to develop a list of assessment items to address this gap. Methods We used a three-round Delphi survey approach to develop consensus of opinion among 155 globally diverse health professionals working in elite field or court sports. This involved surveys that were structured in question format with both closed-response and open-response options. We asked panellists to indicate their agreement about whether or not assessment items should support the RTS decision after an acute lateral ankle sprain injury. The second and third round surveys included quantitative and qualitative feedback from the previous r...</code> | | <code>query: Petit rappel ! ⬇️ « En conclusion, la vaccination contre la COVID-19 est un facteur de risque majeur d'infection chez les patients gravement malades. »</code> | <code>passage: title: Adverse effects of COVID-19 vaccines and measures to prevent them abstract: Abstract Recently, The Lancet published a study on the effectiveness of COVID-19 vaccines and the waning of immunity with time. The study showed that immune function among vaccinated individuals 8 months after the administration of two doses of COVID-19 vaccine was lower than that among the unvaccinated individuals. According to European Medicines Agency recommendations, frequent COVID-19 booster shots could adversely affect the immune response and may not be feasible. The decrease in immunity can be caused by several factors such as N1-methylpseudouridine, the spike protein, lipid nanoparticles, antibody-dependent enhancement, and the original antigenic stimulus. These clinical alterations may explain the association reported between COVID-19 vaccination and shingles. As a safety measure, further booster vaccinations should be discontinued. In addition, the date of vaccination should be recorde...</code> |
- Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim",
"gather_across_devices": false,
"directions": [
"query_to_doc"
],
"partition_mode": "joint",
"hardness_mode": null,
"hardness_strength": 0.0
}Training Hyperparameters
Non-Default Hyperparameters
per_device_train_batch_size: 32num_train_epochs: 10eval_strategy: stepsper_device_eval_batch_size: 32multi_dataset_batch_sampler: round_robin
All Hyperparameters
<details><summary>Click to expand</summary>
per_device_train_batch_size: 32num_train_epochs: 10max_steps: -1learning_rate: 5e-05lr_scheduler_type: linearlr_scheduler_kwargs: Nonewarmup_steps: 0optim: adamwtorchfusedoptim_args: Noneweight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08optim_target_modules: Nonegradient_accumulation_steps: 1average_tokens_across_devices: Truemax_grad_norm: 1label_smoothing_factor: 0.0bf16: Falsefp16: Falsebf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Nonetorch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneuse_liger_kernel: Falseliger_kernel_config: Noneuse_cache: Falseneftune_noise_alpha: Nonetorch_empty_cache_steps: Noneauto_find_batch_size: Falselog_on_each_node: Truelogging_nan_inf_filter: Trueinclude_num_input_tokens_seen: nolog_level: passivelog_level_replica: warningdisable_tqdm: Falseproject: huggingfacetrackio_space_id: trackioeval_strategy: stepsper_device_eval_batch_size: 32prediction_loss_only: Trueeval_on_start: Falseeval_do_concat_batches: Trueeval_use_gather_object: Falseeval_accumulation_steps: Noneinclude_for_metrics: []batch_eval_metrics: Falsesave_only_model: Falsesave_on_each_node: Falseenable_jit_checkpoint: Falsepush_to_hub: Falsehub_private_repo: Nonehub_model_id: Nonehub_strategy: every_savehub_always_push: Falsehub_revision: Noneload_best_model_at_end: Falseignore_data_skip: Falserestore_callback_states_from_checkpoint: Falsefull_determinism: Falseseed: 42data_seed: Noneuse_cpu: Falseaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}parallelism_config: Nonedataloader_drop_last: Falsedataloader_num_workers: 0dataloader_pin_memory: Truedataloader_persistent_workers: Falsedataloader_prefetch_factor: Noneremove_unused_columns: Truelabel_names: Nonetrain_sampling_strategy: randomlength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falseddp_backend: Noneddp_timeout: 1800fsdp: []fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}deepspeed: Nonedebug: []skip_memory_metrics: Truedo_predict: Falseresume_from_checkpoint: Nonewarmup_ratio: Nonelocal_rank: -1prompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robinrouter_mapping: {}learning_rate_mapping: {}
</details>
Training Logs
Training Time
- Training: 2.4 hours
Framework Versions
- Python: 3.12.6
- Sentence Transformers: 5.4.0
- Transformers: 5.5.3
- PyTorch: 2.11.0+cu130
- Accelerate: 1.10.1
- Datasets: 4.8.4
- Tokenizers: 0.22.2
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}MultipleNegativesRankingLoss
@misc{oord2019representationlearningcontrastivepredictive,
title={Representation Learning with Contrastive Predictive Coding},
author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
year={2019},
eprint={1807.03748},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/1807.03748},
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
