Stergios-Konstantinidis/MNLP_M3_tokenizer_tuned
SentenceTransformer based on sucharush/e5stemfinetuned
This is a sentence-transformers model finetuned from sucharush/e5_stem_finetuned. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: sucharush/e5_stem_finetuned <!-- at revision 3013cc95b53a83082746afd556743e81f1da5dff -->
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 768 dimensions
- Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("Stergios-Konstantinidis/MNLP_M3_tokenizer_tuned")
# Run inference
sentences = [
'[{\'from\': \'human\', \'value\': \'Solve the following math problem step-by-step.\\nSimplify your answer as much as possible. Present your final answer as \\\\boxed{Your Answer}.\\nA pentagon is drawn by placing an isosceles right triangle on top of a square as pictured. What percent of the area of the pentagon is the area of the right triangle?\\n\\n[asy]\\nsize(50);\\ndraw((0,0)--(0,-1)--(1,-1)--(1,0)--(0,0)--(.5,.5)--(1,0));\\n[/asy]\'}, {\'from\': \'gpt\', \'value\': "Step 1: We are given a pentagon formed by placing an isosceles right triangle on top of a square.\\nStep 2: To find the percentage of the area of the pentagon that is the right triangle, we need to consider the areas of both shapes and then calculate the ratio of the triangle\'s area to the total pentagon\'s area.\\nStep 3: Let\'s designate the leg length of the isosceles right triangle as $x$. This means the hypotenuse of the triangle has a length of $x\\\\sqrt{2}$.\\nStep 4: Since the hypotenuse of the triangle is a side of the square, the area of the square is $(x\\\\sqrt{2})^2 = 2x^2$. \\nStep 5: On the other hand, the area of the triangle is $(x)(x)/2 = x^2/2$. This is obtained using the standard formula for the area of a triangle, which is base times height divided by 2.\\nStep 6: So, the area of the pentagon is the sum of the areas of the triangle and the square, which gives us $\\\\frac{x^2}{2} + 2x^2 = \\\\frac{5x^2}{2}$.\\nStep 7: Now, to find the percentage of the pentagon\'s area that is inside the triangle, we calculate the ratio of the triangle\'s area to the total pentagon\'s area. This can be computed as $\\\\frac{x^2/2}{5x^2/2} =\\\\frac{x^2}{2}\\\\cdot \\\\frac{2}{5x^2} = \\\\frac15$.\\nStep 8: Therefore, the percentage of the area of the pentagon that is the right triangle is $\\\\frac{1}{5}$, which is equivalent to $\\\\boxed{20\\\\%}$.\\nStep 9: As an alternate solution, we can consider drawing the two diagonals of the square and observe the resulting geometric properties to arrive at the same conclusion.\\n\\nAnswer:\\n\\\\boxed{20\\\\%}"}]',
'[{\'from\': \'human\', \'value\': \'Solve the following math problem step-by-step.\\nSimplify your answer as much as possible. Present your final answer as \\\\boxed{Your Answer}.\\nA pentagon is drawn by placing an isosceles right triangle on top of a square as pictured. What percent of the area of the pentagon is the area of the right triangle?\\n\\n[asy]\\nsize(50);\\ndraw((0,0)--(0,-1)--(1,-1)--(1,0)--(0,0)--(.5,.5)--(1,0));\\n[/asy]\'}, {\'from\': \'gpt\', \'value\': "Step 1: We are given a pentagon formed by placing an isosceles right triangle on top of a square.\\nStep 2: To find the percentage of the area of the pentagon that is the right triangle, we need to consider the areas of both shapes and then calculate the ratio of the triangle\'s area to the total pentagon\'s area.\\nStep 3: Let\'s designate the leg length of the isosceles right triangle as $x$. This means the hypotenuse of the triangle has a length of $x\\\\sqrt{2}$.\\nStep 4: Since the hypotenuse of the triangle is a side of the square, the area of the square is $(x\\\\sqrt{2})^2 = 2x^2$. \\nStep 5: On the other hand, the area of the triangle is $(x)(x)/2 = x^2/2$. This is obtained using the standard formula for the area of a triangle, which is base times height divided by 2.\\nStep 6: So, the area of the pentagon is the sum of the areas of the triangle and the square, which gives us $\\\\frac{x^2}{2} + 2x^2 = \\\\frac{5x^2}{2}$.\\nStep 7: Now, to find the percentage of the pentagon\'s area that is inside the triangle, we calculate the ratio of the triangle\'s area to the total pentagon\'s area. This can be computed as $\\\\frac{x^2/2}{5x^2/2} =\\\\frac{x^2}{2}\\\\cdot \\\\frac{2}{5x^2} = \\\\frac15$.\\nStep 8: Therefore, the percentage of the area of the pentagon that is the right triangle is $\\\\frac{1}{5}$, which is equivalent to $\\\\boxed{20\\\\%}$.\\nStep 9: As an alternate solution, we can consider drawing the two diagonals of the square and observe the resulting geometric properties to arrive at the same conclusion.\\n\\nAnswer:\\n\\\\boxed{20\\\\%}"}]',
'Question: where are the basilar membrane and the organ of corti located, Answer: Organ of Corti The organ of Corti is located in the scala media of the cochlea of the inner ear between the vestibular duct and the tympanic duct and is composed of mechanosensory cells, known as hair cells.[2] Strategically positioned on the basilar membrane of the organ of Corti are three rows of outer hair cells (OHCs) and one row of inner hair cells (IHCs).[4] Separating these hair cells are supporting cells: Deiters cells, also called phalangeal cells, which separate and support both the OHCs and the IHCs.[4]',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 113,450 training samples
- Columns: <code>sentence0</code>, <code>sentence1</code>, and <code>label</code>
- Approximate statistics based on the first 1000 samples: | | sentence0 | sentence1 | label | |:--------|:-------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------|:------------------------------------------------| | type | string | string | int | | details | <ul><li>min: 17 tokens</li><li>mean: 283.42 tokens</li><li>max: 512 tokens</li></ul> | <ul><li>min: 16 tokens</li><li>mean: 281.46 tokens</li><li>max: 512 tokens</li></ul> | <ul><li>0: ~80.00%</li><li>1: ~20.00%</li></ul> |
- Samples: | sentence0 | sentence1 | label | |:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------| | <code>Question: where does rasin in the sun take place, Answer: A Raisin in the Sun A Raisin in the Sun is a play by Lorraine Hansberry that debuted on Broadway in 1959.[1] The title comes from the poem "Harlem" (also known as "A Dream Deferred"[2]) by Langston Hughes. The story tells of a black family's experiences in the Washington Park Subdivision of Chicago's Woodlawn neighborhood as they attempt to "better" themselves with an insurance payout following the death of the father. The New York Drama Critics' Circle named it the best play of 1959.</code> | <code>Question: where does rasin in the sun take place, Answer: A Raisin in the Sun A Raisin in the Sun is a play by Lorraine Hansberry that debuted on Broadway in 1959.[1] The title comes from the poem "Harlem" (also known as "A Dream Deferred"[2]) by Langston Hughes. The story tells of a black family's experiences in the Washington Park Subdivision of Chicago's Woodlawn neighborhood as they attempt to "better" themselves with an insurance payout following the death of the father. The New York Drama Critics' Circle named it the best play of 1959.</code> | <code>1</code> | | <code>Question: when does the movie midnight sun come out, Answer: Midnight Sun (2018 film) Midnight Sun is a 2018 American romantic drama film directed by Scott Speer and written by Eric Kirsten, based on the 2006 Japanese film of the same name. The film stars Bella Thorne, Patrick Schwarzenegger, and Rob Riggle, and follows a teenage girl with the disease xeroderma pigmentosum, which prevents her from going out into sunlight. When she meets a boy, she struggles to decide whether to tell him about her condition or pretend to live a normal life. Principal photography began on October 12, 2015 in Vancouver, British Columbia, Canada. The film was released in the United States on March 23, 2018.</code> | <code>Question: who sang the song unchained melody in the movie ghost, Answer: Unchained Melody "Unchained Melody" reappeared on the Billboard charts in 1990 after The Righteous Brothers' recording was used in the box office blockbuster film Ghost. Two versions charted in the US that year – the original and a new recording. According to Medley, he was interested in having the original recording released due to the renewed interest in the song, but was told that there were licensing issues. Although Hatfield was no longer in the same condition vocally as when he first recorded the song, they decided to re-record the song for Curb Records.[44] The re-recorded version was released as both a cassette single[45] and a CD single.[46] It received minimal airplay but recorded excellent sales, peaking at number 19. The re-recorded version was certified Platinum by the RIAA on January 10, 1991,[47] and received a Grammy Award nomination.[48]</code> | <code>0</code> | | <code>Question: where is the o2 arena in london located, Answer: The O2 Arena The O2 Arena (temporarily the sponsor neutral "North Greenwich Arena", during the 2012 Summer Olympics and 2012 Summer Paralympics), is a multi purpose indoor arena located in the centre of The O2 entertainment complex on the Greenwich Peninsula in south east London.</code> | <code>[{'from': 'human', 'value': "Explain how Archimedes' Principle enables objects to float in water."}, {'from': 'gpt', 'value': "Archimedes' Principle states that the buoyant force acting on an object submerged in water is equal to the weight of the water displaced by the object. If the buoyant force exceeds the weight of the object, it will float."}]</code> | <code>0</code> |
- Loss: <code>ContrastiveTensionLoss</code>
Training Hyperparameters
Non-Default Hyperparameters
per_device_train_batch_size: 10per_device_eval_batch_size: 10num_train_epochs: 10multi_dataset_batch_sampler: round_robin
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: noprediction_loss_only: Trueper_device_train_batch_size: 10per_device_eval_batch_size: 10per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 10max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robin
</details>
Training Logs
<details><summary>Click to expand</summary>
</details>
Framework Versions
- Python: 3.12.8
- Sentence Transformers: 3.4.1
- Transformers: 4.52.4
- PyTorch: 2.6.0+cu126
- Accelerate: 1.3.0
- Datasets: 3.2.0
- Tokenizers: 0.21.0
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}ContrastiveTensionLoss
@inproceedings{carlsson2021semantic,
title={Semantic Re-tuning with Contrastive Tension},
author={Fredrik Carlsson and Amaru Cuba Gyllensten and Evangelia Gogoulou and Erik Ylip{"a}{"a} Hellqvist and Magnus Sahlgren},
booktitle={International Conference on Learning Representations},
year={2021},
url={https://openreview.net/forum?id=Ov_sMNau-PF}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
