minsuas/Misconceptions_1
SentenceTransformer
This is a sentence-transformers model trained. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer <!-- - Base model: Unknown -->
- Maximum Sequence Length: 256 tokens
- Output Dimensionality: 384 dimensions
- Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 256, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("minsuas/Misconceptions_1")
# Run inference
sentences = [
'Subject: Construct Triangle\nConstruct: Construct a triangle using Side-Side-Side\nQuestion: Tom and Katie are arguing about constructing triangles.\n\nTom says you can construct a triangle with lengths \\( 11 \\mathrm{~cm}, 10 \\mathrm{~cm} \\) and \\( 2 \\mathrm{~cm} \\).\n\nKatie says you can construct a triangle with lengths \\( 8 \\mathrm{~cm}, 5 \\mathrm{~cm} \\) and \\( 3 \\mathrm{~cm} \\).\n\nWho is correct?\nIncorrect Answer: Neither is correct',
'Does not realise that the sum of the two shorter sides must be greater than the third side for it to be a possible triangle\nThe passage is discussing a common misconception about the properties required to form a triangle. The misconception is that one might think any three given side lengths can form a triangle. However, for three lengths to actually form a triangle, they must satisfy the triangle inequality theorem. This theorem states that the sum of the lengths of any two sides of a triangle must be greater than the length of the remaining side. This rule must hold true for all three combinations of added side lengths. \n\nTo apply this to the misconception: one does not realize that the sum of the lengths of the two shorter sides must be greater than the length of the longest side to form a possible triangle. This ensures that the sides can actually meet to form a closed figure with three angles.',
'Draws both angles at the same end of the line when constructing a triangle\nThe misconception described refers to a common error in geometry when students are constructing a triangle based on given angles and a line segment. The mistake is to draw both given angles at the same end of the given line segment. This is incorrect because in a triangle, each angle is located at a different vertex, and each vertex connects two sides. To correctly construct the triangle, each given angle should be drawn at different ends of the line segment (if constructing based on one line segment and two angles) or at vertices defined by the construction steps (if additional sides are given). This ensures that the three angles are positioned to form the corners of the triangle, with each angle at a distinct vertex, thereby creating a proper triangle.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 17,405 training samples
- Columns: <code>anchor</code>, <code>positive</code>, and <code>negative</code>
- Approximate statistics based on the first 1000 samples: | | anchor | positive | negative | |:--------|:------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------| | type | string | string | string | | details | <ul><li>min: 32 tokens</li><li>mean: 87.49 tokens</li><li>max: 256 tokens</li></ul> | <ul><li>min: 79 tokens</li><li>mean: 179.08 tokens</li><li>max: 256 tokens</li></ul> | <ul><li>min: 75 tokens</li><li>mean: 180.1 tokens</li><li>max: 256 tokens</li></ul> |
- Samples: | anchor | positive | negative | |:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Subject: Cubics and Reciprocals<br>Construct: Given a positive x value, find the corresponding y value for reciprocal graphs<br>Question: This is a part of the table of values for the equation \( y=\frac{3}{x} \) \begin{tabular}{|l|l|}<br>\hline\( x \) & \( 0.1 \) \\<br>\hline\( y \) & \(\color{gold}\bigstar\) \\<br>\hline<br>\end{tabular} What should replace the star?<br>Incorrect Answer: \( 0.3 \)</code> | <code>Multiplies instead of divides when division is written as a fraction<br>The passage is highlighting a common mistake students sometimes make when dealing with fractions in mathematics. This misconception occurs when a student encounters a fraction (which inherently involves a division operation, i.e., the numerator divided by the denominator) but instead of performing division, the student multiplies the numerator by the denominator. This misunderstanding can lead to incorrect solutions in problems where the correct interpretation and handling of fractions are crucial. For example, if presented with the fraction 8/2, the correct operation is to divide 8 by 2, resulting in 4, not to multiply 8 by 2, which would incorrectly yield 16.</code> | <code>Forgets that a number divided by itself is 1<br>The passage is highlighting a common mistake made in mathematics where a student forgets the fundamental fact that any non-zero number divided by itself equals 1. For example, 5 divided by 5 is 1, or more generally, for any non-zero number n, n/n = 1. This principle is crucial for simplifying fractions, solving equations, and understanding basic arithmetic properties. Forgetting this can lead to errors in calculations and problem-solving scenarios.</code> | | <code>Subject: Angle Facts with Parallel Lines<br>Construct: Identify a transversal<br>Question: What is the name given to the red line that intersects the two dashed lines? ![Shows two straight horizontal dashed lines that are converging and are both intersected by a solid red line]()<br>Incorrect Answer: Parallel</code> | <code>Does not know the meaning of the word parallel<br>The passage is indicating a misconception related to a math problem, specifically one that involves the concept of "parallel." In mathematics, particularly in geometry, "parallel" refers to lines or planes that are equidistant from each other at every point and never intersect, no matter how far they are extended. A misunderstanding or lack of knowledge about this definition can lead to errors when solving problems that involve parallel lines or planes, such as determining angles or distances. Thus, to correctly interpret and solve problems involving parallel lines or planes, one must understand that they maintain a constant distance from each other and never meet.</code> | <code>Does not know the term transversal<br>The passage is indicating a common pitfall in geometry where a student may not be familiar with the term "transversal." A transversal is a line that passes through two or more other lines, often creating several angles with them. When discussing parallel lines and the angles formed when a transversal intersects them, understanding the term and its implications is crucial. The misconception here is likely that without knowing what a transversal is, a student might struggle to identify the relationships between the angles formed (such as alternate interior angles, corresponding angles, etc.), which are fundamental concepts in solving geometry problems involving parallel lines.</code> | | <code>Subject: Sharing in a Ratio<br>Construct: Divide a quantity into two parts for a given a ratio, where each part is an integer<br>Question: Share \( £360 \) in the ratio \( 2: 7 \)<br>Incorrect Answer: \( £ 180: £ 51 \)</code> | <code>Divides total amount by each side of the ratio instead of dividing by the sum of the parts<br>The misconception described refers to a mistake made when dividing a total amount according to a given ratio. For instance, if someone has to divide $100 in the ratio 2:3, a correct approach would be to first add the parts of the ratio (2+3=5) to find the total number of parts. Then, divide the total amount by this sum ($100 ÷ 5 = $20) to determine the value of one part. This $20 can then be multiplied by each number in the ratio (2 and 3) to correctly distribute the $100.<br><br>The misconception occurs when someone divides the total amount ($100) by each individual number in the ratio (2 and 3) rather than by the sum of the parts (5). This method would incorrectly distribute the $100, as it does not account for the proportional relationship that the ratio is meant to establish.</code> | <code>Estimates shares of a ratio instead of calculating<br>The passage is discussing a common mistake made in mathematics, particularly when dealing with ratio problems. The misconception lies in estimating the shares or parts of a ratio rather than calculating them accurately. For example, if a problem involves dividing a quantity in the ratio of 2:3, the misconception would be to guess or estimate what parts of the quantity correspond to 2 and 3, instead of using the correct method to find the exact shares. The correct approach involves first adding the parts of the ratio (in this case, 2 + 3 = 5) and then using this sum to calculate each part's exact share of the total quantity. Thus, it's important to calculate each part of the ratio precisely rather than estimating.</code> |
- Loss: <code>CachedMultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim"
}Training Hyperparameters
Non-Default Hyperparameters
per_device_train_batch_size: 512num_train_epochs: 1lr_scheduler_type: cosinewarmup_ratio: 0.1fp16: True
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: noprediction_loss_only: Trueper_device_train_batch_size: 512per_device_eval_batch_size: 8per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 1max_steps: -1lr_scheduler_type: cosinelr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Truefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: proportional
</details>
Framework Versions
- Python: 3.10.12
- Sentence Transformers: 3.3.1
- Transformers: 4.47.1
- PyTorch: 2.5.1+cu121
- Accelerate: 1.2.1
- Datasets: 3.2.0
- Tokenizers: 0.21.0
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}CachedMultipleNegativesRankingLoss
@misc{gao2021scaling,
title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
year={2021},
eprint={2101.06983},
archivePrefix={arXiv},
primaryClass={cs.LG}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
