CoolFace
Modelpublic

ayushexel/emb-all-MiniLM-L6-v2-squad-6-epochs

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes14downloads
Model Card

SentenceTransformer based on sentence-transformers/all-MiniLM-L6-v2

This is a sentence-transformers model finetuned from sentence-transformers/all-MiniLM-L6-v2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: sentence-transformers/all-MiniLM-L6-v2 <!-- at revision c9745ed1d9f207416be6d2e6f8de32d1f16199bf -->
  • —Maximum Sequence Length: 256 tokens
  • —Output Dimensionality: 384 dimensions
  • —Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 256, 'do_lower_case': False}) with Transformer model: BertModel 
  (1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("ayushexel/emb-all-MiniLM-L6-v2-squad-6-epochs")
# Run inference
sentences = [
    'When did people become able to vote for representation?',
    "After years of demanding greater political autonomy, residents were given the right to directly elect a Head of Government and the representatives of the unicameral Legislative Assembly by popular vote in 1997. Ever since, the left-wing Party of the Democratic Revolution (PRD) has controlled both of them. In recent years, the local government has passed a wave of liberal policies, such as abortion on request, a limited form of euthanasia, no-fault divorce, and same-sex marriage. On January 29, 2016, it ceased to be called the Federal District (Spanish: Distrito Federal or D.F.) and is now in transition to become the country's 32nd federal entity, giving it a level of autonomy comparable to that of a state. Because of a clause in the Mexican Constitution, however, as the seat of the powers of the Union, it can never become a state, lest the capital of the country be relocated elsewhere.",
    'Voting results have been a consistent source of controversy. The mechanism of voting had also aroused considerable criticisms, most notably in season two when Ruben Studdard beat Clay Aiken in a close vote, and in season eight, when the massive increase in text votes (100 million more text votes than season 7) fueled the texting controversy. Concerns about power voting have been expressed from the very first season. Since 2004, votes also have been affected to a limited degree by online communities such as DialIdol, Vote for the Worst (closed in 2013), and Vote for the Girls (started 2010).',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Triplet
MetricValue
cosine_accuracy0.4032

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

Unnamed Dataset
  • —Size: 44,286 training samples
  • —Columns: <code>question</code>, <code>context</code>, and <code>negative</code>
  • —Approximate statistics based on the first 1000 samples: | | question | context | negative | |:--------|:---------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------| | type | string | string | string | | details | <ul><li>min: 6 tokens</li><li>mean: 14.7 tokens</li><li>max: 35 tokens</li></ul> | <ul><li>min: 33 tokens</li><li>mean: 148.76 tokens</li><li>max: 256 tokens</li></ul> | <ul><li>min: 27 tokens</li><li>mean: 146.89 tokens</li><li>max: 256 tokens</li></ul> |
  • —Samples: | question | context | negative | |:------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>What type of deviations are there from the single phoneme to each grapheme general principle?</code> | <code>Although the Estonian orthography is generally guided by phonemic principles, with each grapheme corresponding to one phoneme, there are some historical and morphological deviations from this: for example preservation of the morpheme in declension of the word (writing b, g, d in places where p, k, t is pronounced) and in the use of 'i' and 'j'.[clarification needed] Where it is very impractical or impossible to type š and ž, they are substituted with sh and zh in some written texts, although this is considered incorrect. Otherwise, the h in sh represents a voiceless glottal fricative, as in Pasha (pas-ha); this also applies to some foreign names.</code> | <code>Part of the phonological study of a language therefore involves looking at data (phonetic transcriptions of the speech of native speakers) and trying to deduce what the underlying phonemes are and what the sound inventory of the language is. The presence or absence of minimal pairs, as mentioned above, is a frequently used criterion for deciding whether two sounds should be assigned to the same phoneme. However, other considerations often need to be taken into account as well.</code> | | <code>When did the Act of "Nihil novi nisi commune consensu" happen?</code> | <code>On 3 May 1505 King Alexander I Jagiellon granted the Act of "Nihil novi nisi commune consensu" (Latin: "I accept nothing new except by common consent"). This forbade the king to pass any new law without the consent of the representatives of the nobility, in Sejm and Senat assembled, and thus greatly strengthened the nobility's political position. Basically, this act transferred legislative power from the king to the Sejm. This date commonly marks the beginning of the First Rzeczpospolita, the period of a szlachta-run "Commonwealth".</code> | <code>The National Incident Based Reporting System (NIBRS) crime statistics system aims to address limitations inherent in UCR data. The system is used by law enforcement agencies in the United States for collecting and reporting data on crimes. Local, state, and federal agencies generate NIBRS data from their records management systems. Data is collected on every incident and arrest in the Group A offense category. The Group A offenses are 46 specific crimes grouped in 22 offense categories. Specific facts about these offenses are gathered and reported in the NIBRS system. In addition to the Group A offenses, eleven Group B offenses are reported with only the arrest information. The NIBRS system is in greater detail than the summary-based UCR system. As of 2004, 5,271 law enforcement agencies submitted NIBRS data. That amount represents 20% of the United States population and 16% of the crime statistics data collected by the FBI.</code> | | <code>What is the name of the idea that believes all ethnic Somalis should live in the same country?</code> | <code>Somali people in the Horn of Africa are divided among different countries (Somalia, Djibouti, Ethiopia, and northeastern Kenya) that were artificially and some might say arbitrarily partitioned by the former imperial powers. Pan-Somalism is an ideology that advocates the unification of all ethnic Somalis once part of Somali empires such as the Ajuran Empire, the Adal Sultanate, the Gobroon Dynasty and the Dervish State under one flag and one nation. The Siad Barre regime actively promoted Pan-Somalism, which eventually led to the Ogaden War between Somalia on one side, and Ethiopia, Cuba and the Soviet Union on the other.</code> | <code>The exact number of speakers of Somali is unknown. One source estimates that there are 7.78 million speakers of Somali in Somalia itself and 12.65 million speakers globally. The Somali language is spoken by ethnic Somalis in Greater Somalia and the Somali diaspora.</code> |
  • —Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "cos_sim"
  }

Evaluation Dataset

Unnamed Dataset
  • —Size: 5,000 evaluation samples
  • —Columns: <code>question</code>, <code>context</code>, and <code>negative_1</code>
  • —Approximate statistics based on the first 1000 samples: | | question | context | negative_1 | |:--------|:----------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------| | type | string | string | string | | details | <ul><li>min: 7 tokens</li><li>mean: 14.36 tokens</li><li>max: 36 tokens</li></ul> | <ul><li>min: 28 tokens</li><li>mean: 149.88 tokens</li><li>max: 256 tokens</li></ul> | <ul><li>min: 28 tokens</li><li>mean: 147.66 tokens</li><li>max: 256 tokens</li></ul> |
  • —Samples: | question | context | negative_1 | |:--------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>What century was the architect Le Corbusier in?</code> | <code>On the difference between the ideals of architecture and mere construction, the renowned 20th-century architect Le Corbusier wrote: "You employ stone, wood, and concrete, and with these materials you build houses and palaces: that is construction. Ingenuity is at work. But suddenly you touch my heart, you do me good. I am happy and I say: This is beautiful. That is Architecture".</code> | <code>Cubism was relevant to an architecture seeking a style that needed not refer to the past. Thus, what had become a revolution in both painting and sculpture was applied as part of "a profound reorientation towards a changed world". The Cubo-Futurist ideas of Filippo Tommaso Marinetti influenced attitudes in avant-garde architecture. The influential De Stijl movement embraced the aesthetic principles of Neo-plasticism developed by Piet Mondrian under the influence of Cubism in Paris. De Stijl was also linked by Gino Severini to Cubist theory through the writings of Albert Gleizes. However, the linking of basic geometric forms with inherent beauty and ease of industrial application—which had been prefigured by Marcel Duchamp from 1914—was left to the founders of Purism, Amédée Ozenfant and Charles-Édouard Jeanneret (better known as Le Corbusier,) who exhibited paintings together in Paris and published Après le cubisme in 1918. Le Corbusier's ambition had been to translate the properties o...</code> | | <code>Is the percentage of admixture in the modern Ashkenazi genome higher or lower than previously thought?</code> | <code>A 2010 study by Bray et al., using SNP microarray techniques and linkage analysis found that when assuming Druze and Palestinian Arab populations to represent the reference to world Jewry ancestor genome, between 35 to 55 percent of the modern Ashkenazi genome can possibly be of European origin, and that European "admixture is considerably higher than previous estimates by studies that used the Y chromosome" with this reference point. Assuming this reference point the linkage disequilibrium in the Ashkenazi Jewish population was interpreted as "matches signs of interbreeding or 'admixture' between Middle Eastern and European populations". On the Bray et al. tree, Ashkenazi Jews were found to be a genetically more divergent population than Russians, Orcadians, French, Basques, Italians, Sardinians and Tuscans. The study also observed that Ashkenazim are more diverse than their Middle Eastern relatives, which was counterintuitive because Ashkenazim are supposed to be a subset, not a supe...</code> | <code>A 2010 study by Bray et al., using SNP microarray techniques and linkage analysis found that when assuming Druze and Palestinian Arab populations to represent the reference to world Jewry ancestor genome, between 35 to 55 percent of the modern Ashkenazi genome can possibly be of European origin, and that European "admixture is considerably higher than previous estimates by studies that used the Y chromosome" with this reference point. Assuming this reference point the linkage disequilibrium in the Ashkenazi Jewish population was interpreted as "matches signs of interbreeding or 'admixture' between Middle Eastern and European populations". On the Bray et al. tree, Ashkenazi Jews were found to be a genetically more divergent population than Russians, Orcadians, French, Basques, Italians, Sardinians and Tuscans. The study also observed that Ashkenazim are more diverse than their Middle Eastern relatives, which was counterintuitive because Ashkenazim are supposed to be a subset, not a supe...</code> | | <code>How many doctors are there per 100,000 people?</code> | <code>The Midwest Regional Medical Center located in the suburb of Midwest City; other major hospitals in the city include the Oklahoma Heart Hospital and the Mercy Health Center. There are 347 physicians for every 100,000 people in the city.</code> | <code>Buddhism is practiced by an estimated 488 million,[web 1] 495 million, or 535 million people as of the 2010s, representing 7% to 8% of the world's total population.</code> |
  • —Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "cos_sim"
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: steps
  • —per_device_train_batch_size: 256
  • —per_device_eval_batch_size: 256
  • —num_train_epochs: 6
  • —warmup_ratio: 0.1
  • —fp16: True
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: steps
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 256
  • —per_device_eval_batch_size: 256
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 1
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 5e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 6
  • —max_steps: -1
  • —lr_scheduler_type: linear
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: False
  • —fp16: True
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: False
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —tp_size: 0
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamw_torch
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —dispatch_batches: None
  • —split_batches: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional

</details>

Training Logs

EpochStepTraining LossValidation Lossgooqa-dev_cosine_accuracy
-1-1--0.3238
0.57801000.55450.85360.3902
1.15612000.47740.82670.3982
1.73413000.40180.81290.3980
2.31214000.35940.81630.4018
2.89025000.32230.80500.4038
3.46826000.28370.80190.4050
4.04627000.27110.80470.4096
4.62438000.24140.80690.4050
5.20239000.23340.80650.3996
5.780310000.22540.80460.4048
-1-1--0.4032

Framework Versions

  • —Python: 3.11.0
  • —Sentence Transformers: 4.0.1
  • —Transformers: 4.50.3
  • —PyTorch: 2.6.0+cu124
  • —Accelerate: 1.5.2
  • —Datasets: 3.5.0
  • —Tokenizers: 0.21.1

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->