CoolFace
Modelpublic

vaios-stergio/all-mpnet-base-v2-dblp-aminer-180k-pairs

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes71downloads
Model Card

SentenceTransformer based on sentence-transformers/all-mpnet-base-v2

This is a sentence-transformers model finetuned from sentence-transformers/all-mpnet-base-v2 on the parquet dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: sentence-transformers/all-mpnet-base-v2 <!-- at revision e8c3b32edf5434bc2275fc9bab85f82640a19130 -->
  • —Maximum Sequence Length: 384 tokens
  • —Output Dimensionality: 768 dimensions
  • —Similarity Function: Cosine Similarity
  • —Training Dataset:
  • —parquet <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 384, 'do_lower_case': False, 'architecture': 'MPNetModel'})
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'Positive Semidefiniteness and Positive Definiteness of a Linear Parametric Interval Matrix.We consider a symmetric matrix the entries of which depend linearly on some parameters',
    '. The domains of the parameters are compact real intervals. We investigate the problem of checking whether for each or some setting of the parameters the matrix is positive definite or positive semidefinite. We state a characterization in the form of equivalent conditions and also propose some computationally cheap sufficient necessary conditions. Our results extend the classical results on positive semidefiniteness of interval matrices. They may be useful for checking convexity or nonconvexity in global optimization methods based on branch and bound framework and using interval techniques.',
    '. We show that such reductions exhibit the same desirable properties as their qualitative counterparts and additionally retain the optimality of solutions. Moreover we introduce vertexranked games as a generalpurpose target for quantitative reductions and show how to solve them. In such games the value of a play is determined only by a qualitative winning condition and a ranking of the vertices. 88provide quantitative reductions of quantitative requestresponse games to vertexranked games thus showing ExpTimecompleteness of solving the former games. Furthermore we exhibit the usefulness and flexibility of vertexranked games by showing how to use such games to compute faultresilient strategies for safety specifications. This work lays the foundation for a general study of faultresilient strategies for more complex winning conditions',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.7926, 0.0113],
#         [0.7926, 1.0000, 0.1115],
#         [0.0113, 0.1115, 1.0000]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

parquet
  • —Dataset: parquet
  • —Size: 146,640 training samples
  • —Columns: <code>anchor</code> and <code>positive</code>
  • —Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 16 tokens</li><li>mean: 43.11 tokens</li><li>max: 121 tokens</li></ul> | <ul><li>min: 64 tokens</li><li>mean: 199.08 tokens</li><li>max: 384 tokens</li></ul> |
  • —Samples: | anchor | positive | |:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>The longterm effect of media violence exposure on aggression of youngsters.Abstract The effect of media violence on aggression has always been a trending issue and a better understanding of the psychological mechanism of the impact of media violence on youth aggression is an extremely important research topic for preventing the negative impacts of media violence and juvenile delinquency</code> | <code>. From the perspective of anger this study explored the longterm effect of different degrees of media violence exposure on the aggression of youngsters as well as the role of aggressive emotions. The studies found that individuals with a high degree of media violence exposure HMVE exhibited higher levels of proactive aggression in both irritation situations and higher levels of reactive aggression in lowirritation situations than did participants with a low degree of media violence exposure LMVE. After being provoked the anger of all participants was significantly increased and the anger and proactive aggression levels of the HMVE group were significantly higher than those of the LMVE group. Additionally rumination and anger played a mediating role in the relationship between media violence exposure and aggression. Overall this study enriches the theoretical understanding of the longterm effect of media violence exposure on individual aggression. Second this study deepens our understan...</code> | | <code>Humanoid coworkers How is it like to work with a robot.Humanrobot interaction in corporate workplaces is a research area which remains unexplored</code> | <code>. In this paper we present the results and analysis of a social experiment we conducted by introducing a humanoid robot Nadine into a collaborative social workplace. The humanoidu0027s primary task was to function as a receptionist and provide general assistance to the customers. Moreover the employees who interacted with Nadine were given over a month to get used to her capabilities after which the feedback was collected from the staff on the grounds of influence on productivity affect experienced during interaction and their views on social robots assisting with regular tasks. Our results show that the usage of social robots for assisting with normal daytoday tasks is taken quite positively by the coworkers and that in the near future more capable humanoid social robots can be used in workplaces for assisting with menial tasks. Finally we posit that surveys such as ours could result in constructive opinions based on technological awareness rather than opinions from mediadriven fears ...</code> | | <code>A Common Social Distance Scale for Robots and Humans.From keeping robots as inhome helpers to banning their presence or functions a persons willingness to engage in variably intimate interactions are signals of social distance the degree of felt understanding of and intimacy with an individual or group that characterizes presocial and social connections</code> | <code>. To date social distance has been examined through surrogate metrics not actually representing the construct e.g. selfdisclosure or physical proximity. To address this gap between operations and measurement this project details a fourstage social distance scale development project inclusive of systematic item poolgeneration candidate item ratings for laypersons thinking about social distance testing of candidate items via scalogram and initial validity analyses and final testing for cumulative structure and predictive validity. The final metric yields a 15item 18 counting applications with a none option threedimension scale for physical distance relational distance and conversational distance.</code> |
  • —Loss: <code>CachedMultipleNegativesRankingLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "cos_sim",
      "mini_batch_size": 16,
      "gather_across_devices": false
  }

Evaluation Dataset

parquet
  • —Dataset: parquet
  • —Size: 18,330 evaluation samples
  • —Columns: <code>anchor</code> and <code>positive</code>
  • —Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:-----------------------------------------------------------------------------------|:------------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 14 tokens</li><li>mean: 42.9 tokens</li><li>max: 123 tokens</li></ul> | <ul><li>min: 37 tokens</li><li>mean: 198.6 tokens</li><li>max: 384 tokens</li></ul> |
  • —Samples: | anchor | positive | |:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>Hybrid Filtering for a Class of Quantum Systems with Classical Disturbances.Abstract A filtering problem for a class of quantum systems disturbed by a classical stochastic process is investigated in this paper</code> | <code>. The classical disturbance process which is assumed to be described by a linear stochastic differential equation is modeled by a quantum cavity model. Then the hybrid quantumclassical system is described by a combined quantum system consisting of two quantum cavity subsystems. Quantum filtering theory and a quantum extended Kalman filter method are employed to estimate the states of the combined quantum system. An estimate of the classical stochastic process is derived from the estimate of the combined quantum system.</code> | | <code>Invertibility of graph translation and support of Laplacian Fiedler vectors.The graph Laplacian operator is widely studied in spectral graph theory largely due to its importance in modern data analysis</code> | <code>. Recently the Fourier transform and other timefrequency operators have been defined on graphs using Laplacian eigenvalues and eigenvectors. We extend these results and prove that the translation operator to the iu0027th node is invertible if and only if all eigenvectors are nonzero on the iu0027th node. Because of this dependency on the support of eigenvectors we study the characteristic set of Laplacian eigenvectors. We prove that the Fiedler vector of a planar graph cannot vanish on large neighborhoods and then explicitly construct a family of nonplanar graphs that do exhibit this property.</code> | | <code>PNP as minimization of degree 4 polynomial or Grassmann number problem.While the P vs NP problem is mainly being attacked form the point of view of discrete mathematics this paper propses two reformulations into the field of abstract algebra and of continuous global optimization which advanced tools might bring new perspectives and approaches to attack this problem</code> | <code>. The first one is equivalence of satisfying the 3SAT problem with the question of reaching zero of a nonnegative degree 4 multivariate polynomial. This continuous search between boolean 0 and 1 values could be attacked using methods of global optimization suggesting exponential growth of the number of local minima what might be also a crucial issue for example for adiabatic quantum computers. The second discussed approach is using anticommuting Grassmann numbers thetai making A cdot textrmdiagthetain nonzero only if A has a Hamilton cycle. Hence the PneNP assumption implies exponential growth of matrix representation of Grassmann numbers.</code> |
  • —Loss: <code>CachedMultipleNegativesRankingLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "cos_sim",
      "mini_batch_size": 16,
      "gather_across_devices": false
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: steps
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —num_train_epochs: 1
  • —warmup_ratio: 0.1
  • —fp16: True
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: steps
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 1
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 5e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 1
  • —max_steps: -1
  • —lr_scheduler_type: linear
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: False
  • —fp16: True
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: False
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —parallelism_config: None
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —hub_revision: None
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —liger_kernel_config: None
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional
  • —router_mapping: {}
  • —learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining LossValidation Loss
0.02181000.02810.0073
0.04362000.01790.0062
0.06553000.01460.0069
0.08734000.0220.0072
0.10915000.02860.0087
0.13096000.02020.0090
0.15277000.02610.0095
0.17468000.02230.0080
0.19649000.0180.0085
0.218210000.02470.0073
0.240011000.01730.0079
0.261812000.02080.0064
0.283713000.01760.0063
0.305514000.020.0070
0.327315000.01590.0065
0.349116000.01520.0064
0.370917000.01450.0074
0.392818000.01810.0060
0.414619000.01790.0061
0.436420000.01510.0063
0.458221000.01920.0061
0.480022000.0140.0058
0.501923000.0150.0060
0.523724000.01640.0058
0.545525000.01330.0056
0.567326000.01980.0056
0.589127000.01270.0053
0.611028000.01490.0052
0.632829000.010.0053
0.654630000.01810.0051
0.676431000.01160.0049
0.698232000.01150.0048
0.720133000.01210.0050
0.741934000.00990.0048
0.763735000.00960.0045
0.785536000.00960.0043
0.807337000.0110.0042
0.829238000.01410.0042
0.851039000.00910.0041
0.872840000.010.0040
0.894641000.00880.0040
0.916442000.00650.0040
0.938343000.01140.0039
0.960144000.00630.0038
0.981945000.00880.0038

Framework Versions

  • —Python: 3.11.4
  • —Sentence Transformers: 5.1.1
  • —Transformers: 4.56.2
  • —PyTorch: 2.8.0+cu128
  • —Accelerate: 1.10.1
  • —Datasets: 4.1.1
  • —Tokenizers: 0.22.1

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
CachedMultipleNegativesRankingLoss
bibtex
@misc{gao2021scaling,
    title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
    author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
    year={2021},
    eprint={2101.06983},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->