CoolFace
Modelpublic

Bea-Taylor/objection_fine_tuned_4

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes70downloads
Model Card

SentenceTransformer based on sentence-transformers/all-MiniLM-L6-v2

This is a sentence-transformers model finetuned from sentence-transformers/all-MiniLM-L6-v2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: sentence-transformers/all-MiniLM-L6-v2 <!-- at revision c9745ed1d9f207416be6d2e6f8de32d1f16199bf -->
  • —Maximum Sequence Length: 256 tokens
  • —Output Dimensionality: 384 dimensions
  • —Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 256, 'do_lower_case': False}) with Transformer model: BertModel 
  (1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("Bea-Taylor/objection_fine_tuned_4")
# Run inference
sentences = [
    'It will also reduce/limit/block east facing views of Canary Wharf for Swedish Quays residents next door. In turn, this will have a knock on effect on the value of our property because views of Canary Wharf are sought after by potential buyers.',
    'I support the planning application for the proposed development of the roof space. This project is a vital step toward easing the financial burden on residents and addressing ongoing concerns effectively. Additionally, it brings the added benefit of a positive environmental impact, contributing to a more sustainable and responsible community.',
    'My health has really suffered over the last two years during which, I have had 3 heart attacks and have recently been diagnosed with anaemia and emphysema as well as having stents surgically placed in my arteries. I am on twelve tablets a day for my health and this whole subject is creating all my health conditions to worsen. All I can do is emphasise my objections and hope that Barnet Council decline building permission.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Semantic Similarity
MetricValue
pearson_cosine0.9829
spearman_cosine0.9159

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

Unnamed Dataset
  • —Size: 180,000 training samples
  • —Columns: <code>text1</code>, <code>text2</code>, and <code>label</code>
  • —Approximate statistics based on the first 1000 samples: | | text1 | text2 | label | |:--------|:-----------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------|:---------------------------------------------------------------| | type | string | string | float | | details | <ul><li>min: 3 tokens</li><li>mean: 50.78 tokens</li><li>max: 256 tokens</li></ul> | <ul><li>min: 3 tokens</li><li>mean: 50.72 tokens</li><li>max: 256 tokens</li></ul> | <ul><li>min: 0.0</li><li>mean: 0.36</li><li>max: 1.0</li></ul> |
  • —Samples: | text1 | text2 | label | |:------------------------------------------------------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------| | <code>Loss of Daylight and Sunlight</code> | <code>Fifthly, the increased height of the building with an additional storey will further reduce the available sunlight hours in my first floor flat leading to increased heating costs as well as a reduction in my quality of life. As the sun barely gets above the level of the existing building in the winter months it is likely I will spend much of the winter with the rear of my flat continually in shadow.</code> | <code>0.75</code> | | <code>The existing carpark has a maximum of 16 car parking spaces for the 37 flats, used on a first come first served basis.</code> | <code>As the other comments on this application state, the building works appear to be complete and the application does not appear to demonstrate the scale of the work or the reality of the build.</code> | <code>0.0</code> | | <code>Are you proposing to connect to the existing drainage system?</code> | <code>The design of the types of buildings being proposed is out of character with the area. I object to the removal of the existing footbridge... this is unacceptable and will cut the area of Victoria Park off from Cromer road - a currently safe route which pedestrians and school children use to access without having to cross road and the promise of new access or pedestrian routes will mean walking public pavements around New Barnet via Station Road. I am not convinced by the developers that an alternative safer route will be provided as they will say anything to obtain planning.</code> | <code>0.0</code> |
  • —Loss: <code>CosineSimilarityLoss</code> with these parameters:
json
  {
      "loss_fct": "torch.nn.modules.loss.MSELoss"
  }

Evaluation Dataset

Unnamed Dataset
  • —Size: 20,000 evaluation samples
  • —Columns: <code>text1</code>, <code>text2</code>, and <code>label</code>
  • —Approximate statistics based on the first 1000 samples: | | text1 | text2 | label | |:--------|:-----------------------------------------------------------------------------------|:----------------------------------------------------------------------------------|:---------------------------------------------------------------| | type | string | string | float | | details | <ul><li>min: 3 tokens</li><li>mean: 49.71 tokens</li><li>max: 256 tokens</li></ul> | <ul><li>min: 4 tokens</li><li>mean: 51.7 tokens</li><li>max: 256 tokens</li></ul> | <ul><li>min: 0.0</li><li>mean: 0.33</li><li>max: 1.0</li></ul> |
  • —Samples: | text1 | text2 | label | |:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------| | <code>1. Significant noise and disruption for local residents of rainbow quay and surrounding developments - Including over looking Princes Court and blocking already limited light.</code> | <code>I object to the developer changing the goal posts in order to achieve more profit as it is not the families who need social housing who will benefit or the young people trying to get on the property ladder and can't but the developer who profits from those who can "afford" to pay the "high prices" of Barnet accommodation.</code> | <code>0.0</code> | | <code>Congestion on Camlet way and beach hill , roads which already have traffic issues!</code> | <code>TRAFFIC AND PARKING - Granville Road has limited off road parking and is a busy & important thoroughfare from Ballards Lane, to High Road North Finchley & Summers Lane, linking Finchley & Friern Barnet. The addition of two further flats without parking provision would increase the pressure for parking spaces.</code> | <code>0.75</code> | | <code>It will also obstruct light to my property and garden.</code> | <code>- Health and safety: concern for disruption building works will cause, damage to local infrastructure, increased traffic</code> | <code>0.0</code> |
  • —Loss: <code>CosineSimilarityLoss</code> with these parameters:
json
  {
      "loss_fct": "torch.nn.modules.loss.MSELoss"
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: steps
  • —per_device_train_batch_size: 16
  • —per_device_eval_batch_size: 16
  • —learning_rate: 2e-05
  • —num_train_epochs: 1
  • —warmup_ratio: 0.1
  • —fp16: True
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: steps
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 16
  • —per_device_eval_batch_size: 16
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 1
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 2e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 1
  • —max_steps: -1
  • —lr_scheduler_type: linear
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: False
  • —fp16: True
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: False
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamw_torch
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: batch_sampler
  • —multi_dataset_batch_sampler: proportional

</details>

Training Logs

<details><summary>Click to expand</summary>

EpochStepTraining LossValidation Losssts-dev_spearman_cosine
-1-1--0.3501
0.00891000.10820.11130.3928
0.01782000.10880.10000.4880
0.02673000.09310.08570.5950
0.03564000.07650.07490.6603
0.04445000.07250.06850.6929
0.05336000.0660.06220.7216
0.06227000.05680.05580.7506
0.07118000.05250.04980.7749
0.089000.0480.04530.7926
0.088910000.04380.04120.8091
0.097811000.04470.03750.8239
0.106712000.03910.03360.8357
0.115613000.03590.03050.8486
0.124414000.03070.02720.8565
0.133315000.02890.02510.8621
0.142216000.02560.02330.8667
0.151117000.02850.02210.8702
0.1618000.02290.02060.8743
0.168919000.02280.01930.8781
0.177820000.0220.01820.8814
0.186721000.01970.01730.8827
0.195622000.01850.01670.8837
0.204423000.02020.01620.8850
0.213324000.01780.01510.8883
0.222225000.01740.01470.8896
0.231126000.01710.01460.8891
0.2427000.01630.01360.8921
0.248928000.01470.01310.8934
0.257829000.01490.01290.8953
0.266730000.01520.01220.8966
0.275631000.01380.01200.8969
0.284432000.01280.01140.8977
0.293333000.01280.01110.8991
0.302234000.01170.01060.9005
0.311135000.01260.01040.9009
0.3236000.01180.01020.9020
0.328937000.01150.01000.9016
0.337838000.01150.00980.9019
0.346739000.01160.00920.9035
0.355640000.01130.00900.9042
0.364441000.01170.00900.9043
0.373342000.00970.00840.9054
0.382243000.00980.00870.9052
0.391144000.00980.00850.9054
0.445000.00970.00840.9056
0.408946000.00970.00820.9057
0.417847000.01020.00800.9066
0.426748000.00860.00790.9071
0.435649000.00850.00780.9070
0.444450000.0090.00760.9080
0.453351000.00910.00730.9085
0.462252000.00840.00730.9085
0.471153000.00820.00710.9089
0.4854000.00730.00700.9089
0.488955000.00960.00690.9098
0.497856000.0070.00680.9097
0.506757000.00780.00700.9096
0.515658000.00790.00670.9102
0.524459000.00970.00670.9107
0.533360000.00770.00650.9110
0.542261000.00840.00650.9112
0.551162000.0070.00630.9113
0.5663000.00730.00620.9117
0.568964000.00780.00660.9107
0.577865000.00820.00620.9116
0.586766000.00660.00610.9119
0.595667000.00760.00600.9122
0.604468000.00760.00600.9120
0.613369000.00750.00590.9123
0.622270000.00710.00590.9126
0.631171000.00760.00570.9130
0.6472000.00670.00560.9131
0.648973000.00690.00570.9130
0.657874000.00680.00550.9134
0.666775000.00730.00540.9136
0.675676000.00630.00560.9131
0.684477000.00680.00540.9134
0.693378000.00570.00540.9135
0.702279000.00730.00530.9137
0.711180000.00630.00530.9139
0.7281000.00610.00520.9139
0.728982000.00620.00520.9141
0.737883000.00650.00510.9143
0.746784000.00610.00520.9141
0.755685000.00640.00500.9146
0.764486000.00560.00500.9146
0.773387000.0060.00500.9146
0.782288000.00660.00490.9147
0.791189000.0050.00480.9150
0.890000.00560.00480.9149
0.808991000.00610.00480.9149
0.817892000.00570.00470.9149
0.826793000.00750.00480.9150
0.835694000.00570.00470.9152
0.844495000.00550.00470.9151
0.853396000.00560.00470.9153
0.862297000.00490.00470.9153
0.871198000.00660.00470.9154
0.8899000.00540.00460.9154
0.8889100000.00550.00460.9154
0.8978101000.00550.00460.9155
0.9067102000.00480.00450.9155
0.9156103000.00460.00450.9156
0.9244104000.00630.00450.9157
0.9333105000.00550.00450.9157
0.9422106000.00590.00450.9158
0.9511107000.00490.00450.9158
0.96108000.00580.00450.9158
0.9689109000.00520.00450.9158
0.9778110000.00650.00440.9159
0.9867111000.00530.00440.9159
0.9956112000.00460.00440.9159
-1-1--0.9159

</details>

Framework Versions

  • —Python: 3.10.17
  • —Sentence Transformers: 4.1.0
  • —Transformers: 4.52.4
  • —PyTorch: 2.7.0
  • —Accelerate: 1.7.0
  • —Datasets: 3.6.0
  • —Tokenizers: 0.21.1

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->