LamaDiab/MiniLM-V10Data-256BATCH-SemanticEngine
082
1---2tags:3- sentence-transformers4- sentence-similarity5- feature-extraction6- dense7- generated_from_trainer8- dataset_size:4588309- loss:MultipleNegativesSymmetricRankingLoss10base_model: sentence-transformers/all-MiniLM-L6-v211widget:12- source_sentence: derby cap toe shoes - brown13 sentences:14 - chained strapped block heeled sandals15 - 100% premium natural leather - high quality sole.16 - puppy treats biscuits17- source_sentence: juliette bundle18 sentences:19 - juliette body lotion20 - xiaomi 12 pro21 - reece - serving plate set22- source_sentence: granville original one bite original rice crispy squares23 sentences:24 - ' samsung galaxy z flip case'25 - rice crispy squares dairy-free26 - nivea - rose care micellar water with organic rose water & oil - 400 ml27- source_sentence: rosa fm farha istikana tea cup with plate set28 sentences:29 - premium rug30 - rosa fm farha31 - wall anchor bolt 1444 euro32- source_sentence: jade life necklace33 sentences:34 - duplicate faces35 - protection necklace36 - carefree daily intimate wash37pipeline_tag: sentence-similarity38library_name: sentence-transformers39metrics:40- cosine_accuracy41model-index:42- name: SentenceTransformer based on sentence-transformers/all-MiniLM-L6-v243 results:44 - task:45 type: triplet46 name: Triplet47 dataset:48 name: Unknown49 type: unknown50 metrics:51 - type: cosine_accuracy52 value: 0.959827542304992753 name: Cosine Accuracy54 - type: cosine_accuracy55 value: 0.931085765361785956 name: Cosine Accuracy57---58 59# SentenceTransformer based on sentence-transformers/all-MiniLM-L6-v260 61This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [sentence-transformers/all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2). It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.62 63## Model Details64 65### Model Description66- **Model Type:** Sentence Transformer67- **Base model:** [sentence-transformers/all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) <!-- at revision c9745ed1d9f207416be6d2e6f8de32d1f16199bf -->68- **Maximum Sequence Length:** 256 tokens69- **Output Dimensionality:** 384 dimensions70- **Similarity Function:** Cosine Similarity71<!-- - **Training Dataset:** Unknown -->72<!-- - **Language:** Unknown -->73<!-- - **License:** Unknown -->74 75### Model Sources76 77- **Documentation:** [Sentence Transformers Documentation](https://sbert.net)78- **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers)79- **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers)80 81### Full Model Architecture82 83```84SentenceTransformer(85 (0): Transformer({'max_seq_length': 256, 'do_lower_case': False, 'architecture': 'BertModel'})86 (1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})87 (2): Normalize()88)89```90 91## Usage92 93### Direct Usage (Sentence Transformers)94 95First install the Sentence Transformers library:96 97```bash98pip install -U sentence-transformers99```100 101Then you can load this model and run inference.102```python103from sentence_transformers import SentenceTransformer104 105# Download from the ๐ค Hub106model = SentenceTransformer("LamaDiab/MiniLM-V10Data-256BATCH-SemanticEngine")107# Run inference108sentences = [109 'jade life necklace',110 'protection necklace',111 'duplicate faces',112]113embeddings = model.encode(sentences)114print(embeddings.shape)115# [3, 384]116 117# Get the similarity scores for the embeddings118similarities = model.similarity(embeddings, embeddings)119print(similarities)120# tensor([[1.0000, 0.7006, 0.1425],121# [0.7006, 1.0000, 0.1684],122# [0.1425, 0.1684, 1.0000]])123```124 125<!--126### Direct Usage (Transformers)127 128<details><summary>Click to see the direct usage in Transformers</summary>129 130</details>131-->132 133<!--134### Downstream Usage (Sentence Transformers)135 136You can finetune this model on your own dataset.137 138<details><summary>Click to expand</summary>139 140</details>141-->142 143<!--144### Out-of-Scope Use145 146*List how the model may foreseeably be misused and address what users ought not to do with the model.*147-->148 149## Evaluation150 151### Metrics152 153#### Triplet154 155* Evaluated with [<code>TripletEvaluator</code>](https://sbert.net/docs/package_reference/sentence_transformer/evaluation.html#sentence_transformers.evaluation.TripletEvaluator)156 157| Metric | Value |158|:--------------------|:-----------|159| **cosine_accuracy** | **0.9598** |160 161#### Triplet162 163* Evaluated with [<code>TripletEvaluator</code>](https://sbert.net/docs/package_reference/sentence_transformer/evaluation.html#sentence_transformers.evaluation.TripletEvaluator)164 165| Metric | Value |166|:--------------------|:-----------|167| **cosine_accuracy** | **0.9311** |168 169<!--170## Bias, Risks and Limitations171 172*What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*173-->174 175<!--176### Recommendations177 178*What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*179-->180 181## Training Details182 183### Training Dataset184 185#### Unnamed Dataset186 187* Size: 458,830 training samples188* Columns: <code>anchor</code> and <code>positive</code>189* Approximate statistics based on the first 1000 samples:190 | | anchor | positive |191 |:--------|:----------------------------------------------------------------------------------|:---------------------------------------------------------------------------------|192 | type | string | string |193 | details | <ul><li>min: 3 tokens</li><li>mean: 8.01 tokens</li><li>max: 119 tokens</li></ul> | <ul><li>min: 3 tokens</li><li>mean: 6.27 tokens</li><li>max: 66 tokens</li></ul> |194* Samples:195 | anchor | positive |196 |:-------------------------------------------------------------------------------|:------------------------------|197 | <code>with mountain honey and lemon zest the taste of french childhood.</code> | <code>honey madeleines</code> |198 | <code>acetone</code> | <code>peach acetone</code> |199 | <code>yellow hair oil</code> | <code>argan hair oil</code> |200* Loss: [<code>MultipleNegativesSymmetricRankingLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#multiplenegativessymmetricrankingloss) with these parameters:201 ```json202 {203 "scale": 20.0,204 "similarity_fct": "cos_sim",205 "gather_across_devices": false206 }207 ```208 209### Evaluation Dataset210 211#### Unnamed Dataset212 213* Size: 9,509 evaluation samples214* Columns: <code>anchor</code>, <code>positive</code>, and <code>negative</code>215* Approximate statistics based on the first 1000 samples:216 | | anchor | positive | negative |217 |:--------|:---------------------------------------------------------------------------------|:---------------------------------------------------------------------------------|:---------------------------------------------------------------------------------|218 | type | string | string | string |219 | details | <ul><li>min: 3 tokens</li><li>mean: 9.63 tokens</li><li>max: 43 tokens</li></ul> | <ul><li>min: 2 tokens</li><li>mean: 6.4 tokens</li><li>max: 150 tokens</li></ul> | <ul><li>min: 3 tokens</li><li>mean: 9.37 tokens</li><li>max: 41 tokens</li></ul> |220* Samples:221 | anchor | positive | negative |222 |:---------------------------------------------------------------------|:------------------------------|:------------------------------------------|223 | <code>pilot mechanical pencil progrex h-127 - 0.7 mm</code> | <code>office supplies</code> | <code>hellmann's garlic mayonnaise</code> |224 | <code>superior drawing marker -pen - set of 12 colors - 2 nib</code> | <code> marker pen set </code> | <code>indian sambousak</code> |225 | <code>first person singular author: haruki murakami</code> | <code>english book</code> | <code>elle scarf</code> |226* Loss: [<code>MultipleNegativesSymmetricRankingLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#multiplenegativessymmetricrankingloss) with these parameters:227 ```json228 {229 "scale": 20.0,230 "similarity_fct": "cos_sim",231 "gather_across_devices": false232 }233 ```234 235### Training Hyperparameters236#### Non-Default Hyperparameters237 238- `eval_strategy`: steps239- `per_device_train_batch_size`: 256240- `per_device_eval_batch_size`: 256241- `learning_rate`: 1e-05242- `weight_decay`: 0.001243- `num_train_epochs`: 5244- `warmup_steps`: 2867245- `fp16`: True246- `dataloader_num_workers`: 1247- `dataloader_prefetch_factor`: 2248- `dataloader_persistent_workers`: True249- `push_to_hub`: True250- `hub_model_id`: LamaDiab/MiniLM-V10Data-256BATCH-SemanticEngine251- `hub_strategy`: all_checkpoints252- `batch_sampler`: no_duplicates253 254#### All Hyperparameters255<details><summary>Click to expand</summary>256 257- `overwrite_output_dir`: False258- `do_predict`: False259- `eval_strategy`: steps260- `prediction_loss_only`: True261- `per_device_train_batch_size`: 256262- `per_device_eval_batch_size`: 256263- `per_gpu_train_batch_size`: None264- `per_gpu_eval_batch_size`: None265- `gradient_accumulation_steps`: 1266- `eval_accumulation_steps`: None267- `torch_empty_cache_steps`: None268- `learning_rate`: 1e-05269- `weight_decay`: 0.001270- `adam_beta1`: 0.9271- `adam_beta2`: 0.999272- `adam_epsilon`: 1e-08273- `max_grad_norm`: 1.0274- `num_train_epochs`: 5275- `max_steps`: -1276- `lr_scheduler_type`: linear277- `lr_scheduler_kwargs`: {}278- `warmup_ratio`: 0279- `warmup_steps`: 2867280- `log_level`: passive281- `log_level_replica`: warning282- `log_on_each_node`: True283- `logging_nan_inf_filter`: True284- `save_safetensors`: True285- `save_on_each_node`: False286- `save_only_model`: False287- `restore_callback_states_from_checkpoint`: False288- `no_cuda`: False289- `use_cpu`: False290- `use_mps_device`: False291- `seed`: 42292- `data_seed`: None293- `jit_mode_eval`: False294- `use_ipex`: False295- `bf16`: False296- `fp16`: True297- `fp16_opt_level`: O1298- `half_precision_backend`: auto299- `bf16_full_eval`: False300- `fp16_full_eval`: False301- `tf32`: None302- `local_rank`: 0303- `ddp_backend`: None304- `tpu_num_cores`: None305- `tpu_metrics_debug`: False306- `debug`: []307- `dataloader_drop_last`: False308- `dataloader_num_workers`: 1309- `dataloader_prefetch_factor`: 2310- `past_index`: -1311- `disable_tqdm`: False312- `remove_unused_columns`: True313- `label_names`: None314- `load_best_model_at_end`: False315- `ignore_data_skip`: False316- `fsdp`: []317- `fsdp_min_num_params`: 0318- `fsdp_config`: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}319- `fsdp_transformer_layer_cls_to_wrap`: None320- `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}321- `deepspeed`: None322- `label_smoothing_factor`: 0.0323- `optim`: adamw_torch324- `optim_args`: None325- `adafactor`: False326- `group_by_length`: False327- `length_column_name`: length328- `ddp_find_unused_parameters`: None329- `ddp_bucket_cap_mb`: None330- `ddp_broadcast_buffers`: False331- `dataloader_pin_memory`: True332- `dataloader_persistent_workers`: True333- `skip_memory_metrics`: True334- `use_legacy_prediction_loop`: False335- `push_to_hub`: True336- `resume_from_checkpoint`: None337- `hub_model_id`: LamaDiab/MiniLM-V10Data-256BATCH-SemanticEngine338- `hub_strategy`: all_checkpoints339- `hub_private_repo`: None340- `hub_always_push`: False341- `hub_revision`: None342- `gradient_checkpointing`: False343- `gradient_checkpointing_kwargs`: None344- `include_inputs_for_metrics`: False345- `include_for_metrics`: []346- `eval_do_concat_batches`: True347- `fp16_backend`: auto348- `push_to_hub_model_id`: None349- `push_to_hub_organization`: None350- `mp_parameters`: 351- `auto_find_batch_size`: False352- `full_determinism`: False353- `torchdynamo`: None354- `ray_scope`: last355- `ddp_timeout`: 1800356- `torch_compile`: False357- `torch_compile_backend`: None358- `torch_compile_mode`: None359- `include_tokens_per_second`: False360- `include_num_input_tokens_seen`: False361- `neftune_noise_alpha`: None362- `optim_target_modules`: None363- `batch_eval_metrics`: False364- `eval_on_start`: False365- `use_liger_kernel`: False366- `liger_kernel_config`: None367- `eval_use_gather_object`: False368- `average_tokens_across_devices`: False369- `prompts`: None370- `batch_sampler`: no_duplicates371- `multi_dataset_batch_sampler`: proportional372- `router_mapping`: {}373- `learning_rate_mapping`: {}374 375</details>376 377### Training Logs378| Epoch | Step | Training Loss | Validation Loss | cosine_accuracy |379|:------:|:----:|:-------------:|:---------------:|:---------------:|380| -1 | -1 | - | - | 0.9311 |381| 0.0006 | 1 | 2.0553 | - | - |382| 0.5577 | 1000 | - | 1.0153 | 0.9486 |383| 1.0 | 1793 | 1.8693 | - | - |384| 1.1154 | 2000 | - | 0.9366 | 0.9526 |385| 1.6732 | 3000 | - | 0.9148 | 0.9576 |386| 2.0 | 3586 | 1.5407 | - | - |387| 2.2309 | 4000 | - | 0.9049 | 0.9572 |388| 2.7886 | 5000 | - | 0.8995 | 0.9595 |389| 3.0 | 5379 | 1.3949 | - | - |390| 3.3463 | 6000 | - | 0.8901 | 0.9598 |391| 3.9041 | 7000 | - | 0.8891 | 0.9598 |392| 4.0 | 7172 | 1.3221 | - | - |393 394 395### Framework Versions396- Python: 3.11.13397- Sentence Transformers: 5.1.2398- Transformers: 4.53.3399- PyTorch: 2.6.0+cu124400- Accelerate: 1.9.0401- Datasets: 4.4.1402- Tokenizers: 0.21.2403 404## Citation405 406### BibTeX407 408#### Sentence Transformers409```bibtex410@inproceedings{reimers-2019-sentence-bert,411 title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",412 author = "Reimers, Nils and Gurevych, Iryna",413 booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",414 month = "11",415 year = "2019",416 publisher = "Association for Computational Linguistics",417 url = "https://arxiv.org/abs/1908.10084",418}419```420 421<!--422## Glossary423 424*Clearly define terms in order to be accessible across audiences.*425-->426 427<!--428## Model Card Authors429 430*Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*431-->432 433<!--434## Model Card Contact435 436*Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*437-->