amentaga-nttd/example
019
1---2tags:3- sentence-transformers4- sentence-similarity5- feature-extraction6- generated_from_trainer7- dataset_size:108- loss:MatryoshkaLoss9- loss:MultipleNegativesRankingLoss10base_model: nreimers/albert-small-v211widget:12- source_sentence: What processes are used to separate the raw liquid mix from natural13 gas in a gas recycling plant?14 sentences:15 - '2016 ►M43 (*1) ◄ 21 September 2017 ►M43 (*2) ◄ [▼M28](./../../../legal-content/EN/AUTO/?uri=celex:32014R089516 "32014R0895: INSERTED") 23. Formaldehyde, oligomeric reaction products with aniline17 (technical MDA) EC No: 500-036-1 CAS No: 25214-70-4 Carcinogenic (category 1B)18 22 February 2016 ►M43 (*1) ◄ 22 August 2017 ►M43 (*2) ◄ — 24. Arsenic acid EC19 No: 231-901-9 CAS No: 7778-39-4 Carcinogenic (category 1A) 22 February 2016 2220 August 2017 — 25. Bis(2-methoxyethyl) ether (diglyme) EC No: 203-924-4 CAS No:21 111-96-6 Toxic for reproduction (category 1B) 22 February 2016 ►M43 (*1) ◄ 2222 August 2017 ►M43 (*2) ◄ — 26. 1,2-dichloroethane (EDC) EC No: 203-458-1 CAS No:23 107-06-2 Carcinogenic (category 1B) 22 May 2016 22 November 2017 — 27.'24 - '1. Member States shall ensure that their competent authorities establish at least25 one AI regulatory sandbox at national level, which shall be operational by 2 August26 2026. That sandbox may also be established jointly with the competent authorities27 of other Member States. The Commission may provide technical support, advice and28 tools for the establishment and operation of AI regulatory sandboxes.29 30 31 The obligation under the first subparagraph may also be fulfilled by participating32 in an existing sandbox in so far as that participation provides an equivalent33 level of national coverage for the participating Member States.'34 - and that boils in a range of approximately 149 °C to 205 °C.) 649-345-00-4 232-489-335 8052-41-3 P Natural gas condensates (petroleum); Low boiling point naphtha — unspecified36 (A complex combination of hydrocarbons separated as a liquid from natural gas37 in a surface separator by retrograde condensation. It consists mainly of hydrocarbons38 having carbon numbers predominantly in the range of C2 to C20. It is a liquid39 at atmospheric temperature and pressure.) 649-346-00-X 265-047-3 64741-47-5 P40 Natural gas (petroleum), raw liquid mix; Low boiling point naphtha — unspecified41 (A complex combination of hydrocarbons separated as a liquid from natural gas42 in a gas recycling plant by processes such as refrigeration or absorption. It43 consists mainly of44- source_sentence: What should the report on income tax information include as per45 Article 48c?46 sentences:47 - '(d)48 49 50 seal any business premises and books or records for the period of time of, and51 to the extent necessary for, the inspection.52 53 54 3.55 56 57 The undertaking or association of undertakings shall submit to inspections ordered58 by decision of the Commission. The officials and other accompanying persons authorised59 by the Commission to conduct an inspection shall exercise their powers upon production60 of a Commission decision:61 62 63 (a)64 65 66 specifying the subject matter and purpose of the inspection;67 68 69 (b)70 71 72 containing a statement that, pursuant to Article 16, a lack of cooperation allows73 the Commission to take a decision on the basis of the facts that are available74 to it;75 76 77 (c)'78 - 'By way of derogation from Article 10c, the Member States concerned may only give79 transitional free allocation to installations in accordance with that Article80 for investments carried out until 31 December 2024. Any allowances available to81 the Member States concerned in accordance with Article 10c for the period from82 2021 to 2030 that are not used for such investments shall, in the proportion determined83 by the respective Member State:84 85 86 (a)87 88 89 be added to the total quantity of allowances that the Member State concerned is90 to auction pursuant to Article 10(2); or91 92 93 (b)'94 - '7.95 96 97 Member States shall require subsidiary undertakings or branches not subject to98 the provisions of paragraphs 4 and 5 of this Article to publish and make accessible99 a report on income tax information where such subsidiary undertakings or branches100 serve no other objective than to circumvent the reporting requirements set out101 in this Chapter.102 103 104 Article 48c105 106 107 Content of the report on income tax information108 109 110 1.111 112 113 The report on income tax information required under Article 48b shall include114 information relating to all the activities of the standalone undertaking or ultimate115 parent undertaking, including those of all affiliated undertakings consolidated116 in the financial statements in respect of the relevant financial year.117 118 119 2.'120pipeline_tag: sentence-similarity121library_name: sentence-transformers122metrics:123- cosine_accuracy@1124- cosine_accuracy@3125- cosine_accuracy@5126- cosine_accuracy@10127- cosine_precision@1128- cosine_precision@3129- cosine_precision@5130- cosine_precision@10131- cosine_recall@1132- cosine_recall@3133- cosine_recall@5134- cosine_recall@10135- cosine_ndcg@10136- cosine_mrr@10137- cosine_map@100138model-index:139- name: SentenceTransformer based on nreimers/albert-small-v2140 results:141 - task:142 type: information-retrieval143 name: Information Retrieval144 dataset:145 name: Unknown146 type: unknown147 metrics:148 - type: cosine_accuracy@1149 value: 0.7150 name: Cosine Accuracy@1151 - type: cosine_accuracy@3152 value: 0.7153 name: Cosine Accuracy@3154 - type: cosine_accuracy@5155 value: 0.8156 name: Cosine Accuracy@5157 - type: cosine_accuracy@10158 value: 0.9159 name: Cosine Accuracy@10160 - type: cosine_precision@1161 value: 0.7162 name: Cosine Precision@1163 - type: cosine_precision@3164 value: 0.23333333333333334165 name: Cosine Precision@3166 - type: cosine_precision@5167 value: 0.16168 name: Cosine Precision@5169 - type: cosine_precision@10170 value: 0.09171 name: Cosine Precision@10172 - type: cosine_recall@1173 value: 0.7174 name: Cosine Recall@1175 - type: cosine_recall@3176 value: 0.7177 name: Cosine Recall@3178 - type: cosine_recall@5179 value: 0.8180 name: Cosine Recall@5181 - type: cosine_recall@10182 value: 0.9183 name: Cosine Recall@10184 - type: cosine_ndcg@10185 value: 0.7675917633552429186 name: Cosine Ndcg@10187 - type: cosine_mrr@10188 value: 0.7300000000000001189 name: Cosine Mrr@10190 - type: cosine_map@100191 value: 0.739090909090909192 name: Cosine Map@100193---194 195# SentenceTransformer based on nreimers/albert-small-v2196 197This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [nreimers/albert-small-v2](https://huggingface.co/nreimers/albert-small-v2). It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.198 199## Model Details200 201### Model Description202- **Model Type:** Sentence Transformer203- **Base model:** [nreimers/albert-small-v2](https://huggingface.co/nreimers/albert-small-v2) <!-- at revision 18045fa83de53fd7d4548fdc2473862914cbc7d5 -->204- **Maximum Sequence Length:** 512 tokens205- **Output Dimensionality:** 768 dimensions206- **Similarity Function:** Cosine Similarity207<!-- - **Training Dataset:** Unknown -->208<!-- - **Language:** Unknown -->209<!-- - **License:** Unknown -->210 211### Model Sources212 213- **Documentation:** [Sentence Transformers Documentation](https://sbert.net)214- **Repository:** [Sentence Transformers on GitHub](https://github.com/UKPLab/sentence-transformers)215- **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers)216 217### Full Model Architecture218 219```220SentenceTransformer(221 (0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: AlbertModel 222 (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})223)224```225 226## Usage227 228### Direct Usage (Sentence Transformers)229 230First install the Sentence Transformers library:231 232```bash233pip install -U sentence-transformers234```235 236Then you can load this model and run inference.237```python238from sentence_transformers import SentenceTransformer239 240# Download from the 🤗 Hub241model = SentenceTransformer("sentence_transformers_model_id")242# Run inference243sentences = [244 'What should the report on income tax information include as per Article 48c?',245 '7.\n\nMember States shall require subsidiary undertakings or branches not subject to the provisions of paragraphs 4 and 5 of this Article to publish and make accessible a report on income tax information where such subsidiary undertakings or branches serve no other objective than to circumvent the reporting requirements set out in this Chapter.\n\nArticle 48c\n\nContent of the report on income tax information\n\n1.\n\nThe report on income tax information required under Article 48b shall include information relating to all the activities of the standalone undertaking or ultimate parent undertaking, including those of all affiliated undertakings consolidated in the financial statements in respect of the relevant financial year.\n\n2.',246 '(d)\n\nseal any business premises and books or records for the period of time of, and to the extent necessary for, the inspection.\n\n3.\n\nThe undertaking or association of undertakings shall submit to inspections ordered by decision of the Commission. The officials and other accompanying persons authorised by the Commission to conduct an inspection shall exercise their powers upon production of a Commission decision:\n\n(a)\n\nspecifying the subject matter and purpose of the inspection;\n\n(b)\n\ncontaining a statement that, pursuant to Article 16, a lack of cooperation allows the Commission to take a decision on the basis of the facts that are available to it;\n\n(c)',247]248embeddings = model.encode(sentences)249print(embeddings.shape)250# [3, 768]251 252# Get the similarity scores for the embeddings253similarities = model.similarity(embeddings, embeddings)254print(similarities.shape)255# [3, 3]256```257 258<!--259### Direct Usage (Transformers)260 261<details><summary>Click to see the direct usage in Transformers</summary>262 263</details>264-->265 266<!--267### Downstream Usage (Sentence Transformers)268 269You can finetune this model on your own dataset.270 271<details><summary>Click to expand</summary>272 273</details>274-->275 276<!--277### Out-of-Scope Use278 279*List how the model may foreseeably be misused and address what users ought not to do with the model.*280-->281 282## Evaluation283 284### Metrics285 286#### Information Retrieval287 288* Evaluated with [<code>InformationRetrievalEvaluator</code>](https://sbert.net/docs/package_reference/sentence_transformer/evaluation.html#sentence_transformers.evaluation.InformationRetrievalEvaluator)289 290| Metric | Value |291|:--------------------|:-----------|292| cosine_accuracy@1 | 0.7 |293| cosine_accuracy@3 | 0.7 |294| cosine_accuracy@5 | 0.8 |295| cosine_accuracy@10 | 0.9 |296| cosine_precision@1 | 0.7 |297| cosine_precision@3 | 0.2333 |298| cosine_precision@5 | 0.16 |299| cosine_precision@10 | 0.09 |300| cosine_recall@1 | 0.7 |301| cosine_recall@3 | 0.7 |302| cosine_recall@5 | 0.8 |303| cosine_recall@10 | 0.9 |304| **cosine_ndcg@10** | **0.7676** |305| cosine_mrr@10 | 0.73 |306| cosine_map@100 | 0.7391 |307 308<!--309## Bias, Risks and Limitations310 311*What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*312-->313 314<!--315### Recommendations316 317*What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*318-->319 320## Training Details321 322### Training Dataset323 324#### Unnamed Dataset325 326* Size: 10 training samples327* Columns: <code>query_text</code> and <code>doc_text</code>328* Approximate statistics based on the first 10 samples:329 | | query_text | doc_text |330 |:--------|:----------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------|331 | type | string | string |332 | details | <ul><li>min: 17 tokens</li><li>mean: 38.6 tokens</li><li>max: 89 tokens</li></ul> | <ul><li>min: 113 tokens</li><li>mean: 237.7 tokens</li><li>max: 512 tokens</li></ul> |333* Samples:334 | query_text | doc_text |335 |:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|336 | <code>What are the requirements for Member States regarding the establishment of AI regulatory sandboxes, including the timeline for operational readiness and the possibility of joint establishment with other Member States?</code> | <code>1. Member States shall ensure that their competent authorities establish at least one AI regulatory sandbox at national level, which shall be operational by 2 August 2026. That sandbox may also be established jointly with the competent authorities of other Member States. The Commission may provide technical support, advice and tools for the establishment and operation of AI regulatory sandboxes.<br><br>The obligation under the first subparagraph may also be fulfilled by participating in an existing sandbox in so far as that participation provides an equivalent level of national coverage for the participating Member States.</code> |337 | <code>Member States must provide updates on their national energy and climate strategies, detailing the anticipated energy savings from 2021 to 2030. They are also obligated to report on the necessary energy savings and the policies intended to achieve these goals. If assessments reveal that a Member State's measures are inadequate to meet energy savings targets, the Commission may issue recommendations for improvement. Additionally, any shortfall in energy savings must be addressed in subsequent obligation periods.</code> | <code>9. Member States shall apply and calculate the effect of the options chosen under paragraph 8 for the period referred to in paragraph 1, first subparagraph, points (a) and (b)(i), separately:<br><br>(a) for the calculation of the amount of energy savings required for the obligation period referred to in paragraph 1, first subparagraph, point (a), Member States may make use of the options listed in paragraph 8, points (a) to (d). All the options chosen under paragraph 8 taken together shall amount to no more than 25 % of the amount of energy savings referred to in paragraph 1, first subparagraph, point (a); (b) for the calculation of the amount of energy savings required for the obligation period referred to in paragraph 1, first subparagraph, point (b)(i), Member States may make use of the options listed in paragraph 8, points (b) to (g), provided that the individual actions referred to in paragraph 8, point (d), continue to have a verifiable and measurable impact after 31 December 2020. All...</code> |338 | <code>What is the functional definition of a remote biometric identification system, and how does it operate in terms of identifying individuals without their active participation?</code> | <code>(17) The notion of ‘remote biometric identification system’ referred to in this Regulation should be defined functionally, as an AI system intended for the identification of natural persons without their active involvement, typically at a distance, through the comparison of a person’s biometric data with the biometric data contained in a reference database, irrespectively of the particular technology, processes or types of biometric data used. Such remote biometric identification systems are typically used to perceive multiple persons or their behaviour simultaneously in order to facilitate significantly the identification of natural persons without their active involvement. This excludes AI systems intended to be used for biometric</code> |339* Loss: [<code>MatryoshkaLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#matryoshkaloss) with these parameters:340 ```json341 {342 "loss": "MultipleNegativesRankingLoss",343 "matryoshka_dims": [344 768,345 512,346 256,347 128,348 64349 ],350 "matryoshka_weights": [351 1,352 1,353 1,354 1,355 1356 ],357 "n_dims_per_step": -1358 }359 ```360 361### Training Hyperparameters362#### Non-Default Hyperparameters363 364- `eval_strategy`: steps365- `per_device_train_batch_size`: 3366- `per_device_eval_batch_size`: 3367- `learning_rate`: 2e-05368- `num_train_epochs`: 1369- `warmup_ratio`: 0.1370- `fp16`: True371- `load_best_model_at_end`: True372 373#### All Hyperparameters374<details><summary>Click to expand</summary>375 376- `overwrite_output_dir`: False377- `do_predict`: False378- `eval_strategy`: steps379- `prediction_loss_only`: True380- `per_device_train_batch_size`: 3381- `per_device_eval_batch_size`: 3382- `per_gpu_train_batch_size`: None383- `per_gpu_eval_batch_size`: None384- `gradient_accumulation_steps`: 1385- `eval_accumulation_steps`: None386- `torch_empty_cache_steps`: None387- `learning_rate`: 2e-05388- `weight_decay`: 0.0389- `adam_beta1`: 0.9390- `adam_beta2`: 0.999391- `adam_epsilon`: 1e-08392- `max_grad_norm`: 1.0393- `num_train_epochs`: 1394- `max_steps`: -1395- `lr_scheduler_type`: linear396- `lr_scheduler_kwargs`: {}397- `warmup_ratio`: 0.1398- `warmup_steps`: 0399- `log_level`: passive400- `log_level_replica`: warning401- `log_on_each_node`: True402- `logging_nan_inf_filter`: True403- `save_safetensors`: True404- `save_on_each_node`: False405- `save_only_model`: False406- `restore_callback_states_from_checkpoint`: False407- `no_cuda`: False408- `use_cpu`: False409- `use_mps_device`: False410- `seed`: 42411- `data_seed`: None412- `jit_mode_eval`: False413- `use_ipex`: False414- `bf16`: False415- `fp16`: True416- `fp16_opt_level`: O1417- `half_precision_backend`: auto418- `bf16_full_eval`: False419- `fp16_full_eval`: False420- `tf32`: None421- `local_rank`: 0422- `ddp_backend`: None423- `tpu_num_cores`: None424- `tpu_metrics_debug`: False425- `debug`: []426- `dataloader_drop_last`: False427- `dataloader_num_workers`: 0428- `dataloader_prefetch_factor`: None429- `past_index`: -1430- `disable_tqdm`: False431- `remove_unused_columns`: True432- `label_names`: None433- `load_best_model_at_end`: True434- `ignore_data_skip`: False435- `fsdp`: []436- `fsdp_min_num_params`: 0437- `fsdp_config`: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}438- `fsdp_transformer_layer_cls_to_wrap`: None439- `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}440- `deepspeed`: None441- `label_smoothing_factor`: 0.0442- `optim`: adamw_torch443- `optim_args`: None444- `adafactor`: False445- `group_by_length`: False446- `length_column_name`: length447- `ddp_find_unused_parameters`: None448- `ddp_bucket_cap_mb`: None449- `ddp_broadcast_buffers`: False450- `dataloader_pin_memory`: True451- `dataloader_persistent_workers`: False452- `skip_memory_metrics`: True453- `use_legacy_prediction_loop`: False454- `push_to_hub`: False455- `resume_from_checkpoint`: None456- `hub_model_id`: None457- `hub_strategy`: every_save458- `hub_private_repo`: None459- `hub_always_push`: False460- `gradient_checkpointing`: False461- `gradient_checkpointing_kwargs`: None462- `include_inputs_for_metrics`: False463- `include_for_metrics`: []464- `eval_do_concat_batches`: True465- `fp16_backend`: auto466- `push_to_hub_model_id`: None467- `push_to_hub_organization`: None468- `mp_parameters`: 469- `auto_find_batch_size`: False470- `full_determinism`: False471- `torchdynamo`: None472- `ray_scope`: last473- `ddp_timeout`: 1800474- `torch_compile`: False475- `torch_compile_backend`: None476- `torch_compile_mode`: None477- `dispatch_batches`: None478- `split_batches`: None479- `include_tokens_per_second`: False480- `include_num_input_tokens_seen`: False481- `neftune_noise_alpha`: None482- `optim_target_modules`: None483- `batch_eval_metrics`: False484- `eval_on_start`: False485- `use_liger_kernel`: False486- `eval_use_gather_object`: False487- `average_tokens_across_devices`: False488- `prompts`: None489- `batch_sampler`: batch_sampler490- `multi_dataset_batch_sampler`: proportional491 492</details>493 494### Training Logs495| Epoch | Step | cosine_ndcg@10 |496|:-----:|:----:|:--------------:|497| -1 | -1 | 0.7676 |498 499 500### Framework Versions501- Python: 3.11.10502- Sentence Transformers: 4.0.2503- Transformers: 4.49.0504- PyTorch: 2.6.0+cu124505- Accelerate: 0.26.0506- Datasets: 3.1.0507- Tokenizers: 0.21.2508 509## Citation510 511### BibTeX512 513#### Sentence Transformers514```bibtex515@inproceedings{reimers-2019-sentence-bert,516 title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",517 author = "Reimers, Nils and Gurevych, Iryna",518 booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",519 month = "11",520 year = "2019",521 publisher = "Association for Computational Linguistics",522 url = "https://arxiv.org/abs/1908.10084",523}524```525 526#### MatryoshkaLoss527```bibtex528@misc{kusupati2024matryoshka,529 title={Matryoshka Representation Learning},530 author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},531 year={2024},532 eprint={2205.13147},533 archivePrefix={arXiv},534 primaryClass={cs.LG}535}536```537 538#### MultipleNegativesRankingLoss539```bibtex540@misc{henderson2017efficient,541 title={Efficient Natural Language Response Suggestion for Smart Reply},542 author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},543 year={2017},544 eprint={1705.00652},545 archivePrefix={arXiv},546 primaryClass={cs.CL}547}548```549 550<!--551## Glossary552 553*Clearly define terms in order to be accessible across audiences.*554-->555 556<!--557## Model Card Authors558 559*Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*560-->561 562<!--563## Model Card Contact564 565*Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*566-->