CoolFace
Modelpublic

iamkpi/modernbert-embed-base-legal-matryoshka-2

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes18downloads
Model Card

ModernBERT Embed base Legal Matryoshka

This is a sentence-transformers model finetuned from nomic-ai/modernbert-embed-base on the json dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: nomic-ai/modernbert-embed-base <!-- at revision d556a88e332558790b210f7bdbe87da2fa94a8d8 -->
  • —Maximum Sequence Length: 8192 tokens
  • —Output Dimensionality: 768 dimensions
  • —Similarity Function: Cosine Similarity
  • —Training Dataset:
  • —json
  • —Language: en
  • —License: apache-2.0

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 8192, 'do_lower_case': False}) with Transformer model: ModernBertModel 
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("iamkpi/modernbert-embed-base-legal-matryoshka-2")
# Run inference
sentences = [
    'the dispensary, where he went after he was shot.  \nAs a witness for the State, Detective Victor Liu of the Baltimore Police Department \ntestified that, on September 3, 2021, he responded to a report of “a shooting incident in the \n3900 block of Falls Road.”  There, Detective Liu saw an SUV with bullet holes in the back',
    'What did Detective Liu see at the scene of the shooting incident?',
    'Is the Commission considered an agency under § 551(1)?',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
json
  {
      "truncate_dim": 768
  }
MetricValue
cosine_accuracy@10.5209
cosine_accuracy@30.5873
cosine_accuracy@50.7002
cosine_accuracy@100.7651
cosine_precision@10.5209
cosine_precision@30.4956
cosine_precision@50.3913
cosine_precision@100.2331
cosine_recall@10.1928
cosine_recall@30.5012
cosine_recall@50.6439
cosine_recall@100.7571
cosine_ndcg@100.6492
cosine_mrr@100.5807
cosine_map@1000.6256
Information Retrieval
json
  {
      "truncate_dim": 512
  }
MetricValue
cosine_accuracy@10.5317
cosine_accuracy@30.5765
cosine_accuracy@50.6739
cosine_accuracy@100.7558
cosine_precision@10.5317
cosine_precision@30.4992
cosine_precision@50.383
cosine_precision@100.2295
cosine_recall@10.1937
cosine_recall@30.5009
cosine_recall@50.6256
cosine_recall@100.7465
cosine_ndcg@100.6441
cosine_mrr@100.5823
cosine_map@1000.6232
Information Retrieval
json
  {
      "truncate_dim": 256
  }
MetricValue
cosine_accuracy@10.4853
cosine_accuracy@30.5363
cosine_accuracy@50.6306
cosine_accuracy@100.7156
cosine_precision@10.4853
cosine_precision@30.456
cosine_precision@50.3518
cosine_precision@100.2133
cosine_recall@10.1788
cosine_recall@30.4624
cosine_recall@50.5791
cosine_recall@100.6971
cosine_ndcg@100.5963
cosine_mrr@100.5358
cosine_map@1000.5799
Information Retrieval
json
  {
      "truncate_dim": 128
  }
MetricValue
cosine_accuracy@10.4142
cosine_accuracy@30.4683
cosine_accuracy@50.5502
cosine_accuracy@100.6476
cosine_precision@10.4142
cosine_precision@30.3962
cosine_precision@50.3116
cosine_precision@100.1944
cosine_recall@10.1495
cosine_recall@30.3967
cosine_recall@50.5081
cosine_recall@100.6319
cosine_ndcg@100.5275
cosine_mrr@100.4655
cosine_map@1000.5099
Information Retrieval
json
  {
      "truncate_dim": 64
  }
MetricValue
cosine_accuracy@10.2952
cosine_accuracy@30.3416
cosine_accuracy@50.4142
cosine_accuracy@100.493
cosine_precision@10.2952
cosine_precision@30.2849
cosine_precision@50.2291
cosine_precision@100.1485
cosine_recall@10.1081
cosine_recall@30.2885
cosine_recall@50.3776
cosine_recall@100.4825
cosine_ndcg@100.3937
cosine_mrr@100.3396
cosine_map@1000.385

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

json
  • —Dataset: json
  • —Size: 5,822 training samples
  • —Columns: <code>positive</code> and <code>anchor</code>
  • —Approximate statistics based on the first 1000 samples: | | positive | anchor | |:--------|:------------------------------------------------------------------------------------|:----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 26 tokens</li><li>mean: 96.36 tokens</li><li>max: 170 tokens</li></ul> | <ul><li>min: 8 tokens</li><li>mean: 16.47 tokens</li><li>max: 32 tokens</li></ul> |
  • —Samples: | positive | anchor | |:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------| | <code>the same time they would if they were assigned the original requests.” Id. at 11–12. <br>The plaintiff responds by focusing on the factual underpinnings of the CIA’s policy <br>arguments—in particular the CIA’s contentions about “undue burden.” See Pl.’s 443 Cross-Mot. <br>Mem. at 2–7. For example, the plaintiff points out that the CIA waives FOIA fees “‘as an act of</code> | <code>What is one argument the plaintiff critiques regarding the CIA's policy?</code> | | <code>contends that, “[i]n order to be properly withheld [under Exemption 2], the information must be <br>of a relatively trivial nature.” Id. (citing Dep’t of Air Force v. Rose, 425 U.S. 352, 369–70 <br>(1976) and Lesar v. DOJ, 636 F.2d 472, 485 (D.C. Cir. 1980)). This triviality requirement <br>applies, according to plaintiff, because the rationale for Exemption 2 is “that the very task of</code> | <code>What does the plaintiff assert as the rationale for Exemption 2?</code> | | <code>the shooting.2 The video was 1 minute and 51 seconds long. <br>Before admission of the video, Mr. Zimmerman testified that, in the months prior <br>to the shooting, he had suspected Mr. Mooney of sleeping with his girlfriend, but Mr. <br>Mooney had denied the allegation. Mr. Zimmerman testified that, on the night of the</code> | <code>What did Mr. Mooney do in response to the allegation?</code> |
  • —Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "MultipleNegativesRankingLoss",
      "matryoshka_dims": [
          768,
          512,
          256,
          128,
          64
      ],
      "matryoshka_weights": [
          1,
          1,
          1,
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: epoch
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —gradient_accumulation_steps: 16
  • —learning_rate: 2e-05
  • —num_train_epochs: 4
  • —lr_scheduler_type: cosine
  • —warmup_ratio: 0.1
  • —bf16: True
  • —tf32: False
  • —load_best_model_at_end: True
  • —optim: adamwtorchfused
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: epoch
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 16
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 2e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 4
  • —max_steps: -1
  • —lr_scheduler_type: cosine
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: True
  • —fp16: False
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: False
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: True
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —hub_revision: None
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —liger_kernel_config: None
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional

</details>

Training Logs

EpochStepTraining Lossdim_768_cosine_ndcg@10dim_512_cosine_ndcg@10dim_256_cosine_ndcg@10dim_128_cosine_ndcg@10dim_64_cosine_ndcg@10
0.8791105.4341-----
1.012-0.58940.58800.54250.45810.3261
1.7033202.535-----
2.024-0.63100.62750.58760.50390.3711
2.5275301.854-----
3.036-0.64560.64000.59520.52060.3938
3.3516401.7104-----
4.048-0.64920.64410.59630.52750.3937
  • —The bold row denotes the saved checkpoint.

Framework Versions

  • —Python: 3.11.13
  • —Sentence Transformers: 4.1.0
  • —Transformers: 4.54.0
  • —PyTorch: 2.6.0+cu124
  • —Accelerate: 1.9.0
  • —Datasets: 4.0.0
  • —Tokenizers: 0.21.2

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MatryoshkaLoss
bibtex
@misc{kusupati2024matryoshka,
    title={Matryoshka Representation Learning},
    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
    year={2024},
    eprint={2205.13147},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->