CoolFace
Modelpublic

Madhav234/modernbert-embed-base-legal-matryoshka-2

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes27downloads
Model Card

ModernBERT Embed base Legal Matryoshka

This is a sentence-transformers model finetuned from nomic-ai/modernbert-embed-base on the json dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: nomic-ai/modernbert-embed-base <!-- at revision d556a88e332558790b210f7bdbe87da2fa94a8d8 -->
  • —Maximum Sequence Length: 8192 tokens
  • —Output Dimensionality: 768 dimensions
  • —Similarity Function: Cosine Similarity
  • —Training Dataset:
  • —json
  • —Language: en
  • —License: apache-2.0

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 8192, 'do_lower_case': False, 'architecture': 'ModernBertModel'})
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("Madhav234/modernbert-embed-base-legal-matryoshka-2")
# Run inference
sentences = [
    '10  See U.S. Dep’t of Justice, 9-5.000—Issues Related to Discovery Trials, and Other \nProceedings, https://www.justice.gov/jm/jm-9-5000-issues-related-trials-and-other-court-\nproceedings (last visited May 29, 2020) (clarifying how to navigate discovery).   \n29 \n2. \n \nThe Government’s secondary argument is that the Commission falls within FACA’s',
    'What is being clarified in the provided link from the U.S. Department of Justice?',
    'What must the agency provide according to the Larson v. Dep’t of State case?',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.4024, 0.0953],
#         [0.4024, 1.0000, 0.2285],
#         [0.0953, 0.2285, 1.0000]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
json
  {
      "truncate_dim": 768
  }
MetricValue
cosine_accuracy@10.5765
cosine_accuracy@30.6244
cosine_accuracy@50.7125
cosine_accuracy@100.7743
cosine_precision@10.5765
cosine_precision@30.5466
cosine_precision@50.4117
cosine_precision@100.2388
cosine_recall@10.2098
cosine_recall@30.545
cosine_recall@50.6636
cosine_recall@100.7642
cosine_ndcg@100.6788
cosine_mrr@100.6245
cosine_map@1000.6631
Information Retrieval
json
  {
      "truncate_dim": 512
  }
MetricValue
cosine_accuracy@10.5688
cosine_accuracy@30.6121
cosine_accuracy@50.6924
cosine_accuracy@100.7743
cosine_precision@10.5688
cosine_precision@30.5389
cosine_precision@50.3997
cosine_precision@100.238
cosine_recall@10.206
cosine_recall@30.5385
cosine_recall@50.6448
cosine_recall@100.7606
cosine_ndcg@100.6707
cosine_mrr@100.6152
cosine_map@1000.6539
Information Retrieval
json
  {
      "truncate_dim": 256
  }
MetricValue
cosine_accuracy@10.5394
cosine_accuracy@30.5889
cosine_accuracy@50.6801
cosine_accuracy@100.7403
cosine_precision@10.5394
cosine_precision@30.5162
cosine_precision@50.3901
cosine_precision@100.2283
cosine_recall@10.1931
cosine_recall@30.513
cosine_recall@50.6297
cosine_recall@100.7295
cosine_ndcg@100.6421
cosine_mrr@100.5871
cosine_map@1000.6279
Information Retrieval
json
  {
      "truncate_dim": 128
  }
MetricValue
cosine_accuracy@10.4621
cosine_accuracy@30.507
cosine_accuracy@50.5904
cosine_accuracy@100.6754
cosine_precision@10.4621
cosine_precision@30.4415
cosine_precision@50.3391
cosine_precision@100.207
cosine_recall@10.165
cosine_recall@30.4373
cosine_recall@50.5447
cosine_recall@100.6602
cosine_ndcg@100.5661
cosine_mrr@100.5088
cosine_map@1000.5533
Information Retrieval
json
  {
      "truncate_dim": 64
  }
MetricValue
cosine_accuracy@10.3462
cosine_accuracy@30.3833
cosine_accuracy@50.4699
cosine_accuracy@100.5487
cosine_precision@10.3462
cosine_precision@30.3313
cosine_precision@50.2643
cosine_precision@100.1669
cosine_recall@10.1233
cosine_recall@30.3266
cosine_recall@50.4252
cosine_recall@100.5292
cosine_ndcg@100.442
cosine_mrr@100.3891
cosine_map@1000.4352

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

json
  • —Dataset: json
  • —Size: 5,822 training samples
  • —Columns: <code>positive</code> and <code>anchor</code>
  • —Approximate statistics based on the first 1000 samples: | | positive | anchor | |:--------|:------------------------------------------------------------------------------------|:----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 44 tokens</li><li>mean: 96.74 tokens</li><li>max: 170 tokens</li></ul> | <ul><li>min: 7 tokens</li><li>mean: 16.44 tokens</li><li>max: 40 tokens</li></ul> |
  • —Samples: | positive | anchor | |:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------------------------------------| | <code>uploaded to the classified system” after being located, First Viscuso Decl. ¶ 5 (emphasis added). <br>Relatedly, if responsive records are located elsewhere before being “scanned and uploaded to the <br>classified system,” see First Viscuso Decl. ¶ 5, but “all CIA systems of records are located on the <br>Agency’s classified system,” Second Viscuso Decl. ¶ 4, the CIA does not explain in which</code> | <code>In which paragraph of First Viscuso Declaration does it mention uploading to the classified system?</code> | | <code>1 <br> <br>UNITED STATES DISTRICT COURT <br>FOR THE DISTRICT OF COLUMBIA <br> <br>NATIONAL SECURITY COUNSELORS, <br> <br>Plaintiff, <br> <br>v. <br> <br>CENTRAL INTELLIGENCE AGENCY, et al., <br> <br>Defendants. <br> <br> <br> <br>Civil Action Nos. 11-443, 11-444, 11-445 <br>(BAH) <br> <br>Judge Beryl A. Howell <br> <br>MEMORANDUM OPINION <br>The plaintiff National Security Counselors (“NSC”) brought these three related actions</code> | <code>Who is the plaintiff in the case?</code> | | <code>although “agencies should not be forced to provide such a detailed justification that would itself <br>compromise the secret nature of potentially exempt information,” agencies “must be required to <br>provide the reasons behind their conclusions in order that they may be challenged by FOIA <br>plaintiffs and reviewed by the courts.” Id. To this end, the Circuit has said that “[i]n addition to</code> | <code>What are agencies not forced to provide?</code> |
  • —Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "MultipleNegativesRankingLoss",
      "matryoshka_dims": [
          768,
          512,
          256,
          128,
          64
      ],
      "matryoshka_weights": [
          1,
          1,
          1,
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: epoch
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —gradient_accumulation_steps: 16
  • —learning_rate: 2e-05
  • —num_train_epochs: 4
  • —lr_scheduler_type: cosine
  • —warmup_ratio: 0.1
  • —bf16: True
  • —load_best_model_at_end: True
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: epoch
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 16
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 2e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 4
  • —max_steps: -1
  • —lr_scheduler_type: cosine
  • —lr_scheduler_kwargs: None
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —bf16: True
  • —fp16: False
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: True
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —parallelism_config: None
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —project: huggingface
  • —trackio_space_id: trackio
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —hub_revision: None
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: no
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —liger_kernel_config: None
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: True
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional
  • —router_mapping: {}
  • —learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining Lossdim_768_cosine_ndcg@10dim_512_cosine_ndcg@10dim_256_cosine_ndcg@10dim_128_cosine_ndcg@10dim_64_cosine_ndcg@10
0.8791105.6803-----
1.012-0.63450.62640.59710.51850.3870
1.7033202.7279-----
2.024-0.66560.65990.64230.55140.4163
2.5275301.8881-----
3.036-0.67690.67060.64110.56360.4395
3.3516401.6804-----
4.048-0.67880.67070.64210.56610.442
  • —The bold row denotes the saved checkpoint.

Framework Versions

  • —Python: 3.12.12
  • —Sentence Transformers: 5.2.0
  • —Transformers: 4.57.6
  • —PyTorch: 2.9.1+cu128
  • —Accelerate: 1.12.0
  • —Datasets: 4.5.0
  • —Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MatryoshkaLoss
bibtex
@misc{kusupati2024matryoshka,
    title={Matryoshka Representation Learning},
    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
    year={2024},
    eprint={2205.13147},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->