CoolFace
Modelpublic

digo-prayudha/test-modernbert-embed-base-legal-matryoshka-2

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes78downloads
Model Card

ModernBERT Embed base Legal Matryoshka

This is a sentence-transformers model finetuned from nomic-ai/modernbert-embed-base on the json dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • —Model Type: Sentence Transformer
  • —Base model: nomic-ai/modernbert-embed-base <!-- at revision d556a88e332558790b210f7bdbe87da2fa94a8d8 -->
  • —Maximum Sequence Length: 8192 tokens
  • —Output Dimensionality: 768 dimensions
  • —Similarity Function: Cosine Similarity
  • —Training Dataset:
  • —json
  • —Language: en
  • —License: apache-2.0

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 8192, 'do_lower_case': False, 'architecture': 'ModernBertModel'})
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("digo-prayudha/test-modernbert-embed-base-legal-matryoshka-2")
# Run inference
sentences = [
    'are “broad and vague descriptions” of the information that make it impossible for the Court to \nconduct its requisite de novo review over the Department’s decision to withhold this information \nas “critical infrastructure security information.”  See Prop. of the People, Inc. v. Off. of Mgmt. & \nBudget, 330 F. Supp. 3d 373, 388 (D.D.C. 2018).',
    'What type of review is the Court unable to conduct due to broad and vague descriptions?',
    'What did the court express skepticism about?',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.4837, 0.0878],
#         [0.4837, 1.0000, 0.3375],
#         [0.0878, 0.3375, 1.0000]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
json
  {
      "truncate_dim": 768
  }
MetricValue
cosine_accuracy@10.5425
cosine_accuracy@30.5981
cosine_accuracy@50.7079
cosine_accuracy@100.7713
cosine_precision@10.5425
cosine_precision@30.5126
cosine_precision@50.3966
cosine_precision@100.2362
cosine_recall@10.2031
cosine_recall@30.5176
cosine_recall@50.6473
cosine_recall@100.7621
cosine_ndcg@100.6606
cosine_mrr@100.5967
cosine_map@1000.641
Information Retrieval
json
  {
      "truncate_dim": 512
  }
MetricValue
cosine_accuracy@10.5224
cosine_accuracy@30.5765
cosine_accuracy@50.6862
cosine_accuracy@100.7713
cosine_precision@10.5224
cosine_precision@30.492
cosine_precision@50.3805
cosine_precision@100.234
cosine_recall@10.1949
cosine_recall@30.4992
cosine_recall@50.6256
cosine_recall@100.7563
cosine_ndcg@100.6452
cosine_mrr@100.5773
cosine_map@1000.6225
Information Retrieval
json
  {
      "truncate_dim": 256
  }
MetricValue
cosine_accuracy@10.5085
cosine_accuracy@30.5487
cosine_accuracy@50.6445
cosine_accuracy@100.7512
cosine_precision@10.5085
cosine_precision@30.4781
cosine_precision@50.3632
cosine_precision@100.2287
cosine_recall@10.1877
cosine_recall@30.4813
cosine_recall@50.5922
cosine_recall@100.7353
cosine_ndcg@100.6249
cosine_mrr@100.5591
cosine_map@1000.6024
Information Retrieval
json
  {
      "truncate_dim": 128
  }
MetricValue
cosine_accuracy@10.4389
cosine_accuracy@30.4838
cosine_accuracy@50.5688
cosine_accuracy@100.6615
cosine_precision@10.4389
cosine_precision@30.4173
cosine_precision@50.3255
cosine_precision@100.1995
cosine_recall@10.1601
cosine_recall@30.415
cosine_recall@50.5274
cosine_recall@100.6446
cosine_ndcg@100.5452
cosine_mrr@100.4867
cosine_map@1000.5329
Information Retrieval
json
  {
      "truncate_dim": 64
  }
MetricValue
cosine_accuracy@10.3431
cosine_accuracy@30.3663
cosine_accuracy@50.456
cosine_accuracy@100.544
cosine_precision@10.3431
cosine_precision@30.3174
cosine_precision@50.2498
cosine_precision@100.1623
cosine_recall@10.127
cosine_recall@30.3226
cosine_recall@50.4149
cosine_recall@100.5283
cosine_ndcg@100.4373
cosine_mrr@100.3836
cosine_map@1000.4284

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

json
  • —Dataset: json
  • —Size: 5,822 training samples
  • —Columns: <code>positive</code> and <code>anchor</code>
  • —Approximate statistics based on the first 1000 samples: | | positive | anchor | |:--------|:------------------------------------------------------------------------------------|:----------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 26 tokens</li><li>mean: 97.49 tokens</li><li>max: 160 tokens</li></ul> | <ul><li>min: 8 tokens</li><li>mean: 16.58 tokens</li><li>max: 41 tokens</li></ul> |
  • —Samples: | positive | anchor | |:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------------------------| | <code>policy” or “are in fact properly classified pursuant to such Executive order” are exempt from <br>production under the FOIA. See 5 U.S.C. § 552(b)(1). “[I]n the FOIA context, [the D.C. Circuit <br>has] consistently deferred to executive affidavits predicting harm to the national security, and <br>have found it unwise to undertake searching judicial review.” Ctr. for Nat’l Sec. Studies, 331</code> | <code>What has the D.C. Circuit consistently deferred to in the FOIA context?</code> | | <code>42 The plaintiff states in its briefing that it challenges the CIA’s withholding of two records, in part, in No. 11-443, <br>see Pl.’s First 443 Opp’n at 14, and six documents, in part, in No. 11-444, see Pl.’s First 444 Opp’n at 30, 35. The <br>plaintiff does not specify, however, exactly which Exemption 3 withholdings it challenges in No. 11-445, where the</code> | <code>How many records does the plaintiff challenge the withholding of in part in No. 11-443?</code> | | <code>let alone as a vexing subject of intense legal debate. <br>¶ 46 <br> <br>Indeed, the question of anonymity has taken on increased significance as court records <br>have become readily available to the general public through even casual Internet searches. As <br>the appellant notes in his brief, a Google search of a litigant’s name can produce an untold</code> | <code>In which paragraph is the issue of anonymity discussed?</code> |
  • —Loss: <code>MatryoshkaLoss</code> with these parameters:
json
  {
      "loss": "MultipleNegativesRankingLoss",
      "matryoshka_dims": [
          768,
          512,
          256,
          128,
          64
      ],
      "matryoshka_weights": [
          1,
          1,
          1,
          1,
          1
      ],
      "n_dims_per_step": -1
  }

Training Hyperparameters

Non-Default Hyperparameters
  • —eval_strategy: epoch
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —gradient_accumulation_steps: 16
  • —learning_rate: 2e-05
  • —num_train_epochs: 4
  • —lr_scheduler_type: cosine
  • —warmup_ratio: 0.1
  • —bf16: True
  • —load_best_model_at_end: True
  • —batch_sampler: no_duplicates
All Hyperparameters

<details><summary>Click to expand</summary>

  • —overwrite_output_dir: False
  • —do_predict: False
  • —eval_strategy: epoch
  • —prediction_loss_only: True
  • —per_device_train_batch_size: 32
  • —per_device_eval_batch_size: 16
  • —per_gpu_train_batch_size: None
  • —per_gpu_eval_batch_size: None
  • —gradient_accumulation_steps: 16
  • —eval_accumulation_steps: None
  • —torch_empty_cache_steps: None
  • —learning_rate: 2e-05
  • —weight_decay: 0.0
  • —adam_beta1: 0.9
  • —adam_beta2: 0.999
  • —adam_epsilon: 1e-08
  • —max_grad_norm: 1.0
  • —num_train_epochs: 4
  • —max_steps: -1
  • —lr_scheduler_type: cosine
  • —lr_scheduler_kwargs: {}
  • —warmup_ratio: 0.1
  • —warmup_steps: 0
  • —log_level: passive
  • —log_level_replica: warning
  • —log_on_each_node: True
  • —logging_nan_inf_filter: True
  • —save_safetensors: True
  • —save_on_each_node: False
  • —save_only_model: False
  • —restore_callback_states_from_checkpoint: False
  • —no_cuda: False
  • —use_cpu: False
  • —use_mps_device: False
  • —seed: 42
  • —data_seed: None
  • —jit_mode_eval: False
  • —use_ipex: False
  • —bf16: True
  • —fp16: False
  • —fp16_opt_level: O1
  • —half_precision_backend: auto
  • —bf16_full_eval: False
  • —fp16_full_eval: False
  • —tf32: None
  • —local_rank: 0
  • —ddp_backend: None
  • —tpu_num_cores: None
  • —tpu_metrics_debug: False
  • —debug: []
  • —dataloader_drop_last: False
  • —dataloader_num_workers: 0
  • —dataloader_prefetch_factor: None
  • —past_index: -1
  • —disable_tqdm: False
  • —remove_unused_columns: True
  • —label_names: None
  • —load_best_model_at_end: True
  • —ignore_data_skip: False
  • —fsdp: []
  • —fsdp_min_num_params: 0
  • —fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • —fsdp_transformer_layer_cls_to_wrap: None
  • —accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • —parallelism_config: None
  • —deepspeed: None
  • —label_smoothing_factor: 0.0
  • —optim: adamwtorchfused
  • —optim_args: None
  • —adafactor: False
  • —group_by_length: False
  • —length_column_name: length
  • —ddp_find_unused_parameters: None
  • —ddp_bucket_cap_mb: None
  • —ddp_broadcast_buffers: False
  • —dataloader_pin_memory: True
  • —dataloader_persistent_workers: False
  • —skip_memory_metrics: True
  • —use_legacy_prediction_loop: False
  • —push_to_hub: False
  • —resume_from_checkpoint: None
  • —hub_model_id: None
  • —hub_strategy: every_save
  • —hub_private_repo: None
  • —hub_always_push: False
  • —hub_revision: None
  • —gradient_checkpointing: False
  • —gradient_checkpointing_kwargs: None
  • —include_inputs_for_metrics: False
  • —include_for_metrics: []
  • —eval_do_concat_batches: True
  • —fp16_backend: auto
  • —push_to_hub_model_id: None
  • —push_to_hub_organization: None
  • —mp_parameters:
  • —auto_find_batch_size: False
  • —full_determinism: False
  • —torchdynamo: None
  • —ray_scope: last
  • —ddp_timeout: 1800
  • —torch_compile: False
  • —torch_compile_backend: None
  • —torch_compile_mode: None
  • —include_tokens_per_second: False
  • —include_num_input_tokens_seen: False
  • —neftune_noise_alpha: None
  • —optim_target_modules: None
  • —batch_eval_metrics: False
  • —eval_on_start: False
  • —use_liger_kernel: False
  • —liger_kernel_config: None
  • —eval_use_gather_object: False
  • —average_tokens_across_devices: False
  • —prompts: None
  • —batch_sampler: no_duplicates
  • —multi_dataset_batch_sampler: proportional
  • —router_mapping: {}
  • —learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining Lossdim_768_cosine_ndcg@10dim_512_cosine_ndcg@10dim_256_cosine_ndcg@10dim_128_cosine_ndcg@10dim_64_cosine_ndcg@10
0.8791105.6715-----
1.012-0.62810.60630.56950.48890.3602
1.7033202.5845-----
2.024-0.66440.64670.61440.53790.4200
2.5275302.0086-----
3.036-0.66270.64550.62280.54710.4365
3.3516401.6748-----
4.048-0.66060.64520.62490.54520.4373
  • —The bold row denotes the saved checkpoint.

Framework Versions

  • —Python: 3.12.11
  • —Sentence Transformers: 5.1.0
  • —Transformers: 4.56.1
  • —PyTorch: 2.8.0+cu126
  • —Accelerate: 1.10.1
  • —Datasets: 4.0.0
  • —Tokenizers: 0.22.0

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MatryoshkaLoss
bibtex
@misc{kusupati2024matryoshka,
    title={Matryoshka Representation Learning},
    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
    year={2024},
    eprint={2205.13147},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}
MultipleNegativesRankingLoss
bibtex
@misc{henderson2017efficient,
    title={Efficient Natural Language Response Suggestion for Smart Reply},
    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
    year={2017},
    eprint={1705.00652},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->