vab46/nomic-embed-text-v1.5_Clinical-Trials_Matryoshka_final
nomic-embed text-v1.5 Clinical-Trials Matryoshka
This is a sentence-transformers model finetuned from nomic-ai/nomic-embed-text-v1.5 on the dataset from Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: nomic-ai/nomic-embed-text-v1.5 <!-- at revision e9b6763023c676ca8431644204f50c2b100d9aab -->
- Maximum Sequence Length: 8192 tokens
- Output Dimensionality: 768 dimensions
- Similarity Function: Cosine Similarity
- Supported Modality: Text
- Training Dataset: Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final
- Language: en
- License: apache-2.0
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'NomicBertModel'})
(1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("vab46/nomic-embed-text-v1.5_Clinical-Trials_Matryoshka_final")
# Run inference
documents = [
'TITLE: Safety and Efficacy of Six-Channel Radiofrequency Ablation System for Renal Denervation in Patients With Untreated Grade I Hypertension: a Pilot Study\nSUMMARY: Prospective, Multi-Center, Randomized, shame-Controlled, Uptake clinical trial to evaluate the efficacy and safety of the six-channel radio-frequency(RF) renal denervation system-comprising the six-channel RF generator (specification model: 25D1G, software release version: SRG-V1) and the disposable ultra-guiding RF denervation catheter (specification model: 25C6W127F115T)-for renal denervation in patients with grade I hypertension and without taking antihypertensive medicines.\nINCLUSION_CRITERIA: 1. Male or female, aged 18 to 65 years inclusive\n2. Hypertension duration longer than 3 months\n3. Hypertensive subjects who have been stopped taking antihypertensive drugs continuously and stably for at least 4 weeks or who do not take antihypertensive drugs , with office systolic/diastolic blood pressure still ≥140/90 mmHg and \\<160/100 mmHg, and 24-hour ambulatory mean systolic /diastolic pressure ≥130/80 mmHg and \\<140/90 mmHg;\n4. The subject or his/her legal representative fully understands the content of the informed consent form for this trial and voluntarily signs the written informed consent form.',
]
queries = [
"I have high blood pressure, but I'm not on any medication. Can I participate in this study?",
'What is the primary goal of this clinical trial regarding pain management in premature newborns?',
"I'm on a ventilator and feel really short of breath. Can anything be done to make me feel more comfortable?",
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 768] [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[0.5783, 0.0218, 0.0940]])<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
Evaluation
Metrics
Information Retrieval
- Dataset:
dim_768,dim_512,dim_256,dim_128,dim_64 - Evaluated with <code>InformationRetrievalEvaluator</code> with these parameters:
Comparison with the baseline model performance on same data
For the list of Sequential Evaluator score on nested mashtroyska embeddings [768, 512, 256, 128, 64]:-
Initial(pre-trained) score:- 0.405----------->Final(fine-tuned) score:- 0.56
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 7,079 training samples
- Columns: <code>positive</code> and <code>anchor</code>
- Approximate statistics based on the first 100 samples: | | positive | anchor | |:---------|:-------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------| | type | string | string | | modality | text | text | | details | <ul><li>min: 72 tokens</li><li>mean: 290.18 tokens</li><li>max: 675 tokens</li></ul> | <ul><li>min: 11 tokens</li><li>mean: 21.08 tokens</li><li>max: 42 tokens</li></ul> |
- Samples: | positive | anchor | |:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------| | <code>TITLE: Risk Classification and Prediction of Histopathological Subtypes in Basal Cell Carcinoma Using a CNN-Based Artificial Intelligence Model on Dermoscopic Images<br>SUMMARY: This retrospective observational study aims to develop and evaluate a convolutional neural network (CNN)-based artificial intelligence model for risk classification and histopathological subtype prediction of basal cell carcinoma (BCC) using clinical and dermoscopic images. Histopathologically confirmed BCC cases from a dermatology archive will be included. The primary objective is to assess the diagnostic performance of the CNN model in classifying BCC as low-risk or high-risk. Secondary objectives include predicting histopathological subtypes and comparing the model's performance with that of dermatology physicians. Histopathological diagnosis will serve as the reference standard. All archived data will be anonymized before analysis.<br>INCLUSIONCRITERIA:<br>Patients with histopathologically confirmed basal cell carc...</code> | <code>Are dermoscopic images with sufficient image quality and resolution required for participation in this trial?</code> | | <code>TITLE: Postural Control Mechanism During a Stationary Wheelchair Wheelie<br>SUMMARY: The purpose of this study was to investigate the postural control during a stationary wheelchair wheelie. A group of participants was recruited for this observational study. Participants were asked to perform and maintain the stationary phase of a wheelchair wheelie, during which kinematic and kinetic data were recorded.<br>INCLUSIONCRITERIA:<br>Spinal cord injury at or below the T1 (first thoracic) level.<br>Between 20 and 65 years of age.<br>Currently using a manual wheelchair as the primary means of mobility.<br>Able to perform and maintain a wheelchair wheelie for at least 20 seconds</code> | <code>Is there an age restriction for participants in this clinical trial, and if so, what is it?</code> | | <code>TITLE: The Effect of Mandala Art Therapy on Fear and Anxiety of Childbirth in Expectant Fathers Waiting Outside the Delivery Room: A Randomized Controlled Study<br>SUMMARY: This randomized controlled trial is designed to evaluate the effectiveness of mandala art therapy in reducing childbirth fear and state anxiety among fathers whose partners are nulliparous pregnant women and who are waiting in the obstetric ward waiting area during labor. Eligible participants will be randomly allocated to either the intervention or the control group. Fathers assigned to the intervention group will participate in a single 20-25-minute mandala art therapy session, whereas those in the control group will receive standard care without any additional intervention. Childbirth fear and state anxiety will be assessed using the Fathers' Birth Fear Scale and the State subscale of the Spielberger State-Trait Anxiety Inventory, respectively, at baseline and immediately following the intervention. The findings of ...</code> | <code>Could a father with a partner expecting their first child qualify for this study?</code> |
- Loss: <code>MatryoshkaLoss</code> with these parameters:
{
"loss": "MultipleNegativesRankingLoss",
"matryoshka_dims": [
768,
512,
256,
128,
64
],
"matryoshka_weights": [
1,
1,
1,
1,
1
],
"n_dims_per_step": -1
}Training Hyperparameters
Non-Default Hyperparameters
num_train_epochs: 4learning_rate: 2e-05lr_scheduler_type: cosinewarmup_steps: 0.1gradient_accumulation_steps: 8fp16: Trueper_device_eval_batch_size: 32load_best_model_at_end: Truebatch_sampler: no_duplicates
All Hyperparameters
<details><summary>Click to expand</summary>
per_device_train_batch_size: 8num_train_epochs: 4max_steps: -1learning_rate: 2e-05lr_scheduler_type: cosinelr_scheduler_kwargs: Nonewarmup_steps: 0.1optim: adamwtorchfusedoptim_args: Noneweight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08optim_target_modules: Nonegradient_accumulation_steps: 8average_tokens_across_devices: Truemax_grad_norm: 1.0label_smoothing_factor: 0.0bf16: Falsefp16: Truebf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Nonetorch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneuse_liger_kernel: Falseliger_kernel_config: Noneuse_cache: Falseneftune_noise_alpha: Nonetorch_empty_cache_steps: Noneauto_find_batch_size: Falselog_on_each_node: Truelogging_nan_inf_filter: Trueinclude_num_input_tokens_seen: nolog_level: passivelog_level_replica: warningdisable_tqdm: Falseproject: huggingfacetrackio_space_id: Nonetrackio_bucket_id: Nonetrackio_static_space_id: Noneper_device_eval_batch_size: 32prediction_loss_only: Trueeval_on_start: Falseeval_do_concat_batches: Trueeval_use_gather_object: Falseeval_accumulation_steps: Noneinclude_for_metrics: []batch_eval_metrics: Falsesave_only_model: Falsesave_on_each_node: Falseenable_jit_checkpoint: Falsepush_to_hub: Falsehub_private_repo: Nonehub_model_id: Nonehub_strategy: every_savehub_always_push: Falsehub_revision: Noneload_best_model_at_end: Trueignore_data_skip: Falserestore_callback_states_from_checkpoint: Falsefull_determinism: Falseseed: 42data_seed: Noneuse_cpu: Falseaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}parallelism_config: Nonedataloader_drop_last: Falsedataloader_num_workers: 0dataloader_pin_memory: Truedataloader_persistent_workers: Falsedataloader_prefetch_factor: Nonedataloader_multiprocessing_context: Nonedataloader_in_order: Trueremove_unused_columns: Truelabel_names: Nonetrain_sampling_strategy: randomlength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falseddp_static_graph: Noneddp_backend: Noneddp_timeout: 1800fsdp: Nonefsdp_config: Nonedeepspeed: Nonedebug: []skip_memory_metrics: Truedo_predict: Falseresume_from_checkpoint: Nonelocal_rank: -1prompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportionalrouter_mapping: {}learning_rate_mapping: {}warmup_ratio: None
</details>
Training Logs
- The bold row denotes the saved checkpoint.
Training Time
- Training: 31.6 minutes
Framework Versions
- Python: 3.13.15
- Sentence Transformers: 5.7.0
- Transformers: 5.15.0
- PyTorch: 2.11.0+cu128
- Accelerate: 1.14.0
- Datasets: 4.0.0
- Tokenizers: 0.22.2
Additional Resources
- Training and Finetuning Embedding Models with Sentence Transformers: the end-to-end guide for training or finetuning Sentence Transformer models.
- Introduction to Matryoshka Embedding Models: variable-size embeddings that can be truncated with minimal quality loss.
- Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval: post-training compression of embedding vectors.
- Multimodal Embedding & Reranker Models with Sentence Transformers: use text, image, audio, and video models through the same API.
- Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers: train multimodal embedding models, with a Visual Document Retrieval walkthrough.
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}MatryoshkaLoss
@misc{kusupati2024matryoshka,
title={Matryoshka Representation Learning},
author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
year={2024},
eprint={2205.13147},
archivePrefix={arXiv},
primaryClass={cs.LG}
}MultipleNegativesRankingLoss
@misc{oord2019representationlearningcontrastivepredictive,
title={Representation Learning with Contrastive Predictive Coding},
author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
year={2019},
eprint={1807.03748},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/1807.03748},
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
