CoolFace
Modelpublic

MinhPhuc0804/jina-docling-checkthat-task1-v1

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes27downloads
Model Card

SentenceTransformer based on jinaai/jina-embeddings-v5-text-nano-retrieval

This is a sentence-transformers model finetuned from jinaai/jina-embeddings-v5-text-nano-retrieval. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: jinaai/jina-embeddings-v5-text-nano-retrieval <!-- at revision ac5d898c8d382b17167c33e5c8af644a3519b47d -->
  • Maximum Sequence Length: 256 tokens
  • Output Dimensionality: 768 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'EuroBertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'lasttoken', 'include_prompt': True})
  (2): Normalize({})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

bash
pip install -U sentence-transformers

Then you can load this model and run inference.

python
from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("MinhPhuc0804/jina-docling-checkthat-task1-v1")
# Run inference
queries = [
    "query: @MikeHudema Check out this review article for further discussion on why fossil fuels will keep supplying the bulk of global energy and why electric vehicles aren't exactly eco‑friendly",
]
documents = [
    'passage: title: Is it the end of combustion and engine combustion research? Should it be?\nabstract: The dominant narrative in the affluent west is that climate change poses an "existential threat" and very rapid cuts in greenhouse gas (GHG) emissions and hence fossil fuel use are needed to avoid it. Simultaneously oil, gas, coal, aviation, steel and cement industries and livestock farming have to be largely shut down to eliminate GHG. This review argues that globally all this will not happen by 2050, let alone 2030 because the scale of the problem is too large. Transport is particularly difficult to decarbonize and current policies focusing entirely on battery electric vehicles will not and must not succeed. GHG levels are unlikely to come down significantly in the next several decades and even if they did, extreme weather events will not disappear. It is better to recognize such realities and make societies more resilient to the effects of climate change. Humanity will have to adapt to any further warming as it has successfully done with the previous warming of about 1.1 C over the past century. Combustion research, particularly of fossil fuels and in internal combustion engines is currently seen as unnecessary in many countries.',
    'passage: title: SARS-CoV-2 infection confers greater immunity than shots\nabstract: Study from Israel, the largest of its kind, also finds infection combined with a single jab is highly protective',
    'passage: title: Neurologic Involvement in Children and Adolescents Hospitalized in the United States for COVID-19 or Multisystem Inflammatory Syndrome\nabstract: Coronavirus disease 2019 (COVID-19) affects the nervous system in adult patients. The spectrum of neurologic involvement in children and adolescents is unclear.To understand the range and severity of neurologic involvement among children and adolescents associated with COVID-19.Case series of patients (age <21 years) hospitalized between March 15, 2020, and December 15, 2020, with positive severe acute respiratory syndrome coronavirus 2 test result (reverse transcriptase-polymerase chain reaction and/or antibody) at 61 US hospitals in the Overcoming COVID-19 public health registry, including 616 (36%) meeting criteria for multisystem inflammatory syndrome in children. Patients with neurologic involvement had acute neurologic signs, symptoms, or diseases on presentation or during hospitalization.',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 768] [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[ 0.5678, -0.0192, -0.0984]])

<!--

Direct Usage (Transformers)

<details><summary>Click to see the direct usage in Transformers</summary>

</details> -->

<!--

Downstream Usage (Sentence Transformers)

You can finetune this model on your own dataset.

<details><summary>Click to expand</summary>

</details> -->

<!--

Out-of-Scope Use

List how the model may foreseeably be misused and address what users ought not to do with the model. -->

Evaluation

Metrics

Information Retrieval
MetricValue
cosine_accuracy@10.5299
cosine_accuracy@30.7273
cosine_accuracy@50.787
cosine_accuracy@100.8447
cosine_precision@10.5299
cosine_precision@30.2424
cosine_precision@50.1574
cosine_precision@100.0845
cosine_recall@10.5299
cosine_recall@30.7273
cosine_recall@50.787
cosine_recall@100.8447
cosine_ndcg@100.6898
cosine_mrr@100.6399
cosine_map@1000.6446

<!--

Bias, Risks and Limitations

What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->

<!--

Recommendations

What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->

Training Details

Training Dataset

Unnamed Dataset
  • Size: 17,319 training samples
  • Columns: <code>query</code> and <code>positive</code>
  • Approximate statistics based on the first 1000 samples: | | query | positive | |:--------|:------------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 21 tokens</li><li>mean: 53.57 tokens</li><li>max: 144 tokens</li></ul> | <ul><li>min: 27 tokens</li><li>mean: 183.22 tokens</li><li>max: 256 tokens</li></ul> |
  • Samples: | query | positive | |:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>query: @user Abgesehen von den bekannten Nebenwirkungen wie Thrombosen und Neurologischen Schäden gibt es zum Beispiel jedoch begutachtete Anhaltspunkte, dass das Protein die DNA schädigt. Da habe ich keine Lust darauf.</code> | <code>passage: its DNA damage repair by impeding key DNA repair protein BRCA1 and 53BP1 recruitment to the damage site.<br><br>title: RETRACTED: SARS–CoV–2 Spike Impairs DNA Damage Repair and Inhibits V(D)J Recombination In Vitro<br>Our findings reveal a potential molecular mechanism by which the spike protein might impede adaptive immunity and underscore the potential side effects of full-length spike-based vaccines.</code> | | <code>query: de plus le Japon emploie également les échantillons salivaires ; ci le lien vers l’article de Yale, je vous prie de vérifier les points que je présente et de corriger. 3/</code> | <code>passage: ngeal and saliva samples from confirmed COVID-19 patients and self-collected samples from healthcare workers on COVID-19 wards.<br><br>title: Saliva is more sensitive for SARS-CoV-2 detection in COVID-19 patients than nasopharyngeal swabs<br>When we compared SARS-CoV-2 detection from patient-matched nasopharyngeal and saliva samples, we found that saliva yielded greater detection sensitivity and consistency throughout the course of infection. Furthermore, we report less variability in self-sample collection of saliva. Taken together, our findings demonstrate that saliva is a viable and more sensitive alternative to nasopharyngeal swabs and could enable at-home self-administered sample collection for accurate large-scale SARS-CoV-2 testing.</code> | | <code>query: Are psychedelics the road to comprehension? Interpersonal Dynamics in Ayahuasca circles of Palestinians and Israelis</code> | <code>passage: title: Relational Processes in Ayahuasca Groups of Palestinians and Israelis abstract: Psychedelics are used in many group contexts. However, most phenomenological research on psychedelics is focused on personal experiences. This paper presents a phenomenological investigation centered on intersubjective and intercultural relational processes, exploring how an intercultural context affects both the group and individual process. Through 31 in-depth interviews, ceremonies in which Palestinians and Israelis drink ayahuasca together have been investigated. The overarching question guiding this inquiry was how psychedelics might contribute to processes of peacebuilding, and in particular how an intercultural context, embedded in a protracted conflict, would affect the group's psychedelic process in a relational sense. Analysis of the interviews was based on grounded theory. Three relational themes about multilocal participatory events which occurred during ayahuasca rituals have em...</code> |
  • Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
json
  {
      "scale": 20.0,
      "similarity_fct": "cos_sim",
      "gather_across_devices": false,
      "directions": [
          "query_to_doc"
      ],
      "partition_mode": "joint",
      "hardness_mode": null,
      "hardness_strength": 0.0
  }

Training Hyperparameters

Non-Default Hyperparameters
  • per_device_train_batch_size: 32
  • learning_rate: 8e-06
  • num_train_epochs: 10
  • warmup_steps: 542
  • bf16: True
  • dataloader_drop_last: True
  • load_best_model_at_end: True
All Hyperparameters

<details><summary>Click to expand</summary>

  • overwrite_output_dir: False
  • do_predict: False
  • prediction_loss_only: True
  • per_device_train_batch_size: 32
  • per_device_eval_batch_size: 8
  • per_gpu_train_batch_size: None
  • per_gpu_eval_batch_size: None
  • gradient_accumulation_steps: 1
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 8e-06
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1.0
  • num_train_epochs: 10
  • max_steps: -1
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: {}
  • warmup_ratio: 0.0
  • warmup_steps: 542
  • log_level: passive
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • save_safetensors: True
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • no_cuda: False
  • use_cpu: False
  • use_mps_device: False
  • seed: 42
  • data_seed: None
  • jit_mode_eval: False
  • use_ipex: False
  • bf16: True
  • fp16: False
  • fp16_opt_level: O1
  • half_precision_backend: auto
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • local_rank: 0
  • ddp_backend: None
  • tpu_num_cores: None
  • tpu_metrics_debug: False
  • debug: []
  • dataloader_drop_last: True
  • dataloader_num_workers: 0
  • dataloader_prefetch_factor: None
  • past_index: -1
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: True
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_min_num_params: 0
  • fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}
  • fsdp_transformer_layer_cls_to_wrap: None
  • accelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}
  • parallelism_config: None
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamwtorchfused
  • optim_args: None
  • adafactor: False
  • group_by_length: False
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • use_legacy_prediction_loop: False
  • push_to_hub: False
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_private_repo: None
  • hub_always_push: False
  • hub_revision: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_inputs_for_metrics: False
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • fp16_backend: auto
  • push_to_hub_model_id: None
  • push_to_hub_organization: None
  • mp_parameters:
  • auto_find_batch_size: False
  • full_determinism: False
  • torchdynamo: None
  • ray_scope: last
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • include_tokens_per_second: False
  • include_num_input_tokens_seen: False
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: False
  • use_liger_kernel: False
  • liger_kernel_config: None
  • eval_use_gather_object: False
  • average_tokens_across_devices: False
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

</details>

Training Logs

EpochStepTraining Loss10-percent-dev-split_cosine_ndcg@10
0.92425000.6731-
1.0018542-0.6814
1.848410000.3426-
2.00371084-0.6826
2.772615000.2901-
3.00551626-0.6871
3.696920000.2571-
4.00742168-0.6886
4.621125000.2435-
5.00922710-0.6887
5.545330000.2322-
6.01113252-0.6896
6.469535000.2207-
7.01293794-0.6902
7.393740000.2199-
8.01484336-0.6898
8.317945000.2194-
9.01664878-0.6898
9.242150000.2129-
  • The bold row denotes the saved checkpoint.

Training Time

  • Training: 16.2 minutes

Framework Versions

  • Python: 3.12.6
  • Sentence Transformers: 5.4.1
  • Transformers: 4.56.0
  • PyTorch: 2.8.0+cu129
  • Accelerate: 1.10.1
  • Datasets: 4.8.4
  • Tokenizers: 0.22.0

Citation

BibTeX

Sentence Transformers
bibtex
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
MultipleNegativesRankingLoss
bibtex
@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}

<!--

Glossary

Clearly define terms in order to be accessible across audiences. -->

<!--

Model Card Authors

Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->

<!--

Model Card Contact

Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->