dataisgod/bge-large-fiqa-financial
SentenceTransformer based on BAAI/bge-large-en-v1.5
This is a sentence-transformers model finetuned from BAAI/bge-large-en-v1.5. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: BAAI/bge-large-en-v1.5 <!-- at revision d4aa6901d3a41ba39fb536a557fa166f842b0e09 -->
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 1024 dimensions
- Similarity Function: Cosine Similarity
- Supported Modality: Text <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
(1): Pooling({'embedding_dimension': 1024, 'pooling_mode': 'cls', 'include_prompt': True})
(2): Normalize({})
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("dataisgod/bge-large-fiqa-financial")
# Run inference
queries = [
'Why is day trading considered riskier than long-term trading?',
]
documents = [
"In day trading, you're trying to predict the immediate fluctuations of an essentially random system. In long-term investing, you're trying to assess the strength of a company over a period of time. You also have frequent opportunities to assess your position and either add to it or get out.",
"It is a general truism but the reasons are that the rules change dramatically when you simply have more capital. Here are some examples, limited to particular kinds of markets: Under $2,000 in capital Nobody is going to offer you a margin account, and if you do get one it isn't with the best broker on commissions and other capabilities. So this means cash only trading, enjoy your 3 business day settlement periods. This means no shorting, confining a trader to only buy and hold strategies, making them more dependent on luck than a more capable trader. This means it is more expensive to buy stock, since you have to put down 100% of the cash to hold a share, whereas someone with more money puts down less capital to hold the exact same number of shares. This means no covered options strategies or spreads, again limiting the market directions where a trader could earn Under $25,000 in capital In the stock market, the pattern day trader rule applies to retail margin accounts with a balance under $25,000 and this severally limits the kinds of trades you are able to take because of the limit in the number of trades you can take in a given time period. Forget managing a multi-leg option position when the market isn't moving your direction. Under $125,000 in capital Worse margin rules. You excluded portfolio margin from your post, but it is a key part of the answer Over $1,000,000 in capital Participate in private placements, regulation D offerings reserved for accredited investors. These days, as buy and hold investments, these generally have more growth potential than publicly traded offerings. Over $5,000,000 in capital You can easily get the compliance and risk manager to turn the other way on margin rules. This is not conjecture, leverage up to infinity, try not to bankrupt yourself and the trading firm.",
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 1024] [2, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[0.7313, 0.5647]])<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
Evaluation
Metrics
Information Retrieval
- Dataset:
fiqa-test - Evaluated with <code>InformationRetrievalEvaluator</code>
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 14,166 training samples
- Columns: <code>anchor</code> and <code>positive</code>
- Approximate statistics based on the first 100 samples: | | anchor | positive | |:---------|:----------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------| | type | string | string | | modality | text | text | | details | <ul><li>min: 8 tokens</li><li>mean: 16.55 tokens</li><li>max: 31 tokens</li></ul> | <ul><li>min: 14 tokens</li><li>mean: 191.89 tokens</li><li>max: 512 tokens</li></ul> |
- Samples: | anchor | positive | |:-----------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>New to investing — I have $20,000 cash saved, what should I do with it?</code> | <code>My advice to you is not to take any advice from anyone when it comes to investing, especially when you don't know much about what you are investing in. mbhunter is correct, take your time to learn about what you want to invest in. If your goal at the moment is short term don't invest in stocks unless you really know what you are doing. Put your money where you can get the highest interest rate, continue saving and do a lot of research on the house you wish to buy. Even if you are not ready to buy a house yet, start looking so that by the time you are ready to buy, you know how much the house is really worth. Before buying our house we spent about 7 months looking and researching and looked at more than 100 houses.</code> | | <code>Are you preparing for a possible dollar (USD) collapse? (How?)</code> | <code>"Buying gold, silver, palladium, copper and platinum. The first two I am thinking about new currencies. The last three for the perpetual need for the metals in industry. I also have invested in Numismatic coins. They are small portable and easy to hide around the house. I only collect silver coins, so even if the world really blows up and numismatics goes out the window, I can depend on them forming a barter system through the content value of the silver. The problem with collectable items is that they are easy to see. For example, a nice painting just shouts out ""steal me!"". I don't buy large gold coins. As long as the coin is below 1/4 Oz gold I collect it. If the dollar does finaly collapse, to be honest it will be so bad that I think weapons will be order of the day. Do I think it will collapse...nah never."</code> | | <code>Can an unmarried couple buy a home together with only one person on the mortgage?</code> | <code>In this case can the title of the home still be held by both? Yes, it is possible to have additional people on title that are not on the mortgage. Would the lender (bank) have any reservations about this since a party not on the mortgage has ownership of the property? Possibly, but there is a very simple way to avoid this. Clayton could simply purchase the home himself, and add Emma to the title after closing by recording a quitclaim deed. The lender can't stop that, and from their point of view it's actually better, since they have two people to go after in the case of default. (But despite it being better they often make it difficult to purchase Tip, when you have an attorney draft the quitclaim document, have them draft the reverse document too. (Emma relinquishing the property back to Clayton.) There is usually no extra charge for this and then you have it if you need it. For example, you may need to file the reverse forms if you want to refinance. As a side note, I agree with Gra...</code> |
- Loss: <code>CachedMultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim",
"mini_batch_size": 8,
"mini_batch_num_tokens": null,
"gather_across_devices": false,
"directions": [
"query_to_doc"
],
"partition_mode": "joint",
"hardness_mode": null,
"hardness_strength": 0.0
}Training Hyperparameters
Non-Default Hyperparameters
per_device_train_batch_size: 64num_train_epochs: 1learning_rate: 2e-05warmup_steps: 0.1fp16: Truebatch_sampler: no_duplicates
All Hyperparameters
<details><summary>Click to expand</summary>
per_device_train_batch_size: 64num_train_epochs: 1max_steps: -1learning_rate: 2e-05lr_scheduler_type: linearlr_scheduler_kwargs: Nonewarmup_steps: 0.1optim: adamwtorchfusedoptim_args: Noneweight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08optim_target_modules: Nonegradient_accumulation_steps: 1average_tokens_across_devices: Truemax_grad_norm: 1.0label_smoothing_factor: 0.0bf16: Falsefp16: Truebf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Nonetorch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneuse_liger_kernel: Falseliger_kernel_config: Noneuse_cache: Falseneftune_noise_alpha: Nonetorch_empty_cache_steps: Noneauto_find_batch_size: Falselog_on_each_node: Truelogging_nan_inf_filter: Trueinclude_num_input_tokens_seen: nolog_level: passivelog_level_replica: warningdisable_tqdm: Falseproject: huggingfacetrackio_space_id: Nonetrackio_bucket_id: Nonetrackio_static_space_id: Noneper_device_eval_batch_size: 8prediction_loss_only: Trueeval_on_start: Falseeval_do_concat_batches: Trueeval_use_gather_object: Falseeval_accumulation_steps: Noneinclude_for_metrics: []batch_eval_metrics: Falsesave_only_model: Falsesave_on_each_node: Falseenable_jit_checkpoint: Falsepush_to_hub: Falsehub_private_repo: Nonehub_model_id: Nonehub_strategy: every_savehub_always_push: Falsehub_revision: Noneload_best_model_at_end: Falseignore_data_skip: Falserestore_callback_states_from_checkpoint: Falsefull_determinism: Falseseed: 42data_seed: Noneuse_cpu: Falseaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}parallelism_config: Nonedataloader_drop_last: Falsedataloader_num_workers: 0dataloader_pin_memory: Truedataloader_persistent_workers: Falsedataloader_prefetch_factor: Nonedataloader_multiprocessing_context: Nonedataloader_in_order: Trueremove_unused_columns: Truelabel_names: Nonetrain_sampling_strategy: randomlength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falseddp_static_graph: Noneddp_backend: Noneddp_timeout: 1800fsdp: Nonefsdp_config: Nonedeepspeed: Nonedebug: []skip_memory_metrics: Truedo_predict: Falseresume_from_checkpoint: Nonelocal_rank: -1prompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportionalrouter_mapping: {}learning_rate_mapping: {}warmup_ratio: None
</details>
Training Logs
Training Time
- Training: 51.5 minutes
Framework Versions
- Python: 3.13.15
- Sentence Transformers: 5.7.0
- Transformers: 5.16.1
- PyTorch: 2.11.0+cu128
- Accelerate: 1.14.0
- Datasets: 3.6.0
- Tokenizers: 0.23.1
Additional Resources
- Training and Finetuning Embedding Models with Sentence Transformers: the end-to-end guide for training or finetuning Sentence Transformer models.
- Introduction to Matryoshka Embedding Models: variable-size embeddings that can be truncated with minimal quality loss.
- Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval: post-training compression of embedding vectors.
- Multimodal Embedding & Reranker Models with Sentence Transformers: use text, image, audio, and video models through the same API.
- Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers: train multimodal embedding models, with a Visual Document Retrieval walkthrough.
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}CachedMultipleNegativesRankingLoss
@misc{gao2021scaling,
title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
year={2021},
eprint={2101.06983},
archivePrefix={arXiv},
primaryClass={cs.LG}
}MultipleNegativesRankingLoss
@misc{oord2019representationlearningcontrastivepredictive,
title={Representation Learning with Contrastive Predictive Coding},
author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
year={2019},
eprint={1807.03748},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/1807.03748},
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
