shuaih777/music-crs-biencoder-personalized
SentenceTransformer based on BAAI/bge-large-en-v1.5
This is a sentence-transformers model finetuned from BAAI/bge-large-en-v1.5. It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: BAAI/bge-large-en-v1.5 <!-- at revision d4aa6901d3a41ba39fb536a557fa166f842b0e09 -->
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 1024 dimensions
- Similarity Function: Cosine Similarity <!-- - Training Dataset: Unknown --> <!-- - Language: Unknown --> <!-- - License: Unknown -->
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': True}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformersThen you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
'[20s | female | Germany]\nUser: Suggest some classic pop or dance tracks that bring back good memories.\nAssistant: Awesome! Glad you enjoyed "Thriller"! If you\'re looking to keep that energetic vibe going, how about\nUser: Oh, the Spice Girls! That definitely brings back memories, classic 90s pop. I like the energetic vibe. Can you suggest something similar, maybe still 90s but with a bit more of a dance beat?\nAssistant: Awesome, so glad you loved the Spice Girls! Sticking with that fun, upbeat 90s vibe, how about "MMMB\nUser: MMMBop is super nostalgic and fun, I remember that one! But it\'s not quite hitting that dance beat I was looking for. Could we try something with a stronger rhythm, still from the 90s or early 2000s, that\'s more geared for dancing?',
'The Nights by Avicii from The Days / Nights, Full Circle: The Live Anthology [Dance, 2014 releases, 2014 single, Pop, pop dance, house, playlista rmf fm, epic, 2014, 2010s, dance, hino, chillout, robbie williams, avicii fig] (2014)',
'Ziggy Stardust - 2012 Remastered Version by David Bowie from The Rise And Fall Of Ziggy Stardust And The Spiders From Mars, The Rise And Fall Of Ziggy Stardust And The Spiders From Mars (2012 Remastered Version) [Pop, classic rock] (1972)',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]<!--
Direct Usage (Transformers)
<details><summary>Click to see the direct usage in Transformers</summary>
</details> -->
<!--
Downstream Usage (Sentence Transformers)
You can finetune this model on your own dataset.
<details><summary>Click to expand</summary>
</details> -->
<!--
Out-of-Scope Use
List how the model may foreseeably be misused and address what users ought not to do with the model. -->
<!--
Bias, Risks and Limitations
What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model. -->
<!--
Recommendations
What are recommendations with respect to the foreseeable issues? For example, filtering explicit content. -->
Training Details
Training Dataset
Unnamed Dataset
- Size: 121,592 training samples
- Columns: <code>sentence0</code> and <code>sentence1</code>
- Approximate statistics based on the first 1000 samples: | | sentence0 | sentence1 | |:--------|:-------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------| | type | string | string | | details | <ul><li>min: 19 tokens</li><li>mean: 319.55 tokens</li><li>max: 512 tokens</li></ul> | <ul><li>min: 15 tokens</li><li>mean: 61.09 tokens</li><li>max: 112 tokens</li></ul> |
- Samples: | sentence0 | sentence1 | |:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | <code>[30s | male | Australia]<br>User: Can you help me find popular Christian music from around 2000-2010 that fans of the genre were really into, not just mainstream hits?<br>Assistant: Awesome, glad you liked "Who We Are"! Since you're enjoying the uplifting vibe, how about "Miracles"<br>User: That's a good one, "Miracles" by Newsboys definitely fits the Christian rock vibe and timeframe I'm looking for. Can you give me some more popular Christian rock and also some Christian hip-hop from that 2000-2010 era, like something really well-known by TobyMac or Lecrae?<br>Assistant: Glad you enjoyed "Miracles"! If you liked that beat and uplifting vibe, you've got to check out "Jes<br>User: These Christian rock suggestions like 'Miracles' and 'Jesus Freak' are exactly what I was looking for in terms of community favorites from that early 2000s-2010s era! Thank you for that. Now, could you give me some strong Christian hip-hop recommendations from that same period, perhaps by TobyMac or Lecrae, or similar artis...</code> | <code>Ignition by tobyMac from Portable Sounds With Bonus Remixes [christian hard rock, hard rock, cross rhythms, pump it up, lurv eet, very favorite, alternative rock, elimination chamber, christian rapcore, tobymac, cross2life, awesome, wwe ppv theme, epic, rock] (2008)</code> | | <code>[20s | male | Brazil]<br>User: I'm trying to remember a really popular Latin song from the 2000s, it had a unique sound and was everywhere.<br>Assistant: Awesome, glad you liked it! Let's keep the Legião Urbana train rolling. How about "Andrea Doria" nex<br>User: I love Legião Urbana, thanks for that! But I'm actually looking for something different. I'm trying to remember a Latin pop-rock song, definitely from the 2000s, with a very recognizable female vocalist.<br>Assistant: Awesome! Glad to hear you're liking the Legião Urbana picks. Let's keep it going with "Os Barcos," a<br>User: Okay, I really do love Legião Urbana, but I need to clarify – I'm looking for a different kind of Latin music, not Brazilian rock right now. Think more broadly Latin America, and definitely a pop-rock sound from the 2000s, with a prominent female singer.</code> | <code>Antes Das Seis by Legião Urbana from Mais Do Mesmo, Uma Outra Estação [pop rock nacional, rock alternativo, lyrics i love, uma outra estacao, rock brasilero, rockpop, 90s, brazilian rock, boo hoo, seis, balada, rock, para se matar de triste, canto junto, 10s collections] (1997)</code> | | <code>[60+ | female | Poland]<br>User: What are some popular alternative rock songs I might know from the 2000s or 2010s?<br>Assistant: Absolutely! Speaking of popular alternative rock from the 2010s, you might know "Sweater Weather" by<br>User: Yes, that's a good one, very recognizable! Are there any other tracks that were big hits for alternative bands around that time, perhaps something with a strong, recognizable chorus?</code> | <code>Sex on Fire by Kings of Leon from Balance Le Son, Only By The Night [usa, energetic, fucking awesome, alternative rock, 00s, obsessive repeated listening, amazing, completely awesome, tunes i have played for hours, catchy, favourites, gossip girl, american, british, awesome] (2008)</code> |
- Loss: <code>MultipleNegativesRankingLoss</code> with these parameters:
{
"scale": 20.0,
"similarity_fct": "cos_sim"
}Training Hyperparameters
Non-Default Hyperparameters
per_device_train_batch_size: 64per_device_eval_batch_size: 64multi_dataset_batch_sampler: round_robin
All Hyperparameters
<details><summary>Click to expand</summary>
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: noprediction_loss_only: Trueper_device_train_batch_size: 64per_device_eval_batch_size: 64per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 3max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'minnumparams': 0, 'xla': False, 'xlafsdpv2': False, 'xlafsdpgrad_ckpt': False}tp_size: 0fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'splitbatches': False, 'dispatchbatches': None, 'evenbatches': True, 'useseedablesampler': True, 'nonblocking': False, 'gradientaccumulationkwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robin
</details>
Training Logs
Framework Versions
- Python: 3.11.10
- Sentence Transformers: 3.3.1
- Transformers: 4.50.3
- PyTorch: 2.4.1+cu124
- Accelerate: 1.14.0
- Datasets: 5.0.0
- Tokenizers: 0.21.4
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}MultipleNegativesRankingLoss
@misc{henderson2017efficient,
title={Efficient Natural Language Response Suggestion for Smart Reply},
author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
year={2017},
eprint={1705.00652},
archivePrefix={arXiv},
primaryClass={cs.CL}
}<!--
Glossary
Clearly define terms in order to be accessible across audiences. -->
<!--
Model Card Authors
Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction. -->
<!--
Model Card Contact
Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors. -->
