CoolFace
Modelpublic

jinaai/jina-embedding-b-en-v1

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
8likes1.5kdownloads
Model Card

<br><br>

<p align="center"> <img src="https://huggingface.co/datasets/jinaai/documentation-images/resolve/main/logo.webp" alt="Jina AI: Your Search Foundation, Supercharged!" width="150px"> </p>

<p align="center"> <b>The text embedding set trained by <a href="https://jina.ai/"><b>Jina AI</b></a></b> </p>

Intented Usage & Model Info

jina-embedding-b-en-v1 is a language model that has been trained using Jina AI's Linnaeus-Clean dataset. This dataset consists of 380 million pairs of sentences, which include both query-document pairs. These pairs were obtained from various domains and were carefully selected through a thorough cleaning process. The Linnaeus-Full dataset, from which the Linnaeus-Clean dataset is derived, originally contained 1.6 billion sentence pairs.

The model has a range of use cases, including information retrieval, semantic textual similarity, text reranking, and more.

With a standard size of 110 million parameters, the model enables fast inference while delivering better performance than our small model. It is recommended to use a single GPU for inference. Additionally, we provide the following options:

Data & Parameters

Please checkout our technical blog.

Metrics

We compared the model against all-minilm-l6-v2/all-mpnet-base-v2 from sbert and text-embeddings-ada-002 from OpenAI:

Nameparamdimension
all-minilm-l6-v223m384
all-mpnet-base-v2110m768
ada-embedding-002Unknown/OpenAI API1536
jina-embedding-t-en-v114m312
jina-embedding-s-en-v135m512
jina-embedding-b-en-v1110m768
jina-embedding-l-en-v1330m1024
NameSTS12STS13STS14STS15STS16STS17TRECOVIDQuoraSciFact
all-minilm-l6-v20.7240.8060.7560.8540.790.8760.4730.8760.645
all-mpnet-base-v20.7260.8350.780.8570.80.9060.5130.8750.656
ada-embedding-0020.6980.8330.7610.8610.860.9030.6850.8760.726
jina-embedding-t-en-v10.7170.7730.7310.8290.7770.8600.4820.8400.522
jina-embedding-s-en-v10.7430.7860.7380.8370.800.8750.5230.8570.524
jina-embedding-b-en-v10.7510.8090.7610.8560.8120.8900.6060.8760.594
jina-embedding-l-en-v10.7450.8320.7810.8690.8370.9020.5730.8810.598

Usage

Usage with Jina AI Finetuner:

python
!pip install finetuner
import finetuner

model = finetuner.build_model('jinaai/jina-embedding-b-en-v1')
embeddings = finetuner.encode(
    model=model,
    data=['how is the weather today', 'What is the current weather like today?']
)
print(finetuner.cos_sim(embeddings[0], embeddings[1]))

Use with sentence-transformers:

python
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim

sentences = ['how is the weather today', 'What is the current weather like today?']

model = SentenceTransformer('jinaai/jina-embedding-b-en-v1')
embeddings = model.encode(sentences)
print(cos_sim(embeddings[0], embeddings[1]))

Fine-tuning

Please consider Finetuner.

Plans

  1. 1.The development of jina-embedding-s-en-v2 is currently underway with two main objectives: improving performance and increasing the maximum sequence length.
  2. 2.We are currently working on a bilingual embedding model that combines English and X language. The upcoming model will be called jina-embedding-s/b/l-de-v1.

Contact

Join our Discord community and chat with other community members about ideas.

Citation

If you find Jina Embeddings useful in your research, please cite the following paper:

latex
@misc{günther2023jina,
      title={Jina Embeddings: A Novel Set of High-Performance Sentence Embedding Models},
      author={Michael Günther and Louis Milliken and Jonathan Geuter and Georgios Mastrapas and Bo Wang and Han Xiao},
      year={2023},
      eprint={2307.11224},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}