CoolFace
Datasetpublic

stephantulkens/msmarco-mxbai-pooled

Embedpress: mixedbread large on MsMarco This is the full MsMarco corpus, embedded with Mixedbread AI's mixedbread-ai/mxbai-embed-large-v1. For each document, we take the first 510 tokens (the model's max length -2 special tokens), and embed it, not using any instructions. Because the model was trained using Matryoshka Representation Learning, these embeddings can safely be truncated. These are mainly useful for large-scale knowledge distillation. The dataset consists of 8.8… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/msmarco-mxbai-pooled.

sourceHugging Facemitupdated 1y agoView on Hugging Face
1likes232downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
stephantulkens/msmarco-mxbai-pooled · CoolFace