stephantulkens/msmarco-mxbai-pooled
Embedpress: mixedbread large on MsMarco This is the full MsMarco corpus, embedded with Mixedbread AI's mixedbread-ai/mxbai-embed-large-v1. For each document, we take the first 510 tokens (the model's max length -2 special tokens), and embed it, not using any instructions. Because the model was trained using Matryoshka Representation Learning, these embeddings can safely be truncated. These are mainly useful for large-scale knowledge distillation. The dataset consists of 8.8… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/msmarco-mxbai-pooled.
1232
